Selling to coding agents instead?Go to Gauge Agents
Blogs

3 minSeptember 4, 2026Author:Evan DoyleEvan Doyle
Open Models Are Useful Proxies for Claude and Codex Preferences

Most people study coding agents by running sessions through Claude Code and Codex, with their default model selected. This is a very expensive way to get the signal you need.

Our latest Agent Preference study suggests a more practical design: run a high volume of open-model sessions to map the market, then use a smaller volume of frontier sessions to calibrate and confirm the result.

Across thousands of head-to-head sessions, the open agents produced preference patterns that were an effective proxy for Claude Code and Codex.

Open models are a useful proxy

We ran coding sessions using Claude Code, Codex, and a selection of 5 open models under 2 harnesses each (Pi and OpenCode). Each session had the agent choose between two products. These choices covered 18 software markets, including highly contested ones like auth providers, managed databases, and observability platforms.

Claude Code with Opus 5 and Codex with GPT-5.6 Sol were 80% similar to each other. The closest open configuration (Pi w/ MiniMax M3) averaged 88% similarity to the two frontier references. The average agreement was 85%.

If two frontier agents only agree at 80%, expecting an open model to behave identically to both is the wrong standard. The useful question is whether an open panel preserves enough of the preference pattern to identify likely winners, disagreements, and decisions worth investigating with frontier models.

The answer from this sample is clearly yes.

You can check out the data below. Click any of the agents to sort the results.

Explore pairwise results

Preference-similarity matrix

Reference Claude · Opus 5

Click any row or column label to make it the reference and re-sort the matrix. Each cell compares estimated choice distributions across 18 prompts; the diagonal is identity, not repeated-run agreement.

Click an axis to sort
100%93%90%89%88%87%85%84%84%82%81%80%
93%100%86%91%89%84%86%88%81%80%82%78%
90%86%100%86%83%87%82%84%79%86%79%83%
89%91%86%100%88%88%84%86%86%84%83%81%
88%89%83%88%100%84%91%88%90%84%87%79%
87%84%87%88%84%100%88%90%86%82%87%89%
85%86%82%84%91%88%100%88%90%86%91%83%
84%88%84%86%88%90%88%100%88%80%91%82%
84%81%79%86%90%86%90%88%100%88%95%79%
82%80%86%84%84%82%86%80%88%100%83%77%
81%82%79%83%87%87%91%91%95%83%100%77%
80%78%83%81%79%89%83%82%79%77%77%100%
76% lower96% higher
Similarity is one minus the mean absolute gap between Jeffreys-smoothed choice probabilities, with prompts weighted equally. Hover a cell for coverage and uncertainty.

Run many open sessions and fewer frontier sessions

Open models make it affordable to test agent behavior in far more situations. You can cheaply survey a massive landscape with open models, then narrow in with Claude Code and Codex.

A useful operating model is:

  1. Start broad with open models. Run a diverse open panel, using models from Kimi, GLM, Qwen, DeepSeek and MiniMax. Vary the repository, the user persona, the installed skills, MCPs, and anything else in context.
  2. Use frontier models to resolve ambiguity and lock in findings. Run Claude Code and Codex when open models don't show a clear picture, and when you want to confirm a commercially relevant finding.
  3. Recalibrate regularly. Run a sample of your open model sessions through frontier models, and flag any significant deviations.

The practical takeaway

Open models do not need to reproduce frontier agents perfectly to be valuable. They need to cheaply preserve enough of the preference landscape to show you where to look.

In this study, they did. They tracked frontier preferences closely, leaned more Claude-like than Codex-like, and remained useful even when individual model rankings shifted. That supports a simple allocation strategy: use open models for breadth and repetition; use frontier models for calibration and confirmation.

This mixed-model strategy lets teams find opportunities and derive insights at the scale that they need, with far better returns-per-token.