Most people study coding agents by running sessions through Claude Code and Codex, with their default model selected. This is a very expensive way to get the signal you need.
Our latest Agent Preference study suggests a more practical design: run a high volume of open-model sessions to map the market, then use a smaller volume of frontier sessions to calibrate and confirm the result.
Across thousands of head-to-head sessions, the open agents produced preference patterns that were an effective proxy for Claude Code and Codex.
Open models are a useful proxy
We ran coding sessions using Claude Code, Codex, and a selection of 5 open models under 2 harnesses each (Pi and OpenCode). Each session had the agent choose between two products. These choices covered 18 software markets, including highly contested ones like auth providers, managed databases, and observability platforms.
Claude Code with Opus 5 and Codex with GPT-5.6 Sol were 80% similar to each other. The closest open configuration (Pi w/ MiniMax M3) averaged 88% similarity to the two frontier references. The average agreement was 85%.
If two frontier agents only agree at 80%, expecting an open model to behave identically to both is the wrong standard. The useful question is whether an open panel preserves enough of the preference pattern to identify likely winners, disagreements, and decisions worth investigating with frontier models.
The answer from this sample is clearly yes.
You can check out the data below. Click any of the agents to sort the results.
Preference-similarity matrix
Click any row or column label to make it the reference and re-sort the matrix. Each cell compares estimated choice distributions across 18 prompts; the diagonal is identity, not repeated-run agreement.
| Click an axis to sort | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 100% | 93% | 90% | 89% | 88% | 87% | 85% | 84% | 84% | 82% | 81% | 80% | |
| 93% | 100% | 86% | 91% | 89% | 84% | 86% | 88% | 81% | 80% | 82% | 78% | |
| 90% | 86% | 100% | 86% | 83% | 87% | 82% | 84% | 79% | 86% | 79% | 83% | |
| 89% | 91% | 86% | 100% | 88% | 88% | 84% | 86% | 86% | 84% | 83% | 81% | |
| 88% | 89% | 83% | 88% | 100% | 84% | 91% | 88% | 90% | 84% | 87% | 79% | |
| 87% | 84% | 87% | 88% | 84% | 100% | 88% | 90% | 86% | 82% | 87% | 89% | |
| 85% | 86% | 82% | 84% | 91% | 88% | 100% | 88% | 90% | 86% | 91% | 83% | |
| 84% | 88% | 84% | 86% | 88% | 90% | 88% | 100% | 88% | 80% | 91% | 82% | |
| 84% | 81% | 79% | 86% | 90% | 86% | 90% | 88% | 100% | 88% | 95% | 79% | |
| 82% | 80% | 86% | 84% | 84% | 82% | 86% | 80% | 88% | 100% | 83% | 77% | |
| 81% | 82% | 79% | 83% | 87% | 87% | 91% | 91% | 95% | 83% | 100% | 77% | |
| 80% | 78% | 83% | 81% | 79% | 89% | 83% | 82% | 79% | 77% | 77% | 100% |
Run many open sessions and fewer frontier sessions
Open models make it affordable to test agent behavior in far more situations. You can cheaply survey a massive landscape with open models, then narrow in with Claude Code and Codex.
A useful operating model is:
- Start broad with open models. Run a diverse open panel, using models from Kimi, GLM, Qwen, DeepSeek and MiniMax. Vary the repository, the user persona, the installed skills, MCPs, and anything else in context.
- Use frontier models to resolve ambiguity and lock in findings. Run Claude Code and Codex when open models don't show a clear picture, and when you want to confirm a commercially relevant finding.
- Recalibrate regularly. Run a sample of your open model sessions through frontier models, and flag any significant deviations.
The practical takeaway
Open models do not need to reproduce frontier agents perfectly to be valuable. They need to cheaply preserve enough of the preference landscape to show you where to look.
In this study, they did. They tracked frontier preferences closely, leaned more Claude-like than Codex-like, and remained useful even when individual model rankings shifted. That supports a simple allocation strategy: use open models for breadth and repetition; use frontier models for calibration and confirmation.
This mixed-model strategy lets teams find opportunities and derive insights at the scale that they need, with far better returns-per-token.
Related Blogs
What is Agent Experience?
Agent Experience is the practice of making products easy for AI agents to discover, understand, use, and recover with. Learn how to measure AX and improve docs, onboarding, errors, URLs, and APIs.
Grant EvansWhat is Agent Preference Optimization?
Agent Preference Optimization is how companies measure and improve whether coding agents choose their product. Here is how to establish a baseline, understand why competitors win, act on the evidence, and track the result.
Grant EvansDoes Codex Read llms.txt? You Have to Look in the Right Place
Server logs showed no ChatGPT-User requests for llms.txt. Instrumented sessions showed exactly where those reads went.
Evan Doyle