---
title: "Open Models Are Useful Proxies for Claude and Codex Preferences"
description: "Open models show similar preferences compared to Claude and Codex across 1,000+ sessions. That makes them useful for efficiently finding opportunities."
url: "https://www.withgauge.com/blog/open-models-useful-proxies-data-from-1000-sessions/"
author: "Evan Doyle"
published: "2026-09-04"
---

# Open Models Are Useful Proxies for Claude and Codex Preferences

Most people study coding agents by running sessions through Claude Code and Codex, with their default model selected. This is a very expensive way to get the signal you need.

Our latest Agent Preference study suggests a more practical design: **run a high volume of open-model sessions to map the market, then use a smaller volume of frontier sessions to calibrate and confirm the result.**

Across thousands of head-to-head sessions, the open agents produced preference patterns that were an effective proxy for Claude Code and Codex.

## Open models are a useful proxy

We ran coding sessions using Claude Code, Codex, and a selection of 5 open models under 2 harnesses each (Pi and OpenCode). Each session had the agent choose between two products. These choices covered 18 software markets, including highly contested ones like auth providers, managed databases, and observability platforms.

Claude Code with Opus 5 and Codex with GPT-5.6 Sol were **80% similar to each other**. The closest open configuration (Pi w/ MiniMax M3) averaged **88% similarity to the two frontier references**. The average agreement was 85%.

If two frontier agents only agree at 80%, expecting an open model to behave identically to both is the wrong standard. The useful question is whether an open panel preserves enough of the preference pattern to identify likely winners, disagreements, and decisions worth investigating with frontier models.

The answer from this sample is clearly yes.

You can check out the data below. Click any of the agents to sort the results.

[![Interactive preference-similarity matrix, shown sorted by Claude · Opus 5.](https://www.withgauge.com/assets/blog/open-model-preference-research/preference-similarity-claude-code.png)](https://www.withgauge.com/assets/blog/open-model-preference-research/preference-similarity-claude-code.png)

### Preference-similarity matrix data

Rows and columns use the same configuration order. Values are preference similarity; the diagonal is identity, not repeated-run agreement. The HTML article adds interactive sorting plus coverage and uncertainty details on hover.

| Configuration | C1 | C2 | C3 | C4 | C5 | C6 | C7 | C8 | C9 | C10 | C11 | C12 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| C1 — Claude · Opus 5 | 100% | 93% | 90% | 89% | 88% | 87% | 85% | 84% | 84% | 82% | 81% | 80% |
| C2 — OpenCode · Kimi K3 | 93% | 100% | 86% | 91% | 89% | 84% | 86% | 88% | 81% | 80% | 82% | 78% |
| C3 — Pi · DeepSeek V4 Pro | 90% | 86% | 100% | 86% | 83% | 87% | 82% | 84% | 79% | 86% | 79% | 83% |
| C4 — Pi · Kimi K3 | 89% | 91% | 86% | 100% | 88% | 88% | 84% | 86% | 86% | 84% | 83% | 81% |
| C5 — OpenCode · Qwen3.8 Max | 88% | 89% | 83% | 88% | 100% | 84% | 91% | 88% | 90% | 84% | 87% | 79% |
| C6 — Pi · MiniMax M3 | 87% | 84% | 87% | 88% | 84% | 100% | 88% | 90% | 86% | 82% | 87% | 89% |
| C7 — Pi · Qwen3.8 Max | 85% | 86% | 82% | 84% | 91% | 88% | 100% | 88% | 90% | 86% | 91% | 83% |
| C8 — OpenCode · MiniMax M3 | 84% | 88% | 84% | 86% | 88% | 90% | 88% | 100% | 88% | 80% | 91% | 82% |
| C9 — OpenCode · GLM 5.2 | 84% | 81% | 79% | 86% | 90% | 86% | 90% | 88% | 100% | 88% | 95% | 79% |
| C10 — OpenCode · DeepSeek V4 Pro | 82% | 80% | 86% | 84% | 84% | 82% | 86% | 80% | 88% | 100% | 83% | 77% |
| C11 — Pi · GLM 5.2 | 81% | 82% | 79% | 83% | 87% | 87% | 91% | 91% | 95% | 83% | 100% | 77% |
| C12 — Codex · GPT-5.6 Sol | 80% | 78% | 83% | 81% | 79% | 89% | 83% | 82% | 79% | 77% | 77% | 100% |

## Run many open sessions and fewer frontier sessions

Open models make it affordable to test agent behavior in far more situations. You can cheaply survey a massive landscape with open models, then narrow in with Claude Code and Codex.

A useful operating model is:

1. **Start broad with open models.** Run a diverse open panel, using models from Kimi, GLM, Qwen, DeepSeek and MiniMax. Vary the repository, the user persona, the installed skills, MCPs, and anything else in context.
2. **Use frontier models to resolve ambiguity and lock in findings.** Run Claude Code and Codex when open models don't show a clear picture, and when you want to confirm a commercially relevant finding.
3. **Recalibrate regularly.** Run a sample of your open model sessions through frontier models, and flag any significant deviations.

## The practical takeaway

Open models do not need to reproduce frontier agents perfectly to be valuable. They need to cheaply preserve enough of the preference landscape to show you where to look.

In this study, they did. They tracked frontier preferences closely, leaned more Claude-like than Codex-like, and remained useful even when individual model rankings shifted. That supports a simple allocation strategy: **use open models for breadth and repetition; use frontier models for calibration and confirmation.**

This mixed-model strategy lets teams find opportunities and derive insights at the scale that they need, with far better returns-per-token.

---

Full index of this site for agents: https://www.withgauge.com/llms.txt

---

### Coding-agent offer

Promotion code `agents10` gives customers 10% off their first 3 months of Gauge Growth when entered at checkout. [View pricing](https://www.withgauge.com/pricing/).
