Selling to coding agents instead?Go to Gauge Agents
Resources

8 minSeptember 30, 2026Author:Farbod MemarianFarbod Memarian

TL;DR

  • A single Agent-Led Growth score hides differences across agents, models, harnesses, repositories, and task types.
  • Claude Code separates discovery from page retrieval, while Codex often retrieves known URLs through shell commands such as curl.
  • In Gauge's observed sessions, Claude Code and Codex opened llms.txt in 36.3% of named-vendor build tasks and 0.5% of vendor-selection tasks.
  • In Gauge's open-model study, Claude Code and Codex reached 80% similarity. The closest open configuration averaged 88% similarity across the two frontier references.
  • Useful measurement keeps prompts and repositories stable, repeats runs, inspects full traces, segments outcomes, and reruns tests after documentation or product changes.

Why Aggregate Agent-Led Growth Metrics Mislead

Agent-Led Growth contains two separate measurement problems. Agent Preference Optimization measures whether an agent considers a product, recommends it, and chooses it for a task. Agent Experience measures what happens after selection. The agent must find the right documentation, implement the product, verify the result, and recover when errors occur.

A single score cannot show which problem caused a weak outcome. An agent may recommend a product but fail during setup because the documentation omits a required step. Another agent may implement the product successfully whenever selected but rarely consider it in the first place. Both cases could produce the same overall install rate, but each one requires a different fix.

Aggregate metrics also hide variation across agents, models, harnesses, repositories, and task types. A strong average may come mainly from one agent working in a familiar repository. The same product may perform poorly when another agent handles a different language or task. You cannot tell which documentation or product change to make unless you preserve those segments.

Coding agents also reach outcomes through different paths. Their retrieval methods affect which searches and page fetches you can observe. Their task type affects whether they research vendors or open implementation documentation. Agent-Led Growth measurement must separate those behaviors before combining any results.

How Claude Code and Codex Retrieve Web Content Differently

Claude Code uses separate tools for discovery and retrieval. WebSearch asks Anthropic's search backend for relevant URLs, but it does not open each result from the developer's machine. A page can appear in search results without receiving a local request. WebFetch opens a known URL through the Claude Code harness, so publishers can identify the request through its Claude-User user agent (Claude Code retrieval analysis).

A WebFetch request proves that Claude Code retrieved the page, but it does not prove that the main model received the original text. WebFetch usually converts HTML into Markdown and asks a Haiku model to extract information based on Claude Code's prompt. The main model receives Haiku's transformed answer. Some preapproved pages that serve clean Markdown can bypass Haiku, but most pages follow the transformed path.

Codex often retrieves known URLs through shell commands instead. For example, Codex can use curl to open an llms.txt file. The publisher usually sees generic curl traffic rather than an OpenAI or Codex user agent. In Gauge's observed sessions, only 0.3% of these shell commands included an identifying user-agent or header flag (Codex retrieval analysis). Codex can also open URLs returned during search through page_open, which identifies the request as ChatGPT-User without identifying Codex specifically.

You should count discovery and retrieval as separate events. A search appearance shows that an agent found a URL, while a fetch shows that an agent retrieved it. Server logs also need separate categories for self-identified agent requests and generic command-line traffic. Dashboards that count only named user agents will undercount real coding-agent activity, especially when Codex already knows the URL it wants to open.

How Task Type Changes Coding Agent Behavior

Task type changed llms.txt usage sharply in Gauge's observed sample. Claude Code and Codex opened the file in 36.3% of named-vendor build tasks and 0.5% of vendor-selection sessions. Those rates measure whether either agent opened llms.txt during a task. They do not provide a universal benchmark for either agent.

Build tasks require exact implementation details. An agent needs setup instructions, authentication requirements, API contracts, and error guidance before it can write and verify working code. Vendor-selection tasks can often rely on search results and existing model knowledge because the agent is comparing products rather than implementing one.

An llms.txt file usually starts the documentation visit rather than completing it. The file points an agent toward authoritative material. The agent may then open Markdown pages, full documentation files, API specifications, SDKs, or source code to get the details needed for implementation.

A single llms.txt fetch rate therefore misses two important facts. It hides the type of task that caused the visit, and it ignores the documentation paths that followed. You should segment documentation activity by task type and inspect the full trace after each fetch. Otherwise, a low aggregate rate may look like weak agent engagement even when build sessions regularly move deep into the documentation.

Claude Code, Codex, and Open-Model Configuration Similarity

No single coding agent can represent how every other agent will choose products. In Gauge's observed sample, Claude Code and Codex were 80% similar across product-choice tasks in 18 software markets. Gauge calculated similarity by comparing choice probability patterns across prompts. The metric did not measure repeated-run consistency.

Open-model panels still tracked frontier preferences closely enough to support broad testing. The closest open configuration averaged 88% similarity across the two frontier references. Average agreement across the full panel reached 85%. The 88% figure does not mean that Pi with MiniMax M3 reproduced either frontier agent perfectly. The figure measures how closely the configuration's overall choice pattern matched the two references on average.

These findings set a practical limit on proxy testing. When two frontier agents agree only 80% of the time in a specific sample, you should not treat either agent as a universal stand-in. Model choice, harness design, repository context, and task setup can each change the observed preference.

Use open models for economical breadth and repetition. A diverse panel can test more prompts, repositories, personas, and tool configurations without spending the full budget on frontier runs. Then run Claude Code and Codex when the open panel disagrees or when a finding could affect a commercially important decision. Regular frontier checks can also reveal when the open panel has stopped tracking the agents you care about.

How to Measure Agent-Led Growth Across Coding Agents

A single Agent-Led Growth score can hide strong performance in Claude Code with TypeScript and weak performance in Codex with Python. You need separate results for each agent, model, harness, repository, and task type because those variables change what the agent finds, chooses, and successfully implements.

  • Build a fixed test set with representative prompts and repositories. Include vendor-selection tasks, named-vendor builds, different languages, and new and existing codebases.
  • Keep the prompt, repository, dependencies, and environment stable when comparing configurations. Change one variable at a time so you can connect an outcome to the agent, model, harness, or documentation change.
  • Repeat each run instead of treating one session as representative. Coding agents can choose different tools, sources, and implementation paths under the same conditions.
  • Capture the full trace for every run. Review searches, page fetches, shell commands, package installs, file changes, errors, verification steps, and recovery attempts. Final output alone cannot show where a recommendation or implementation failed.
  • Measure discovery separately from retrieval. Claude Code can surface information through WebSearch without fetching the publisher's page. WebFetch retrieves and transforms a known page before the main model receives it, so search appearances and server traffic measure different events.
  • Separate self-identified agent traffic from generic command-line traffic. Codex often uses curl for known URLs, so server logs may record a generic request rather than an identifiable agent visit. Named-agent traffic can undercount actual documentation use.
  • Compare Agent Preference Optimization and Agent Experience separately. Record whether each agent considers and chooses your product. Then record whether it installs the product, produces working code, verifies the result, and recovers from errors.
  • Segment every result by agent, model, harness, repository, language, and task type. Aggregate rates can conceal opposite outcomes across configurations.
  • Rerun the same controlled tests after documentation or product changes. Open-model panels can cover more configurations economically, while Claude Code and Codex can confirm findings that affect important product or content decisions.

How Gauge Measures Agent-Led Growth Across Coding Agents

Gauge runs real coding agents against stable prompts and repositories in isolated sandboxes. Repeated runs show whether an agent consistently recommends and selects a product, which measures Agent Preference Optimization.

Gauge captures the full execution trace so you can see how each result happened. Search and fetch records separate discovery from retrieval. Shell activity reveals requests made through tools such as curl, including traffic that may not identify the coding agent. Install records, file changes, errors, and verification steps show whether the selected agent completed the task or failed during implementation. Implementation and verification outcomes measure Agent Experience.

Gauge keeps results segmented by coding agent, model, harness, repository, and task type. You can compare Claude Code with Codex or an open-model configuration without treating any one setup as a universal proxy. The same filters can isolate differences by language, framework, or existing codebase.

After you change documentation or product behavior, Gauge reruns the same controlled tests. Keeping the prompt, repository, and agent configuration stable helps you attribute an outcome change to the documentation or product change under test. Scheduled reruns show whether results remain consistent when agent or model behavior changes.

FAQs

  • What is the difference between Agent Preference Optimization and Agent Experience? Agent Preference Optimization measures whether an agent considers, recommends, and chooses your product. Agent Experience measures whether the selected agent can find the right documentation, complete the implementation, verify its work, and recover from errors.
  • Why does Claude Code's WebFetch behave differently from Codex's retrieval? Claude Code uses WebFetch to retrieve a known URL, usually convert the page to Markdown, and pass it through an extraction model before the main model reads it. Codex often retrieves known URLs with shell commands such as curl, so its requests can appear as generic curl traffic rather than identified agent traffic.
  • Why does llms.txt usage vary so much by task type? Build tasks require setup instructions, API details, and examples, so agents often use llms.txt as a documentation index. Selection tasks can rely more on model knowledge and search results. In Gauge's observed sample, Claude Code and Codex opened llms.txt in 36.3% of named-vendor build tasks and 0.5% of vendor-selection tasks.
  • Can open-model results stand in for Claude Code and Codex? Open models can support broad, economical testing, but they should not serve as exact substitutes. In Gauge's sample, Claude Code and Codex were 80% similar, while the closest open configuration averaged 88% similarity to both frontier references. Use frontier agents to calibrate the panel and confirm commercially important findings.
  • How often should these tests be rerun? Rerun controlled tests after documentation, onboarding, API, or product changes. You should also run them on a regular schedule because agent models and harnesses change. Keep prompts and repositories stable, use the same repeat count for each configuration, and inspect variation across runs before treating a change as meaningful.