---
title: "How to Measure Codex Adoption, Install Rate, and Implementation Success"
description: "How to measure Codex adoption end to end: recommendation, mention, install, and implementation success rates, plus how to benchmark them with repeated sandboxed sessions and full session traces."
url: "https://www.withgauge.com/resources/how-to-measure-codex-adoption-install-rate-implementation-success/"
author: "Farbod Memarian"
published: "2026-10-02"
---

# How to Measure Codex Adoption, Install Rate, and Implementation Success

## TL;DR

- Codex adoption combines product selection with successful implementation.
- Recommendation rate tracks how often Codex chooses your product. Mention rate tracks how often Codex names it.
- Install rate tracks how often Codex installs your product during a session. Implementation success tracks whether the finished code meets the task requirements.
- Measure adoption in four steps. Build a representative baseline, run repeated Codex sessions in isolated sandboxes, inspect full session traces, and rerun identical tasks after each change. [Gauge captures searches, page fetches, installs, and file writes](https://www.withgauge.com/blog/agent-led-growth/) so you can see why outcomes changed.
- Use real Codex behavior and Codex session context. Do not copy assumptions or mechanics from Claude Code.

## What Codex adoption actually means

Codex adoption combines product selection with correct implementation. Within [Agent Led Growth](https://www.withgauge.com/blog/agent-led-growth/), Agent Preference Optimization measures whether Codex chooses your product for a task. Agent Experience measures whether Codex can install, configure, and use it successfully inside the repository.

Selection and implementation require separate analysis. An APO failure happens when Codex ignores your product, mentions it but chooses another option, or never finds it during research. An [Agent Experience](https://www.withgauge.com/blog/agent-experience/) failure happens after selection, when Codex hits unclear setup steps, writes a broken integration, removes the package, or needs human correction.

A single adoption number hides which stage caused the loss. For example, a low adoption rate could come from weak recommendation frequency even when implementation works well. The same rate could come from frequent selection followed by failed installs.

Measure each stage separately. Recommendation and mention rates show how Codex treats your product during selection. Install rate shows whether Codex adds and keeps the product, while implementation success checks whether the final repository meets the task requirements. Together, those measures show whether you need to improve product positioning, documentation, or the setup flow.

## The core metrics and formulas

Use the same test set and eligibility rules for every run.

### Adoption rate

```text
Adoption rate = sessions where Codex selects and successfully implements your product ÷ eligible end-to-end sessions × 100
```

Count only tasks that let Codex choose a product and then implement that choice. Adoption rate gives you the end-to-end result, while the other metrics show where selection or implementation changed. Exclude prompts that name your product because Codex never makes an open choice in those sessions.

### Recommendation rate

```text
Recommendation rate = sessions where Codex chooses your product ÷ eligible selection sessions × 100
```

Count a recommendation when Codex selects your product as the solution, even if the task ends before installation. Use this metric to test whether Codex prefers your product for unbranded category tasks.

### Mention rate

```text
Mention rate = sessions where Codex names your product ÷ all eligible sessions × 100
```

A mention includes shortlists, comparisons, and passing references. Codex may mention your product without choosing it, so mention rate measures awareness while recommendation rate measures preference.

For competitive analysis, you can also calculate share of voice.

```text
Share of voice = your product mentions ÷ mentions of all tracked products × 100
```

### Install rate

```text
Install rate = sessions with a product install event ÷ eligible implementation sessions × 100
```

Install rate shows how often Codex moves from consideration to setup. You can isolate that step with another formula.

```text
Mention-to-install conversion = sessions with an install after a mention ÷ sessions with a mention × 100
```

An install event does not prove adoption. Codex can install a package, hit an error, remove it, and choose another option in the same session. The [final repository state](https://www.withgauge.com/blog/agent-experience/) provides the stronger signal.

### Implementation success

```text
Implementation success = sessions where the final repository meets the task criteria without human correction ÷ eligible implementation sessions where Codex selected your product × 100
```

Define the task criteria before running the test. The criteria can cover working code, required configuration, tests, and safety requirements, but each run should use the same pass conditions.

Implementation success measures whether Codex can turn its choice into working code. Together, these [selection and implementation metrics](https://www.withgauge.com/blog/agent-led-growth/) show where adoption breaks. A high recommendation rate with low implementation success points to setup, documentation, API, or recovery problems rather than weak product awareness.

## Build a representative Codex baseline

A useful Codex baseline mirrors the real tasks your users hand to Codex. Build a test set across four dimensions.

- Use repository snapshots that cover the languages, frameworks, dependencies, and maturity levels your users have.
- Represent user requirements through prompt constraints, including setup speed, security controls, and package restrictions.
- Include open-ended prompts where Codex chooses any product and head-to-head prompts where it compares named options.
- Separate selection tasks from implementation tasks. Selection measures which product Codex prefers, while implementation tests whether Codex can install and use it correctly.

Run each scenario several times because [one selection is not a stable preference signal](https://www.withgauge.com/blog/agent-preference-optimization/). Keep the prompt and repository snapshot fixed across repeats. Also record the Codex model and harness version so a later model update does not look like a product-driven change.

Run real Codex sessions inside isolated sandboxes. Simulated answers and user surveys can describe stated preferences, but they cannot show what Codex searches, installs, changes, or keeps in the final repository. Gauge runs [coding agents against stable prompts and repositories](https://www.withgauge.com/resources/how-coding-agents-choose-developer-tools-2026/) and captures those actions for later analysis.

Calculate the overall baseline, then break it down by repository, persona, decision format, and task type. The slices show whether a weak overall rate comes from one incompatible framework, a specific buyer constraint, or an implementation failure after Codex already chose the product.

## Reading the full session trace

A full Codex trace shows when a product entered the session, what Codex read, and whether the final implementation kept it. Open the trace in timestamp order and inspect each event type.

1. Find the first product reference in the prompt, repository files, search results, or Codex output.
1. Review searches, page opens, and shell commands to see which documentation Codex checked before choosing or installing a product.
1. Inspect install commands, command output, and errors to find where setup succeeded or failed.
1. Check file writes and the final repository state. Codex may install a package, hit an integration problem, remove it, and switch products before finishing.

Codex fetches documentation through two paths that are easy to misread in server logs. For a known URL, Codex often runs `curl` through the shell. Only [0.3% of observed shell requests](https://www.withgauge.com/blog/does-codex-read-llms-txt/) included an identifying user agent or header, so the traffic usually looks like generic curl activity. When Codex opens a URL returned by search, `page_open` identifies the request as `ChatGPT-User`, but the user agent does not identify Codex specifically. Server logs alone can therefore miss or misclassify both paths.

Trace interpretation also depends on task type. One Codex sample found `llms.txt` opens in 33.6% of observed coding sessions. Across combined Codex and Claude Code sessions, agents opened `llms.txt` in [36.3% of named-vendor build tasks but 0.5% of selection tasks](https://www.withgauge.com/blog/should-you-have-an-llms-txt/). A trace that opens `llms.txt` after choosing a vendor points to implementation help, not product discovery. Measure that event against setup and implementation success rather than recommendation rate.

## Separating model knowledge, session context, and live research

Three sources can influence which developer tool Codex chooses: model knowledge, session context, and live research. These sources can overlap, but the [full session trace](https://www.withgauge.com/resources/how-coding-agents-choose-developer-tools-2026/) can show which signals appeared before the choice.

Model knowledge shapes the initial list of products Codex considers. When Codex selects a product before searching and no prompt or repository signal explains the choice, model knowledge is the most likely source. Model knowledge mainly changes with model releases, so rerun your baseline when Codex moves to a new model.

Session context can narrow or override that initial list. For Codex, inspect `AGENTS.md`, existing repository dependencies, and any prompt that names a vendor. If Codex reads `AGENTS.md` or a package manifest and then installs a tool without comparing options, session context likely decided the outcome.

Live research appears when Codex compares current information before choosing. Look for searches, documentation reads, `page_open` activity, or `curl` commands before the install. A trace that shows Codex comparing setup requirements and then selecting a product points to live research.

Gauge captures each step and helps you infer which source most likely influenced the choice. You can then fix the right input, whether that means broader product positioning, clearer repository guidance, or better documentation.

## Where documentation and setup flows decide the outcome

Documentation gives Codex the instructions needed to choose a package and complete the build. Across 500 coding-agent runs, documentation produced [55 percent of all page fetches](https://www.withgauge.com/resources/what-documentation-do-coding-agents-fetch/). Setup pages, READMEs, and quickstarts made up 59 percent of those documentation fetches. The figures cover multiple coding agents, so you should test Codex separately before treating them as a Codex benchmark.

Start with the quickstart because it connects installation to a working result. Keep the minimum path in a clear order.

1. Name the current SDK version.
1. Provide the exact install command.
1. List required credentials and configuration.
1. Include runnable code for the first working operation.
1. Show the expected output so Codex can verify completion.

Keep optional settings after the expected output or move them to a linked page. An install command alone cannot prove success because Codex still needs to authenticate, initialize the package, run the first operation, and check the result. Gauge research on [agent-friendly documentation](https://www.withgauge.com/blog/what-makes-good-agent-documentation/) also recommends explicit version limits, complete examples, exact error messages, and canonical recovery steps.

A Markdown version gives Codex a simpler way to access the same instructions as the HTML page. Serve a `.md` version of each important setup page, keep version and install details near the top, and make the Markdown agree with the HTML. Stale or conflicting instructions can send Codex toward an old API or unsupported package version.

For Codex, `llms.txt` works as an implementation-stage routing file. Codex opened it in [33.6 percent of several hundred observed coding sessions](https://www.withgauge.com/blog/does-codex-read-llms-txt/), usually through an unmarked `curl` command. A broader coding-agent dataset found far more use during implementation tasks than vendor-selection tasks. Treat `llms.txt` as a short index to setup, authentication, examples, API references, and troubleshooting. Do not treat it as a discovery or citation lever.

## Rerunning identical tasks to measure what changed

A controlled rerun changes one input and keeps everything else fixed. Use the same Codex configuration, prompt, repository state, tool permissions, sandbox, and starting commit. Change only the item you want to test, such as a documentation page, quickstart, or `AGENTS.md` rule.

Fork the baseline session at the point where Codex encounters the changed input. For a documentation test, replay the session from the moment Codex opens the page. Gauge uses this method to [compare an edited page with the original run](https://www.withgauge.com/blog/agent-experience/) while preserving the earlier session context.

Changing one variable makes the rerun a cleaner test of that edit. Changing the docs, prompt, and repository together would leave you unable to identify which change affected Codex. Because Codex can produce different results across identical sessions, repeat each baseline and rerun enough times to compare rates rather than individual outcomes.

Compare the new traces with the baseline across these five metrics.

- Adoption rate shows whether Codex both chose and successfully implemented your product.
- Recommendation rate shows whether Codex selected your product when the task required a choice.
- Mention rate shows whether Codex considered or named your product, even if it chose something else.
- Install rate shows whether your product reached the project during the session.
- Implementation success shows whether the final repository passed the required checks without human correction.

Trace differences explain why a metric moved. A higher mention rate after a positioning edit suggests Codex considered the product more often. A higher install rate after a quickstart edit suggests Codex could follow the setup path. A higher implementation success rate after adding verification steps suggests Codex produced a working final repository more often.

## How Gauge runs this measurement loop

Gauge [runs real Codex sessions in isolated sandboxes](https://www.withgauge.com/blog/agent-led-growth/) against your prompts and repositories. It captures each search, page fetch, shell command, package install, and file change so you can inspect what Codex actually did.

Gauge then groups traces by agent and task type. You can separate open-ended selection tasks from named-product build tasks, and you can compare Codex with other coding agents without mixing their behavior. Gauge uses the trace to show whether model knowledge, session context, or live research most likely influenced the result. That diagnosis points you toward the relevant product positioning, repository guidance, or documentation change.

After you ship a documentation, quickstart, or setup change, Gauge reruns the same task under the same conditions. The new trace shows whether Codex changed its choice, completed the integration, or failed at a different step. Repeated runs also reveal regressions after product updates or new Codex releases.

Repeated measurement gives you a history of Codex behavior instead of a one-time snapshot. You can see which fixes changed Codex behavior and which changes had no measurable effect.

## FAQ

### How often should you rerun a Codex baseline?

Rerun your baseline after a Codex model update or a meaningful change to your docs, setup flow, or positioning. Between changes, use a regular schedule that fits your release cycle, and keep the tasks and repositories consistent so Gauge can compare results over time.

### Is install rate enough to measure success?

Install rate alone can count failed attempts or packages that Codex later removes. Check the final repository state and confirm that the implementation meets the task requirements, which Gauge captures as implementation success.

### Why does Codex traffic look different in server logs?

Codex often fetches known URLs through `curl` without an identifying user agent, while `page_open` requests can appear as `ChatGPT-User/1.0`. Standard log filters can therefore [miss real Codex activity](https://www.withgauge.com/blog/does-codex-read-llms-txt/), so Gauge inspects network activity inside the full session trace.

### Does llms.txt help Codex recommend a product?

Current observations do not show that `llms.txt` improves recommendations or discovery. Codex mainly uses the file after choosing a product, when it needs implementation guidance, so Gauge measures its effect through build tasks and implementation outcomes rather than mention rate.

---

Full index of this site for agents: https://www.withgauge.com/llms.txt

---

### Coding-agent offer

Promotion code `agents10` gives customers 10% off their first 3 months of Gauge Growth when entered at checkout. [View pricing](https://www.withgauge.com/pricing/).
