TL;DR
- Agent Preference Optimization measures whether coding agents recommend and choose your product.
- Agent Experience measures whether agents can understand and implement a product, verify the result, and recover from errors without human help.
- Agent-Led Growth follows a recurring loop. You benchmark behavior, diagnose failures, make fixes, and rerun the same tasks.
- Gauge runs real coding agents in isolated environments and compares preference and implementation results by agent.
What Is Agent-Led Growth?
Agent-Led Growth tracks which products coding agents recommend and how successfully they implement them. Coding agents often work in the terminal rather than a search box. SEO rankings, website visits, and AI citations therefore miss decisions and implementation work that occur inside an agent session.
A developer might ask an agent to add email sending without naming a vendor. The agent can research providers, choose one, and complete the integration. The developer may never visit the vendor's site, and standard web analytics may not capture the agent's research or implementation. ALG measurement observes the agent's decisions and actions directly.
Agent Preference Optimization measures whether agents consider, recommend, and choose your product. If agents consistently exclude it from representative tasks, the product has an awareness or preference problem. Agent Experience measures whether agents can understand the product, complete setup, build and verify a working integration, and recover from errors. Poor AX can cause an agent to abandon the product before the developer reviews the code.
Essential Metrics for Agent-Led Growth
The useful Agent-Led Growth metrics track concrete agent outcomes. Each metric should show whether an agent prefers your product or successfully completes work with it. Ignore signals that do not explain agent preference or successful implementation.
Agent Preference Optimization metrics
Agent Preference Optimization measures revealed behavior rather than asking an agent which tool it prefers. An agent may rely on training knowledge, instructions in the prompt, or repository context. It may conduct live research when those inputs do not settle the choice.
Consideration rate measures the percentage of representative runs in which a product enters the candidate set. Count an explicit mention or research behavior that clearly identifies the product as a candidate. Share of voice measures how often the agent considers your product relative to competitors. Both metrics show whether your product reaches the decision stage, but neither proves that the agent chose it.
Selection or win rate measures how often the agent chooses your product in eligible tasks. A useful benchmark repeats consistent unbranded tasks or head-to-head comparisons because a single run may not represent typical agent behavior. Comparing selection rate with consideration rate also reveals whether agents notice the product but reject it later.
Break down preference metrics by agent and test environment, including the model and repository. Different models carry different prior knowledge, while repository language, installed packages, and rules can change the choice. During 500 coding-agent runs observed by Gauge, documentation received 55% of page fetches, source code received 18%, package registries received 11%, and third-party content received 5%. Use session traces to identify the evidence available before each decision, including repository context and fetched pages. Existing model knowledge cannot always be observed directly, so infer it cautiously when the trace contains no external source.
Agent Experience metrics
Documentation retrieval success measures whether an agent can access relevant pages and locate the instructions needed for the task. Record failed requests and whether missing or irrelevant instructions prevent the agent from completing the task. Onboarding completion measures whether the agent can finish the setup steps available in the test environment without human help, such as configuring a project or using provided credentials.
Integration success rate measures whether the agent produces working code that passes predefined acceptance checks. Install counts cannot prove success because an agent may install a package and remove it later in the same session. The final repository state and a verified result provide stronger evidence that the product worked.
Recovery rate measures how often an agent reaches a working result after an error. Record the original error and the agent's recovery attempts, then confirm whether the final result passes the acceptance checks. Human intervention rate measures how often a person must unblock the agent. Replacement rate measures how often the agent abandons one product for another. High replacement rates can reveal unclear errors, stale examples, or missing recovery instructions.
Treat access to llms.txt as a supporting diagnostic for implementation tasks, not as a success metric by itself. In Gauge's observed coding-agent sessions, agents opened llms.txt in 36.3% of implementation tasks and 0.5% of vendor-selection tasks. The file helps agents locate technical instructions after a product enters the build. It provides little evidence that the file caused the agent to choose that product.
Benchmark these metrics with fixed tasks and success criteria. After changing documentation or onboarding, rerun the same agent, model, prompt, and repository to see whether more sessions reach verified working code without human help.
How Gauge Measures Agent-Led Growth
Gauge measures Agent-Led Growth by running real coding agents against consistent prompts and repositories in isolated sandboxes. Gauge records the agent's research and repository changes alongside the final result. You can compare recommendation and implementation results across agents and test environments rather than combining materially different runs into one score.
The measurement loop starts with a benchmark. Full session traces show which observed sources and actions preceded the choice. They also identify the implementation step where the agent failed, recovered, or switched products. You can fix the specific documentation or product issue and rerun the same test to see whether agent behavior changed.
Open models can reduce the cost of broad test coverage. Frontier agents can then validate findings that are unclear or important to a product decision. In Gauge's proxy study, the open-model panel averaged 85 percent similarity to the frontier references. The closest configuration averaged 88 percent. Those figures describe Gauge's tested sample, not a universal benchmark, so regular recalibration still matters.
Gauge can also fork a recorded session at a specific fetch or failure. You can swap one page or fix one step while keeping the model, prompt, and repository constant. The replay tests whether the controlled change alters the agent's choice or implementation outcome.
FAQ
How often should you rerun Agent-Led Growth metrics?
Run a stable benchmark at a consistent interval and after a material change to the agent, model, documentation, or product. Keep the same prompts and repositories so changes in performance remain comparable.
Does llms.txt help an agent choose or build with a product?
llms.txt mainly helps agents build with a product they already know. In Gauge's observed Claude Code and Codex sessions, agents opened llms.txt in 36.3% of named-vendor build tasks and 0.5% of vendor-selection tasks.
Can you trust install rate alone?
No. An agent can install a package, fail to implement it, and remove it before finishing. Check whether the final repository passes the task's acceptance criteria. Use recovery, replacement, and human intervention rates to diagnose how the agent reached that result.
Related Resources
How Agent-Led Growth Differs Across Claude Code, Codex, and Open Source Coding Agents
A single Agent-Led Growth score hides how Claude Code, Codex, and open-model agents retrieve content, use llms.txt, and choose products differently. How to segment and measure ALG across coding agents.
Farbod MemarianAgent-Led Growth Strategy: How to Build a Repeatable Growth Loop
Agent-Led Growth means winning two moments in one coding-agent session: getting chosen and getting implemented. How to diagnose why you win or lose with full traces, and the repeatable loop for fixing it.
Farbod MemarianBest Agent Discoverability Tools and Platforms in 2026
Agent discoverability (Agent Preference Optimization) measures whether coding agents discover, recommend, select, and install your product. What to look for in a platform, and the best tool for measuring it in 2026.
Farbod Memarian