Selling to coding agents instead?Go to Gauge Agents
Resources

8 minSeptember 1, 2026Author:Farbod MemarianFarbod Memarian

Introduction

Coding agents now pick products on a developer's behalf, without a person comparing options or making the final call. Agent Preference Optimization measures and improves how often a coding agent chooses your product when it makes that decision.

The product an agent picks gets installed. Every alternative disappears with no impression, click, or lost-deal record to show for it. You need a tool that can observe those decisions under realistic conditions and tell you what to fix. Gauge is the tool built for that job, and this list covers it.

What is Agent Preference Optimization (APO)?

Agent Preference Optimization (APO) measures and improves how often a coding agent chooses your product while completing a task for a developer. APO measures actual behavior rather than asking a model which brands it prefers. You give coding agents realistic tasks, repositories, and constraints, then record which products they consider and install.

APO focuses on a product choice that an agent acts on directly. When Claude Code or Codex selects a provider, installs its package, and writes the integration, the developer may never see the alternatives that lost.

Model knowledge shapes the first stage of an agent’s choice. Each model starts with beliefs learned during training, including which products fit a category and which tools commonly work with a language or framework. A model may choose a familiar provider without searching, or it may rely on outdated information about another product.

Session context shapes the choice around the specific task. The prompt, repository, installed packages, rules files, and saved preferences can favor one provider before the agent compares any options. For example, an existing dependency may win because the agent can use it with fewer code changes.

Live research happens only when model knowledge and session context do not settle the choice. The agent may run task-specific searches, open documentation, inspect package registries, and compare compatibility before selecting a provider. AEO can help a product get found during that research, but research is not guaranteed. APO covers the full choice, including decisions made before an agent searches the web.

What Makes a Good Agent Preference Optimization (APO) Tool?

A good APO baseline represents real developer decisions. The benchmark should use realistic repositories, developer personas, and task formats. It should include open-ended tasks and structured head-to-head comparisons across multiple coding agents and models. Repeated runs with the same prompt and repository help separate stable preferences from one-off choices.

A good APO tool records the full session trace. A final selection tells you who won, but it does not explain why. The tool should show whether the agent researched the task, which searches it ran, which sources it opened, and how it described each product. Full traces also reveal whether the agent considered your product and rejected it or never mentioned it at all.

A good APO tool reruns benchmarks on a schedule. Model releases can change what an agent already knows, while new documentation can change what the agent finds during research. Scheduled runs let you compare the same scenarios over time and see whether preference changes after a product, content, or documentation update. The benchmark should preserve its core prompts and repositories so each comparison remains meaningful.

Top Tools for Agent Preference Optimization

Gauge

Best for

Companies that want to measure and improve how often coding agents choose their products.

What it is

Gauge offers an Agent Led Growth product built specifically for Agent Preference Optimization. Gauge runs real coding agents such as Claude Code and Codex against real repositories inside isolated sandboxes. Each agent receives a coding task, reads the repository, researches when needed, selects a product, and attempts the implementation.

Gauge can vary the repository, persona, prompt, agent, and model. You can run open-ended tasks without a named vendor or structured comparisons between selected products. Repeated runs help separate stable preferences from one-off choices.

Gauge records the full session rather than reporting only the winner. Each trace includes searches, fetched pages, installed packages, file writes, extracted decisions, and implementation outcomes. You can inspect whether an agent chose from existing model knowledge, relied on repository context, or researched the category before deciding.

The reporting layer turns those runs into measurable preference data. Brand rankings show install rate, mention rate, and rank movement. Per-agent pick rates reveal when Claude Code and Codex treat the same product differently. Filters let you narrow results by persona, prompt, and time window.

Gauge also supports recurring schedules. You can bundle prompts, select agents and repository targets, and rerun the same benchmark after a model release or documentation update. Repeated gaps can become tracked action items, which gives you a record of what needs fixing and whether the next run changed agent behavior.

Pros

  • Gauge tests actual agent behavior against real repositories instead of asking a model which product it prefers.

  • Full traces connect each choice to the searches, sources, packages, code changes, and outcomes that produced it.

  • Repository, persona, prompt, and agent controls support a representative baseline rather than a single generic test.

  • Scheduled runs show how pick rates change after documentation edits, product updates, or model releases.

  • Brand rankings and per-agent comparisons make preference patterns easier to inspect across a larger set of runs.

Cons

  • Gauge focuses on coding-agent decisions, so it does not serve companies whose products cannot be selected or installed through a software development task.

  • Useful results depend on realistic prompts, repositories, and personas. A poorly chosen benchmark can measure scenarios that do not match customer behavior.

How to Improve Agent Preference with Gauge

Gauge turns APO into a four-step operating loop. You repeat the same core scenarios over time so you can compare agent behavior after product updates, documentation changes, model releases, and competitor moves.

1. Establish a representative baseline. Gauge runs Claude Code, Codex, and other coding agents against real repositories in isolated sandboxes. You can vary repositories, developer personas, decision formats, and agents. Repeated runs help separate stable preferences from one-off choices.

2. Find out why agents choose each product. Gauge records the full session trace rather than stopping at the winner. You can inspect searches, fetched pages, installed packages, and file writes. The trace also shows how the agent described each product and what happened during implementation.

3. Fix what repeated runs reveal. Gauge turns recurring gaps into tracked action items. You might update documentation that omits a needed integration, create content for a task-shaped search, or correct framing that places your product in the wrong category. Rerunning the same scenario shows whether the targeted change affected the agent’s choice.

4. Keep measuring preference. Gauge schedules the benchmark to run again across the same prompts, repositories, personas, and agents. Brand rankings, install rates, mention rates, rank movement, and per-agent pick rates show how preference changes over time.

APO works as a repeatable measurement and improvement program. A one-time audit captures agent behavior at one moment, while scheduled runs show whether your changes continue to influence real choices.

Conclusion

Agent Preference Optimization measures whether a coding agent chooses your product and acts on that choice, while AEO measures whether an AI answer cites you for a person to evaluate. An agent’s choice starts with model knowledge and changes with session context, while live research contributes only when needed. A useful APO program builds a representative baseline and captures full session traces. You can then fix repeated gaps and rerun the benchmark as models and competitors change. Gauge puts this repeatable approach into practice so you can become the product agents choose.

FAQs

How does Agent Preference Optimization (APO) differ from AEO and SEO?

APO improves the chance that a coding agent selects and installs your product. AEO helps AI answers cite your content for a person to evaluate. SEO helps your pages appear in search results when an agent or person searches the web.

Do coding agents always research before choosing a product?

Coding agents sometimes choose using model knowledge or repository context without searching. Gauge captures the full session so you can see whether Claude Code or Codex searched before making a choice. You can then focus on documentation when research shapes the decision and other factors when it does not.

Why are one prompt and one run not enough?

One run shows what one agent chose under one set of conditions. Gauge repeats tasks across repositories, personas, prompts, agents, and models to identify stable patterns. Repeated runs help you separate a real preference from a one-off result.

Does good SEO or AEO automatically win agent preference?

No. Search visibility only matters at the live research stage, and agents often decide before they ever search, based on model knowledge or session context alone. Gauge captures every stage of the decision, not just the research step, so you can tell whether a loss happened before an agent ran a single search.

How do I benchmark my product against competitors in Agent Preference Optimization (APO)?

Run structured head-to-head tasks alongside the open-ended ones, naming your product and the competitors that matter under the same repository and persona conditions. Gauge supports this format directly, so you can see which competitor wins a given scenario and why, not just whether your product was considered at all.

How do I improve my product's Agent Preference Optimization (APO) performance?

Improvement starts with the runs, not a guess. Once Gauge's traces show a repeated gap, whether a missing framework guide, a weak product description, or stale compatibility information, you fix that specific issue and rerun the same benchmark to confirm agent behavior actually changed.

What does Gauge capture during a coding agent's session?

Gauge records the full trace: the searches an agent runs, the pages and documentation it fetches, the packages it considers and installs, the files it writes, and the final implementation outcome. That level of detail shows why a product won or lost, not just which one did.