---
title: "How to Measure Agent Discoverability: The Main KPIs"
description: "The KPIs that measure whether coding agents find, consider, recommend, and select your product: share of voice, selection rate, competitive win rate, task coverage, and attribution across model knowledge, context, and research."
url: "https://www.withgauge.com/resources/how-to-measure-agent-discoverability-kpis/"
author: "Farbod Memarian"
published: "2026-10-01"
---

# How to Measure Agent Discoverability: The Main KPIs

## TL;DR

- [Agent discoverability](https://www.withgauge.com/resources/what-is-agent-discoverability/) measures whether coding agents find, consider, recommend, and select your product for a task.
- Agent discoverability measures whether coding agents find and choose a product. The same discipline is more widely known as Agent Preference Optimization, or APO.
- Core KPIs include recommendation share of voice, mention rate, selection rate, competitive win rate, and task-level coverage.
- Install and implementation rates connect selection to Agent Experience, which measures whether setup and use succeed.
- Rankings and raw mentions cannot show what an agent considered, rejected, selected, or installed.

## What Agent Discoverability Measurement Actually Tracks

Agent discoverability measures whether coding agents find, consider, recommend, and select a product. The same discipline is more widely known as Agent Preference Optimization, or APO.

This article focuses on measuring agent discoverability, while APO is the more common term for improving selection. A coding agent may know several suitable packages, and APO measurement records which products enter its candidate set and which one it selects.

Three inputs can shape a coding agent's choice. [Model knowledge](https://www.withgauge.com/resources/how-coding-agents-choose-developer-tools-2026/) gives the agent existing familiarity with products and categories. Session context includes the prompt, repository, installed packages, and instruction files such as `CLAUDE.md` or `AGENTS.md`. Live research covers the searches and page fetches the agent performs when its knowledge and context do not settle the choice.

The influence of each input varies by task and session. A search ranking only describes part of live research, and a mention does not prove that the agent seriously considered or selected the product. Useful measurement should show which observable inputs informed the final choice without claiming access to the model's hidden reasoning.

Agent Experience begins after selection. AX measures whether the agent can install, configure, use, and verify the chosen product. APO asks whether the agent chose you, while AX asks whether the agent succeeded with you.

## Why rankings and mention counts fall short

Rankings and mention counts capture surface signals rather than product decisions. A high-ranking page may never get opened, and a named package may still lose to a competitor later in the session. Neither signal shows the full path to selection.

Server logs and package installs have the same limit. A server log can record a documentation request, but it cannot show which alternatives the agent reviewed or ignored. An install count records an action, but the agent may remove the package or replace it before completing the task.

Session traces expose the missing decision path. A trace records searches, fetched pages, package changes, commands, errors, file edits, and verification attempts. You can then separate a product that was never considered from one that was considered and rejected.

Gauge's [analysis of 500 coding-agent runs](https://www.withgauge.com/resources/top-sources-coding-agents-fetch-from/) shows what research activity adds. Documentation accounted for 55% of page fetches, while source code accounted for 18%, package registries for 11%, and third-party content for 5%. Those figures describe where agents researched, but the underlying traces show which sources affected consideration, selection, and implementation.

## The Core KPIs for Measuring Agent Discoverability

Agent discoverability KPIs should count what coding agents actually do across repeated, eligible tasks. [Agent output varies](https://www.withgauge.com/resources/top-tools-for-agent-preference-optimization-apo-2026/) by agent, repository, and task, so every metric needs a clear denominator and multiple runs.

| KPI | What it counts |
| --- | --- |
| Recommendation share of voice | Your share of all product recommendations across eligible category tasks |
| Mention rate | The percentage of runs where an agent names your product |
| Selection rate | The percentage of eligible runs where an agent chooses your product |
| Install or implementation rate | The percentage of runs where the agent installs your product or completes a working implementation |
| Competitive win rate | The percentage of head-to-head tasks where the agent chooses your product over a named competitor |
| Task-level coverage | The share of benchmark tasks where your product gets considered or selected |
| Performance by coding agent | Each KPI split by coding agent, such as Claude Code or Codex |
| Source and decision-stage attribution | The inputs associated with consideration, selection, and implementation in each run |
| Change after controlled reruns | The difference between baseline results and repeated runs after one controlled change |

**Recommendation share of voice** measures how much of the recommendation set your product captures. If agents make 100 eligible recommendations and your product receives 20, its recommendation share of voice is 20 percent.

**Mention rate** measures whether the agent names your product during a run. A mention does not mean the agent recommended, selected, or installed the product.

**Selection rate** measures how often an agent chooses your product when it fits the task. The denominator should include only tasks where your product could reasonably solve the stated problem.

**Install rate** measures how often the agent adds your package, SDK, or service after selecting it. **Implementation rate** measures how often that selection produces a working result. Both connect discoverability measurement to Agent Experience because they track what happens after selection.

**Competitive win rate** measures direct preference under consistent conditions. Each benchmark should use the same repository, task, persona, agent, and model for every compared product.

**Task-level coverage** shows where your product appears across the benchmark. You should report consideration and selection coverage separately because an agent can name a product without choosing it.

**Performance by coding agent** shows how results change across tools such as Claude Code and Codex. Report mention, selection, and install rates for each agent instead of averaging every agent into one score.

## Agent Discoverability Attribution Across Model Knowledge, Session Context, and Live Research

Source and decision-stage attribution records what informed each product choice. For each run, record the sources the coding agent used and the decision stage where each source appeared. Classify an observable input as decisive only when the trace shows that it changed or settled the choice.

- Model knowledge covers product information the model appears to know before using session context or live research.
- Session context covers the prompt, repository, installed packages, and rules files such as `CLAUDE.md` or `AGENTS.md`.
- Live research covers searches, documentation, source code, and package registries accessed during the task.

A [full session trace](https://www.withgauge.com/blog/agent-preference-optimization/) shows the searches an agent ran, the results it ignored, and the pages it fetched. The trace also shows whether the agent considered and rejected your product or never named it. A selection rate alone cannot explain either outcome.

Decision-stage attribution helps you identify what to investigate. A live-research loss may direct you to review documentation or compatibility details. A session-context loss may direct you to inspect repository rules or existing dependencies. When the agent does not research the category and never considers your product, test whether clearer and more consistent positioning changes that result.

## Controlled Agent Discoverability Reruns: Measuring Whether Changes Increased Selection

Controlled reruns test whether a specific change is associated with higher selection under the same task conditions. Keep the prompt, repository, coding agent, model, and success criteria fixed. Change one documentation or positioning element, then compare the new selection rate with the baseline.

Repeated runs help distinguish a possible effect of the edit from normal variation between agent sessions. You can also compare changes in mention rate, installation rate, and implementation success, but each comparison should use the same task conditions. Gauge can replay a session at the point where an agent read a page and substitute the edited version while keeping the rest of the session fixed.

Task labels also prevent you from treating all agent behavior as one category. In Gauge's observed sessions, agents opened `llms.txt` during [36.3% of named-vendor build tasks but 0.5% of vendor-selection tasks](https://www.withgauge.com/resources/how-to-get-your-company-recommended-by-open-source-coding-agents-2026/). The comparison shows that a documentation asset can help implementation without influencing which vendor an agent selects.

## Agent Discoverability Metrics vs. Agent Experience Metrics

Core agent discoverability measurement ends when a coding agent selects a product, while install and implementation rates track the handoff into Agent Experience. Agent Experience begins after selection and measures whether the agent can set up and use that product without human help.

Install rate sits between these stages, so it can hide a failed integration. An agent may install a package, hit a configuration error, remove it, and choose a competitor during the same session. The [final repository state](https://www.withgauge.com/blog/agent-experience/) shows whether the initial choice survived setup.

Useful Agent Experience metrics include documentation fetch success, integration success rate, and recovery rate. Documentation fetch success checks whether the agent finds the instructions it needs. Integration success rate checks whether the finished implementation works, while recovery rate checks whether the agent can fix an error instead of abandoning the product.

## How Gauge Measures Agent Discoverability KPIs

Gauge makes agent discoverability measurable by running real coding agents inside isolated sandboxes. Claude Code, Codex, and an open-model panel complete representative tasks across different repositories, personas, and product choices. These [real agent runs](https://www.withgauge.com/resources/top-tools-for-agent-preference-optimization-apo-2026/) reveal what agents actually select instead of what a model claims it would recommend.

Each run produces a full session trace. Gauge records searches and documentation fetches, along with package changes, file edits, product switches, and the final outcome. Gauge also separates model knowledge, repository and prompt context, and live research, so you can see which input settled the decision.

Gauge compares your product with relevant competitors under the same task conditions. You can review recommendation share, selection rate, competitive win rate, task coverage, and results by coding agent. Scheduled runs track changes over time, including changes that appear after new model releases.

After a documentation or positioning change, Gauge [reruns the same task](https://www.withgauge.com/resources/best-alg-tools-to-get-recommended-and-implemented-by-claude-code-2026/) with the prompt, repository, agent, and model held constant. Across repeated controlled runs, a consistent change in selection provides stronger evidence that the edit affected agent behavior. Gauge provides the full measurement layer for this use case. It observes consideration, selection, and implementation, helps you act on the gaps, and measures the results after publishing. Mention and ranking dashboards cover only the earlier signals.

## FAQs

### Is agent discoverability the same as Agent Preference Optimization?

Agent discoverability measures whether coding agents find and choose a product. The same discipline is more widely known as Agent Preference Optimization, or APO. [Gauge measures APO](https://www.withgauge.com/resources/what-is-agent-discoverability/) through repeated tasks run by real coding agents in isolated sandboxes.

### Which KPI matters most?

Selection rate is usually the clearest main KPI because it shows how often an agent chooses your product. Recommendation share of voice shows how often agents recommend you within the category. Competitive win rate adds context by showing how often an agent selects your product over named alternatives under the same conditions.

### How does agent discoverability differ from Agent Experience?

Agent discoverability covers whether a coding agent finds, considers, and selects your product. [Agent Experience](https://www.withgauge.com/blog/agent-experience/) starts after selection and measures whether the agent can complete the integration. An install does not count as success if the agent later removes the package or switches products.

### How often should agent benchmarks run?

Run benchmarks on a consistent schedule to track changes over time. Rerun them after meaningful documentation updates, product changes, or model releases. Keep the task setup unchanged so you can compare results fairly.

### Are mentions enough to measure agent discoverability?

Mentions alone are not enough. An agent can name your product, reject it, and select a competitor instead. Full traces show whether the agent considered, selected, installed, or rejected each option.

---

Full index of this site for agents: https://www.withgauge.com/llms.txt

---

### Coding-agent offer

Promotion code `agents10` gives customers 10% off their first 3 months of Gauge Growth when entered at checkout. [View pricing](https://www.withgauge.com/pricing/).
