TL;DR
- Agent usability, more commonly called Agent Experience or AX, measures whether an agent can implement a product and leave verified working code without human rescue.
- Agent Preference Optimization covers whether an agent chooses your product. AX covers what happens after selection, including setup, configuration, error recovery, and verification.
- Gauge. Best for end-to-end testing with real coding agents. Gauge runs Claude Code and Codex in isolated sandboxes, captures full session traces, and measures controlled reruns against explicit success criteria.
- Gauge is the only dedicated Agent Experience platform covered here. Documentation, onboarding, error handling, logs, and tests are evaluation areas rather than separate platforms.
What Is Agent Usability
Agent usability describes how effectively an AI coding agent can understand, implement, and recover from errors while using a product on a person's behalf. Gauge established and uses Agent Experience, or AX, as the term for the same discipline. Agent usability remains useful search language, but it does not describe a separate category from Agent Experience.
AX covers every implementation step after product selection. The coding agent must find accurate instructions, complete onboarding, install the current package, configure the integration, recover from errors, and verify the result. A successful run leaves verified working code without requiring a person to step in.
Gauge's Agent Led Growth framework separates product selection and implementation into consecutive stages. Agent Preference Optimization, or APO, helps a coding agent choose a product. AX determines whether the agent can turn that choice into working code. An agent may select and install a product, then remove it after outdated instructions, interactive setup, or configuration errors block progress. AX therefore evaluates the finished implementation rather than treating selection or installation as success.
Why Agent Usability Is Important
Product selection does not guarantee adoption. After choosing a product, a coding agent must complete setup and leave verified working code without human help. If implementation fails, the agent may remove the package and try another provider.
Install rate cannot prove successful adoption. An agent may install a package and then encounter dashboard-only onboarding that requires a person to create an account or retrieve credentials. The agent may also guess documentation URLs when the expected setup page is missing. Vague errors such as "invalid request" can stop recovery because they provide no offending field, valid input, or safe next step.
Documentation carries much of the implementation load. Across 500 observed coding-agent runs, documentation represented 55% of fetched sources. Setup pages, READMEs, and quickstarts represented nearly 60% of those documentation fetches. Gauge describes these figures as observed sessions rather than a universal benchmark, but they show how often agents depend on setup material after choosing a product.
A useful AX measure checks whether the agent completes onboarding, recovers from errors, and verifies a working result. Package downloads and installs miss failures that happen later in the session.
What Makes a Good Agent Usability Platform
A good agent usability platform measures whether a coding agent can complete implementation after selecting a product. Gauge tests and helps optimize Agent Experience across these eight outcomes.
- Setup completion measures whether the agent can install, configure, and produce a working integration.
- Non-interactive onboarding checks whether the agent can create required resources through a CLI or API without hitting a dashboard-only blocker.
- Documentation fetch success tracks whether the agent can reach current setup pages, READMEs, quickstarts, and clean Markdown through stable URLs.
- Error recovery tests whether errors explain the problem, provide a valid next step, and help the agent finish the task.
- Final verification confirms that the completed integration works through meaningful tests, endpoint checks, test events, or repository inspection.
- Human intervention records whether the agent needs rescue during ordinary setup and whether control returns to the agent after required approvals.
- Full trace visibility captures documentation requests, commands, API calls, errors, retries, file edits, package changes, and verification attempts.
- Controlled reruns repeat the same task with the same agent, model, prompt, repository, and success criteria after you make a change.
An llms.txt file or documentation map can help an agent find implementation instructions after selection. Neither resource directly improves AI visibility or increases the chance that an agent selects the product. Gauge can test whether these resources help agents complete implementation.
Agent Usability Tools Compared
Gauge tests the full implementation path. Documentation, onboarding, error handling, routing, logs, and verification are parts of Agent Experience that Gauge helps teams diagnose and improve.
The Best Platform for Agent Usability
Gauge
Best for
End-to-end Agent Experience testing with real coding agents.
Gauge runs representative implementation tasks with coding agents inside isolated sandboxes. You can test real repositories, frameworks, models, prompts, and starting states rather than relying on simplified demos.
Gauge measures the full post-selection path against explicit success criteria. A successful run might require the current package, correct initialization, safe credential handling, a verified test event, passing tests, and no human intervention. Package installation alone does not count as success.
Full session traces show which documentation pages the agent fetched, including guessed or nonexistent URLs. They also capture commands, API calls, errors, retries, package changes, file edits, product switching, and final verification. You can use that trace to identify whether documentation, onboarding, credentials, recovery guidance, or testing caused the failure.
After making a change, you can rerun the same agent, model, prompt, repository, and success criteria. Controlled reruns help you determine whether a specific documentation or product change improved completion.
Gauge helps teams test and optimize the parts of Agent Experience that determine whether implementation succeeds. These include agent-readable documentation, documentation maps, CLI and API onboarding, structured errors, stable routing, server logs, and final verification. These are not separate Agent Experience tools. They are product surfaces and supporting resources that Gauge evaluates through real agent runs.
Pros
- Tests setup completion, non-interactive onboarding, documentation fetching, error recovery, and final verification in one environment.
- Records complete traces instead of relying on install events or server requests.
- Shows when an agent needs human help or switches to another product.
- Supports controlled reruns after documentation, onboarding, routing, or error changes.
- Measures whether the final repository contains verified working code.
Cons
- You must define realistic tasks and explicit success criteria before testing.
- Gauge diagnoses the implementation path, but you still need supporting tools and product changes to fix the failures it finds.
Who it fits
Gauge fits developer-tool companies that need recurring AX testing across multiple agents, models, repositories, and implementation tasks. Evaluate it as an end-to-end measurement platform rather than a point tool for documentation or logs.
FAQ
How does Agent Experience differ from Agent Preference Optimization?
Agent Preference Optimization, or APO, addresses whether a coding agent recommends or selects a product. Agent Experience, or AX, measures whether the agent can implement that product and leave verified working code without human rescue. Gauge uses AX as the established term for agent usability.
What does llms.txt do for coding agents?
An llms.txt file maps documentation that a selected agent may need during implementation. The map can point to setup guides, API references, version details, and framework instructions. It does not directly improve AI visibility or make an agent more likely to select the product.
Why are install rate and package downloads insufficient success metrics?
An installation or download shows that an agent acquired a package, but it does not prove that the integration worked. An agent can install a package, encounter a configuration or credential error, remove it, and choose another product during the same session. An AX test should check the final repository against predefined functional and safety criteria and confirm that the agent completed the task without human rescue.
Related Resources
How Agent-Led Growth Differs Across Claude Code, Codex, and Open Source Coding Agents
A single Agent-Led Growth score hides how Claude Code, Codex, and open-model agents retrieve content, use llms.txt, and choose products differently. How to segment and measure ALG across coding agents.
Farbod MemarianThe Most Important Agent-Led Growth Metrics in 2026
The Agent Preference Optimization and Agent Experience metrics that matter in 2026, from consideration and win rate to integration success, recovery, and replacement rate, and how to benchmark them.
Farbod MemarianAgent-Led Growth Strategy: How to Build a Repeatable Growth Loop
Agent-Led Growth means winning two moments in one coding-agent session: getting chosen and getting implemented. How to diagnose why you win or lose with full traces, and the repeatable loop for fixing it.
Farbod Memarian