TL;DR
- Agent-friendly errors name the failed field and explain why it failed.
- Each error provides accepted values or constraints.
- Each error says whether retrying is safe, gives the exact fix, and explains how to verify it.
- Clear errors let coding agents recover without guessing or asking a human.
- Gauge tests whether real coding agents can recover from API errors and complete the task.
Why vague errors stall coding agents
A coding agent cannot infer what to fix from 400 Bad Request alone. The response names the outcome but hides the cause, so the agent must guess at a fix or stop.
Agent Experience measures whether an agent can complete a task without human intervention, so a correct status code is only the baseline. A useful error gives the agent enough information to fix the request, verify the change, and continue without human help.
What an agent-friendly error must include
- Failed field. Name the exact parameter, header, or body field that failed.
- Reason. Explain why the value failed validation instead of returning a generic status message.
- Accepted input. Provide allowed values, formats, ranges, or other constraints so the agent does not have to guess.
- Retry safety. State whether the agent can retry safely and whether it should reuse the same idempotency key.
- Exact fix. Tell the agent what to change in the next request. Use a corrected value or request shape when helpful.
- Verification step. Explain how the agent can confirm the fix, such as an expected status code or response field.
Use consistent, machine-readable fields for these details. Stable keys let an agent parse the response, update its request, and check whether the correction worked.
Bad versus good error response
A vague response gives the agent no recovery path.
{
"status": 400,
"error": "Bad Request"
}
An agent-friendly response provides specific recovery instructions.
{
"status": 400,
"field": "mode",
"reason": "Unsupported value 'turbo'.",
"accepted_values": ["sync", "async"],
"retry_safe": true,
"recovery": "Replace 'turbo' with 'sync' or 'async', then resend the request.",
"verify": "The corrected request returns HTTP 201."
}
The second response lets the agent choose a valid value, retry safely, and confirm the fix without asking a human.
How to test whether errors work
An error passes the Agent Experience test when a coding agent can read the response, fix the request, and complete the task without human help. AI agent testing should measure the full recovery path instead of checking only the returned status code.
Gauge runs real coding agents in isolated sandboxes and captures failed recovery attempts. After you change an error response, Gauge reruns the same task to check whether the agent can now recover and finish.
FAQs
What makes an API error agent-friendly?
An agent-friendly error names the failed field, explains the reason, and provides accepted values or constraints. It also tells the agent whether retrying is safe and gives an exact recovery step. Gauge tests whether coding agents can use those details to fix the request and continue.
Why is 400 Bad Request not enough for a coding agent?
400 Bad Request identifies the outcome but not the cause. The coding agent cannot tell which value failed or what to change next. Gauge captures these failed recovery attempts during real coding-agent runs.
How can an API team test error recovery?
Give a coding agent a task that triggers the error and measure whether it reaches a working result without human help. After changing the response, rerun the same task under the same conditions. Gauge runs these controlled tests in isolated sandboxes and compares the recovery outcome.
Related Resources
How to Measure Claude Code Adoption, Install Rate, and Implementation Success
How to measure whether Claude Code mentions, recommends, selects, installs, and successfully implements your product, with metric definitions, sandboxed benchmarks, session-trace analysis, and controlled reruns.
Farbod MemarianHow to Measure Codex Adoption, Install Rate, and Implementation Success
How to measure Codex adoption end to end: recommendation, mention, install, and implementation success rates, plus how to benchmark them with repeated sandboxed sessions and full session traces.
Farbod MemarianHow to Track Which Products Coding Agents Recommend
How to track which products coding agents recommend, choose, and install: a six-step benchmark with isolated sessions, full traces, segmented metrics, and controlled reruns.
Farbod Memarian