Clerk wanted coding agents to use the CLI to set up authentication end-to-end. The problem was that agents would reach for a now-deprecated path, due to their stale training data.
Growth Engineer Devin Gould and his team tried an unusual intervention: replace the deprecated path with an error that explicitly directed the agent to the CLI. Gauge gave them a way to test whether the instruction would change the agent’s behavior before relying on it in the product.
With Gauge, we could test whether an SDK error would get agents to use our CLI.
In the evaluations Devin described, agents went from not using the CLI to using it consistently. After Clerk shipped the change, the team saw a marked increase in CLI users. The experiment gave Clerk a concrete example of how improving the agent experience could influence adoption of its preferred onboarding path.
Making the next step obvious for the agent
Clerk’s goal is straightforward: a developer should be able to ask an agent to add authentication and get a working, secure sign-in experience with as little intervention as possible.
Getting there requires more than making each individual step easy. Agents may know an older version of the product and miss newer tools.
Easier for a human doesn’t necessarily mean easier for an agent.
The team’s hypothesis was that an explicit error would give agents both a reason to act and a clear next step. Agents already try to fix errors; Clerk could use that moment to introduce the tool it wanted them to use.
Testing the SDK change in Gauge
Devin had a strong expectation that the approach would work, but wanted evidence. The team supplied a snapshot of its SDK in Gauge and asked an agent to install it and get an application running.
The changed SDK produced the error. The agent then had to interpret the guidance and continue setup through the CLI. According to Devin, the evaluations showed a clear shift toward the intended behavior.
After the change shipped, Clerk saw that CLI users made it further through the growth funnel. CLI adoption also rose according to their internal tracking.
The evaluation and the production data answered different questions: the test showed that agents followed the new guidance, while Clerk’s own instrumentation showed increased CLI usage after launch. Together, they gave the team evidence that the intervention was worth pursuing.
Bringing evidence into code review
The result helped make building for agents more concrete inside Clerk. Devin shared it with colleagues who were already considering how to improve the agent experience, and the team began looking for other errors where clearer instructions could help.
On Devin’s team, changes that might affect agents now prompt an early question: can we eval this?
I can bring evaluation results into code review to back up why I think a change works.
Devin starts with local experiments, then uses evaluation results from Gauge to support the proposed change during review. Instead of guessing whether the change will have the intended effect, he can point colleagues to runs showing what happened when an agent tried it.
He sees this as an extension of growth engineering. The familiar question is whether a change improves behavior; agent simulations make it possible to investigate that question before exposing the change to real users.
Finding the friction the team didn’t anticipate
Gauge also helps Clerk look beyond the behaviors it explicitly wants to encourage. A run can reveal an agent attempting a nonexistent URL, expecting an unavailable command, or getting stuck further along the implementation.
Those observations create specific product questions. Would a redirect resolve a repeated dead end? Does the CLI need a command agents already expect? Can an agent complete a production workflow without resorting to the dashboard?
Devin’s broader product principle is that features need interfaces for both people and agents. A human may prefer the dashboard, while an agent needs a direct way to perform the same task through the CLI or API. The SDK still needs to produce code that a developer can understand and review.
What other developer tool teams can take from Clerk’s approach
- Choose a behavior you can observe. Clerk wanted agents to discover and use its CLI during setup.
- Test guidance where the agent encounters it. SDK errors can communicate a next step at the moment an agent needs help.
- Bring runs into the review. Use evaluation results to show how a proposed change affects the journey.
- Check what happens after shipping. Compare the tested behavior with product instrumentation, and investigate friction elsewhere in the flow.
For Clerk, Gauge helped turn a plausible SDK change into a tested onboarding intervention. That experience is shaping how Devin’s team builds: identify the behavior, evaluate the change, and use the results to make the next product decision.
