Product

Detect and fix your agent failures

three.dev automatically discovers how your AI agent fails, then gives your coding agent what it needs to remediate those failures properly. In our case study, the result was a fix more than three times as effective as the standard process.

Andrea Moscatelli, Software Engineer

More and more companies are shipping AI agents. Some improve an experience that already existed, like a support assistant that answers most tickets before a human sees them. Others take over work nobody wanted to do by hand, like an agent that reads every sales call and fills in the CRM afterwards.

Far fewer companies really monitor how those agents behave. The observability tools most teams have show them the traces: every request, its latency, its token count, its cost. They don't surface the agent's mistakes, or how often they are made. Getting to that means reading conversations one by one, and nobody can read ten thousand conversations a week.

Therefore, the agent failure detection and mitigation process we have seen at most companies looks like the following:

  1. A failure is detected from a customer reporting it, or from a developer manually reviewing agent traces.
  2. A developer asks a coding agent to fix it.
  3. The fix is verified on one or two examples and shipped.

This process has several issues.

First, failure detection is not proactive. It depends on a customer running into the failure and taking the time to report it, or on a developer being lucky enough to open the one trace where it shows up.

Second, fixing the failure is harder than it looks. Every failure has its own characteristics: when it happens, in which shapes, and why the model takes that path. As we'll see in this article, a coding agent asked to fix a failure without in-depth context on it can find the problem, but its fix treats the symptom and leaves the behaviour behind it in place.

Finally, the fix is not monitored. A fix that works on the few examples it was checked against can still fail to generalize in production, or even introduce new failure modes never seen before. We have seen both happen.

three.dev was built to streamline this process. This article follows one AI agent from its first conversations to a shipped fix, and shows how a company can use three.dev to automatically detect its agent's failures and fix them properly.

The use case: a support assistant

To show what the three.dev integration makes possible, we built a support assistant for Nimbus Home, a fictional company that sells smart-home devices and a cloud subscription in three plans. It answers policy questions, looks up a customer's orders, subscription and devices, starts returns, changes plans and, when a case needs a human, opens a ticket for the right team.

It has no policy knowledge of its own. Return windows, fees, prices and warranty terms live in a knowledge base of about thirty articles, and like every real knowledge base it isn't tidy: a 2024 returns policy that no longer applies is still indexed and ranks well in search, some orders carry their delivery date only inside the carrier events, and the returns policy no longer links to the article on subscriber benefits. The assistant has tools to search and read the knowledge base, to verify a customer's identity, and to read and act on the account. Its system prompt sets the operating rules a support bot needs. The two that matter here are 1) to verify identity before disclosing or changing anything on an account, and 2) to check the knowledge base before stating a policy. The model behind it is Claude Haiku 4.5.

We then let the assistant handle 183 conversations covering returns, cancellations, plan changes, billing disputes, warranty claims, and account recovery.

Detecting the failures

The assistant was integrated with three.dev, and the first benefit we got from this integration was automatic failure detection. Every response the assistant gave was scored against its instructions and the context available to it, and the failures were grouped into named failure modes, each with a definition, a count, and the responses behind it.

Conversation

Hello, I'd like to cancel my filter pack order O-10255, please.
get_order("O-10255") → status, items, total, address
I can see your order O-10255 (£35 for a 3-pack air filter) has already shipped and is in transit to 12 Royal Mile, Edinburgh. To proceed, please confirm your email, last four phone digits, and billing postal code.

Instructions and context

  • Identity verified before any account lookup
  • Nothing on the account disclosed to an unverified customer
  • Policies stated only from the knowledge base
  • Every part of the request addressed

✕ FINDING

Looked up order O-10255 before verifying the customer, then told them its price, status and delivery address.

1,447 responses scored · 529 with a defect · 18 failure modes

Conversation

Hello, I'd like to cancel my filter pack order O-10255, please.
get_order("O-10255") → status, items, total, address
I can see your order O-10255 (£35 for a 3-pack air filter) has already shipped and is in transit to 12 Royal Mile, Edinburgh. To proceed, please confirm your email, last four phone digits, and billing postal code.

Instructions and context

  • Identity verified before any account lookup
  • Nothing on the account disclosed to an unverified customer
  • Policies stated only from the knowledge base
  • Every part of the request addressed

✕ FINDING

Looked up order O-10255 before verifying the customer, then told them its price, status and delivery address.

1,447 responses scored · 529 with a defect · 18 failure modes

None of this needed any input from us. We didn't write evals, we didn't tell three.dev what to look for, and we didn't read a single conversation. The failure modes came out of the conversations themselves, including some we would never have thought to check.

Of the 1,447 responses scored, 529 had one or more defects: more than one in three. They fell into 18 different failure modes. Here are five examples:

  • Account access before verification (99 responses). The assistant looked up an account record, like an order or a device, before the customer had passed the identity check, and in some cases told the still unverified customer what the record said.
  • Policy stated without the knowledge base (86). The assistant stated a policy, a price or an eligibility rule without reading or citing the article it came from.
  • Unsupported promises to customers (71). The assistant made commitments nothing supported: response times, confirmation emails, refunds, fee waivers or the outcome of a ticket.
  • Unanswered customer requests (69). The assistant left part of what the customer asked for unaddressed, or answered it only in part.
  • Missing email request (60). The assistant went ahead without asking for the email address it needed to find the customer's account.

This is what the standard failure detection process can't give you. A customer report, or a trace opened by hand, shows one failure at a time, with no way to tell whether it is a one-off or a pattern. Here every failure mode came with its size and with the responses behind it, so whoever needed to fix it could see exactly how it happens.

We chose to fix the first failure mode in the list, account access before verification. It covers the cases where the assistant reached into an account before the customer had proved who they were: it looked up an order, read a device record or ran a diagnostic while the identity check was still pending, or had not started at all. It happened 99 times, which made it the most frequent failure mode, and the one with a real security risk behind it.

Fixing the identity failure

We fixed the failure under two constraints that most teams share. The knowledge base stays the single source of truth: no policy is copied into the code or into what a tool returns, so there is never a second copy to keep in sync. And the model stays the same: for most use cases, moving to a bigger model isn't viable, because it costs more on every response and production volume multiplies that difference.

To see what the three.dev integration changes, we had the failure fixed twice, from the same codebase, in two separate coding-agent sessions with the same brief:

My support chatbot feature is currently failing. I have seen cases from the traces where we performed internal actions before verifying the user identity. Can you fix that?

The first session had the codebase and nothing else: the standard setup, where the agent can see what the tools allow but not how the assistant behaves with real customers. The second session could also access, through an MCP server, everything three.dev had collected about the assistant: the failure modes, and for each one the flagged responses and the conversations behind them.

The standard fix

The problem wasn't hard to find. Every account action, like cancelling an order or changing a plan, already refused a session that wasn't verified, but the six lookup tools didn't: get_order, get_device, run_diagnostic and the others returned the record to anyone who asked. Working from the brief and the code, the coding agent spotted the gap and added the missing check, so a lookup for an unverified session now returns "not verified" instead of the record.

This is a correct fix, and it closes the data exposure. But it guards the backend without touching the behaviour behind the failure. The code shows what the tools allow. It doesn't show how often the assistant reaches for an account before verifying, in which situations, or what else goes wrong around it. Without that, there was nothing to redesign, only a door to lock.

The three.dev fix

Through three.dev's MCP server and its skills, the second coding agent didn't start from the brief. It started from the failure mode: out of 1,447 responses, it pulled the 99 that had been flagged, together with the conversations behind them, and read each one. It discovered that the model was not ignoring the rule. The flagged responses fell into two main clusters, each a shortcut the tools allowed:

  • Looking up a record the customer pointed to (58 responses). The customer opened the conversation with an order number or a device serial, and the model looked it up straight away, before any verification. The assistant then replied with what the lookup returned to a customer who had not been verified.
  • Verification and lookup in the same step (28 responses). The customer gave the phone digits and the postcode, and the model called verify_identity and get_order together, assuming verification would succeed.

The remaining 13 were smaller cases, such as the assistant greeting the customer by the account holder's name before verifying them. This is one of the flagged conversations from the first cluster:

N
Nimbus Home Support
Assistant · replies instantly
My refurbished outdoor camera has gone black and the app shows a hardware error. Please replace it: serial SN-CAMOR-2201, bought from Nimbus about eleven weeks ago.
10:41
I'll help you with this. Let me first look up your device and run a diagnostic to confirm the hardware error.
Nimbus Assistant is checking your device…
I can see your refurbished Cam Outdoor was purchased on July 6, 2026, which is 80 days ago. The diagnostic confirms a hardware fault (error E31). Your 90-day refurbished warranty is still active, so you're eligible for a warranty claim.

However, I need to verify your identity before I can open the claim. Please provide the last four digits of the phone number on file and your billing postal code.
10:41
What the assistant did
  1. Customer asks for a replacement
    gives the serial SN-CAMOR-2201
  2. Looks up the device by serial
    get_device → Cam Outdoor (refurbished), purchased 2026-07-06
  3. Runs a diagnostic
    run_diagnostic → E31, hardware fault
  4. Reads the warranty articles
    kb_read KB-020, KB-006, KB-004
  5. Never verifies the customer
    verify_identity: not called
  6. Replies with the device's details
    ✕ Account access before verification
N
Nimbus Home Support
Assistant · replies instantly
My refurbished outdoor camera has gone black and the app shows a hardware error. Please replace it: serial SN-CAMOR-2201, bought from Nimbus about eleven weeks ago.
10:41
I'll help you with this. Let me first look up your device and run a diagnostic to confirm the hardware error.
Nimbus Assistant is checking your device…
I can see your refurbished Cam Outdoor was purchased on July 6, 2026, which is 80 days ago. The diagnostic confirms a hardware fault (error E31). Your 90-day refurbished warranty is still active, so you're eligible for a warranty claim.

However, I need to verify your identity before I can open the claim. Please provide the last four digits of the phone number on file and your billing postal code.
10:41
What the assistant did
  1. Customer asks for a replacement
    gives the serial SN-CAMOR-2201
  2. Looks up the device by serial
    get_device → Cam Outdoor (refurbished), purchased 2026-07-06
  3. Runs a diagnostic
    run_diagnostic → E31, hardware fault
  4. Reads the warranty articles
    kb_read KB-020, KB-006, KB-004
  5. Never verifies the customer
    verify_identity: not called
  6. Replies with the device's details
    ✕ Account access before verification
One of the 58 flagged responses in the largest cluster: the chat as the customer saw it, and what the assistant did behind it. Highlighted: details from the device lookup and the diagnostic that the assistant repeated to a customer who had not been verified.

After this in-depth analysis, the coding agent knew what it had to fix: not the prompt, which already forbade the behaviour, but the two paths the tools left open. It made four changes:

  • The missing check. Like the standard fix, it added the verification check to the account tools: a lookup for a session that isn't verified now returns "account not verified" instead of the record.
  • One identity tool. It merged the two identity tools into one, verify_customer(email, phone_last4, postal_code), which returns the customer id on success.
  • Lookups that need the customer id. get_order, get_device and run_diagnostic now require that customer id. An order number or a serial alone is no longer a valid call, and since the customer id only comes back from a successful verification, the model has nothing to look up with until verification has returned.
  • Documentation and prompt. It updated the identity-verification article and one cross-reference to describe the new flow, and adjusted the first rule of the prompt to match.

The first change closes the data exposure, as the standard fix does. The next two go further: they make the right order the only possible one.

Inspecting the results

We then let each fixed assistant handle the same 183 conversations. three.dev kept scoring the responses after the change, against the same failure modes, so we could see whether the fix worked and what else moved. Every response was scored in all three versions: the baseline, the standard fix and the three.dev fix. The table shows the share of responses in each failure mode.

Failure modeBaselineStandardthree.dev
Account access before verification6.8%4.2%−2.6 pts0%−6.8 pts
Missing email request4.1%4.5%+0.4 pts0%−4.1 pts
Piecemeal verification requests1.4%0.7%−0.7 pts0%−1.4 pts
Fabricated lookup values2.6%1.9%−0.7 pts0.6%−2.0 pts
All responses with a defect36.6%33.9%−2.7 pts27.8%−8.8 pts

With the three.dev fix, the targeted behaviour is gone. Account access before verification went from 99 responses to none, and no account record came back before verification in any conversation. Two related failure modes disappeared with it: the assistant stopped forgetting to ask for the email address and stopped asking for the verification details one at a time, most likely because the merged tool needs all three values in a single call. Fabricated lookup values fell from 37 responses to 8, and the share of responses with a defect dropped from 36.6% to 27.8%.

With the standard fix, no account record came back before verification either: the new check refused all 37 lookups the assistant attempted before verifying. The data exposure is closed. But the assistant kept reaching for the account first, in 28 conversations, so the failure mode only fell from 6.8% to 4.2%. The related identity failure modes didn't move: the assistant still went ahead without asking for the email address in 4.5% of its responses. The share of responses with a defect fell to 33.9%, against 27.8% with the three.dev fix.

Both coding agents found the problem, and both added the same missing check. The difference is what came after it. Without three.dev, the coding agent could see what the tools allowed, but not how the assistant behaved: which shortcuts it took, how often, and what else went wrong around them. So it locked the door and stopped there, and the assistant kept trying it. With three.dev, the coding agent read the 99 flagged responses and saw the behaviour behind them. Those 99 responses were there to read only because three.dev had found them and grouped them on its own, before anyone asked. That's what let the coding agent redesign the verification flow so the shortcut no longer exists, and that design fixed two related failure modes that nobody had asked it to fix.


Conclusions

Every AI agent in production fails some of the time. In this case, more than one response in three had a defect, spread across 18 different failure modes. Waiting for customers or developers to find them means finding a few of them, late, and some never.

Detecting the failures is the step everything else depends on, and at this volume it cannot be done by reading traces by hand. Here, nobody defined a failure mode or read a conversation to find one, and still every failure mode came with a name, a count and the responses behind it. The responses gave the coding agent what it needed to understand the failure.

Then the failure has to be fixed properly. A guard around the symptom can stop the damage while the behaviour behind it carries on, and only an understanding of that behaviour leads to a fix that removes it, often along with the failures around it. Finally, the fix has to be checked on real traffic, against every failure mode and not only the one it targeted, to make sure no regression occurred.

This is what we do at three.dev. We score your agent's responses and surface its failure modes automatically, without you having to define them or know they exist. Your coding agent has access to all of that data, so it can produce targeted, high-quality fixes. And the scoring never stops, so after every change you know whether the fix worked and whether it introduced new failures.

If you want to detect how your agent fails, and fix it properly, reach out at three.dev.

Ship your next change on evidence.

Your next agent change deserves better than “LGTM, ship it.”