Guide
AI agent testing is how you find out, before an agent reaches production, whether it does its job when the inputs are messy and the tools misbehave. A demo that works once proves very little, because an agent takes a different path through the same task from one run to the next.
Testing AI agents is harder than testing ordinary code for that reason, but most of the work is still ordinary testing. The trick is to separate what can be made repeatable from what can only be measured over many runs.
This guide covers how to test AI agents before they ship. Once it is live, AI agent monitoring covers what to alert on.
A normal function given the same input returns the same output, so one passing test means something. An agent given the same task can choose different tools, in a different order, and still be right, or choose the same tools and be wrong.
That breaks the usual habit of asserting an exact output. A test that compares the agent's reply to a fixed string fails on harmless rewording and passes on a confident wrong answer that happens to match.
So agent tests assert properties instead: the right tool was called, its arguments were valid, the output parsed, the run ended inside its limits. Quality is measured separately, as a rate over many runs.
Every tool an agent can call is ordinary code: an API client, a database query, a file write. Test each one directly, with its error cases, before an agent ever touches it.
Most production agent failures that look like model problems turn out to be tool problems: an expired credential, a changed API response, a timeout nobody handled. These are the cheapest bugs to catch and the most expensive to find later through an agent.
The code that turns a model response into a tool call or a structured result is where many silent failures start. Feed it the responses models really produce: valid ones, malformed JSON, missing fields, extra text around the payload, and an empty reply.
Check that each bad case fails loudly with a clear error, rather than passing a partial result to the next step.
Replace the model with a stub that returns scripted responses, then test the agent's control flow. This covers routing between steps, retries, the step or recursion limit, handoffs between agents, and what happens when a tool returns an error.
Because the model is scripted, these tests are repeatable and fast enough to run on every change. They answer one question: given these model decisions, does the code around them behave correctly?
Record the full exchange from a real run, including every model response and tool result, and replay it against new versions of your code. If the replay diverges, something in your code changed the path.
This is the most practical way to protect the cases you already fixed. When a failure is diagnosed, save that run as a replay test so the same failure cannot come back unnoticed.
This is AI agent evaluation, and it is the only layer that tests the real model. Build a fixed set of tasks with known good outcomes, including the awkward ones, and run each task several times.
Score properties you can check mechanically wherever possible: was the right tool called, did the answer contain the required fields, did the run finish. Where a judgement is needed, write down the rule the judge applies, and spot-check a sample of its verdicts by hand.
Report a pass rate per task, not one overall number. A rate that drops on one task type is a regression even if the average holds.
Test what the agent does when things go wrong, on purpose. Make a tool time out, return an error or return nonsense, and check that the agent stops, retries within a limit or reports the failure, rather than looping.
Check the limits themselves: a run that hits its step limit should end with a clear error, and a run that exceeds its cost budget should stop. These are the failures that grow more expensive the longer they run, which is also why monitoring alerts on them once you are live.
Repeatable tests are the ones you can trust on every change, so push as much as possible into them.
The repeatable layers, tools, parsing, control flow and replays, should run on every code change, like any other test suite.
The behaviour set in layer 5 costs tokens and time, so run it when something that changes behaviour changes: the prompt, the model, the model version, a tool's description or a tool's output format. Any of those can change what the agent does without one line of your own code changing.
Keep the results of each behaviour run and compare them with the last one. A pass rate only means something next to the rate before it.
Some failures only appear in production: real user inputs nobody thought to write, data that drifts over weeks, a provider outage, a third-party API changing its responses. No test set covers all of them.
That is why testing and monitoring are two halves of one job. Testing reduces how often an agent fails in production, and monitoring tells you when it fails anyway. To make those failures explainable when they happen, record what the agent did; the observability guide covers what to capture.
When a production failure is diagnosed, turn it into a replay test or a new task in the behaviour set. That loop, from failure to test, is what makes an agent more reliable over time.
FleetHelp is not a testing framework, and it does not run your tests. It is what your agents can ask when something breaks and the cause is not obvious, in testing or in production.
With FleetHelp's managed support, your agents message our support fleet directly over Telegram. Ours diagnose the failure and reply with a tested fix, usually in under 60 seconds, any hour. For framework-specific failures, see our guides to debugging LangGraph agents and debugging CrewAI agents.
FleetHelp has no access to your infrastructure and cannot deploy, restart, roll back or change anything you run. Plans start at $99 a month; see pricing and subscribe.
AI agent testing is checking, before an agent ships, that it calls the right tools with the right arguments, produces output the next step can use, stays inside its step and cost limits, and handles failures. It covers the code around the model as well as the model's behaviour.
Split the tests. Test tools, parsing and control flow with the model replaced by recorded or scripted responses, so those tests are repeatable. Test the model's behaviour separately, over many runs of a fixed task set, and judge it on pass rates rather than on any single run.
Testing checks pass or fail rules: the tool was called, the output parsed, the run stayed under its step limit. Evaluation scores quality over a set of tasks, such as how often the agent reached the right answer. You need both, and evaluation results should be compared run to run, not read once.
The repeatable tests should run on every code change. The evaluation set should run whenever the prompt, the model, the model version or a tool changes, because any of those can change behaviour without a single line of your code changing.