Most agent demos show the happy path: a question comes in, the right tool is called, a correct answer goes out. Production traffic is not shaped like a demo. A share of what arrives will be out of scope, ambiguous, missing the information needed to answer, or asking for something the agent should not do on its own. If the system has no designed way to handle those cases, the model will handle them anyway, by producing its best guess in the same confident voice it uses for everything else.
That is the failure we design against first. Every agent we ship has a refusal path: a defined, typed outcome for "I should not complete this", with a reason, a message the user can act on and, where it matters, a handoff to a person. It is built before the happy path and tested as hard as it.
What counts as a refusal
We use the word broadly. A refusal is any outcome where the agent deliberately stops short of completing the request. In practice the reasons fall into a small number of groups, and naming them is most of the work.
- Out of scope: the request is outside what this agent was built and evaluated for, even if the underlying model could attempt it.
- Insufficient evidence: retrieval came back empty or weak, and an answer would rest on the model's general knowledge rather than the client's documents.
- Needs a human decision: the action is irreversible, above an agreed limit, or the kind of judgement the client has said a person must make.
- Policy: the request conflicts with a rule the client has set, such as not discussing pricing for a named account.
- Dependency failure: a tool, API or data source the answer needs is down, or it returned something malformed.
The list differs per engagement, but it is always short and always written down. A reason that is not on the list is a sign the scope has not been agreed yet.
Make it a type, not a string
The most common way refusals go wrong is that they exist only in the prompt. "If you are unsure, say you don't know" produces a sentence, not a signal. The rest of the system cannot count it, route it or test for it, and the model will phrase it differently each time. We make the refusal part of the agent's output contract instead.
type RefusalReason =
| 'out_of_scope'
| 'insufficient_evidence'
| 'needs_human'
| 'policy'
| 'dependency_failure';
type AgentResult =
| { kind: 'answer'; text: string; sources: SourceRef[] }
| { kind: 'action'; tool: string; args: unknown; confirmed: boolean }
| {
kind: 'refusal';
reason: RefusalReason;
userMessage: string;
handoff?: { queue: string; summary: string };
};With a typed result, the interface can render a refusal differently from an answer, the operations dashboard can count refusals by reason, and the evaluation suite can assert that a given input produces kind: 'refusal' with a specific reason. The model's structured output is validated against the schema, and anything that fails validation is treated as a dependency failure rather than shown to the user.
Where the decision is made
Not every refusal should be left to the model. We put each check at the point where the information to make it first exists, and we prefer deterministic checks wherever they are possible.
- Before any tool call: a lightweight classifier or rule set decides whether the request is in scope at all. It is cheap, and it stops out-of-scope requests from triggering real actions.
- After retrieval: if no source clears the relevance threshold agreed in evaluation, the agent returns insufficient evidence instead of generating. The threshold is a number in configuration, not a feeling in a prompt.
- Before an irreversible action: refunds, payments, outbound messages and record deletions pass through code that checks limits and permissions. The model can propose the action; it cannot waive the check.
- On dependency errors: timeouts and malformed tool responses map to a refusal with a plain message, not a retry loop the user has to watch.
What the user sees
A refusal that just says "I can't help with that" moves the problem to the user without telling them where to take it. We hold every refusal message to three things: say what the agent cannot do, say why in one plain sentence, and say what happens next.
The third part is where the handoff matters. When a refusal has a handoff, the agent writes a short summary of the conversation so far, the reason it stopped and anything it already looked up, and puts that in the queue the client named. The person who picks it up should not need to ask the customer to repeat themselves. In our experience that summary decides whether users treat a refusal as a dead end or as part of the service.
A refusal is not an error state. It is a designed outcome with its own copy, its own route and its own tests.
Testing both directions
There are two ways to get refusals wrong, and an evaluation suite needs cases for both. Under-refusal is the agent answering when it should have stopped: inventing a policy, acting above a limit, answering from general knowledge when the documents were silent. Over-refusal is the agent stopping when it could have answered. It is quieter, but just as costly, because users learn the system is not worth asking.
So the suite carries labelled cases for each refusal reason alongside the cases that should be answered, including near-misses that sit just either side of the boundary. We report refusal precision and recall per reason rather than one blended accuracy figure, and we agree with the client before launch which of the two errors they would rather have more of. For a customer-facing assistant that can issue refunds, the answer is usually to accept some over-refusal. For an internal research tool, it often goes the other way.
After launch
Refusal rate by reason is one of the first panels on every operations dashboard we build. A sudden rise in insufficient-evidence refusals usually means a document source stopped syncing. A steady climb in out-of-scope refusals is often users asking for something the client may want to build next. A fall to zero is worth investigating too: it can mean a prompt change has quietly taught the agent to guess.
None of this is complicated engineering. It is mostly the discipline of deciding, before the first demo, what the agent should do when it should not do the task, and then treating that answer as part of the product.

