An agent you cannot audit is not in production, it is on trial
The second question every operator asks is "what did it do, and why?" If the system cannot answer that, it does not matter how good the demo was.
The second question every operator asks is "what did it do, and why?" If the system cannot answer that, it does not matter how good the demo was.
The first question about an agent is always whether it can do the job. The second one arrives about a week after launch, usually from operations or finance: a customer got an answer nobody recognises, and they want to know what happened.
If your answer is a model name and a prompt, you have a problem. Not because the agent was wrong, but because you cannot prove it was right.
For every run: the input as received, the records the agent read, each tool call with its arguments and result, the decision it reached, and who or what approved the irreversible parts. Timestamped, stored, and searchable by customer rather than only by request id.
That list looks obvious written down. It is also the thing most prototypes skip entirely, because logging is boring and the demo does not need it.
A system that cannot explain a decision it already made is not automation. It is an outsourcing arrangement with something that cannot be questioned.
Once traces exist, evals are cheap. You take the runs that went wrong, turn them into cases, and run them on every change. Nobody argues about whether the new prompt is better, because there is a number.
This is the part that makes an agent maintainable. Without it, every improvement is a gamble, and eventually people stop changing anything because they are afraid of it.
Guardrails are not restrictions on capability. They are the reason anyone lets the thing near a customer. An agent should have a scoped set of tools, an explicit ceiling on the value of actions it can take alone, and a predictable fallback when the model goes off-script.
We usually set that ceiling deliberately low at launch and raise it once the traces show it earning trust. It is much easier to widen a boundary than to recover from a week nobody could explain.
Logs, evals, dashboards, and the runbook for when something looks wrong. Not as a nice extra, but because a system your team cannot inspect is a system your team will quietly stop using.