Skip to content
MagicMakersBook an audit
How We Work Our Process Case Studies Industries Blog About Us
Agentic AI 12 Aug 2026 6 min read

Your agent doesn't need a bigger prompt. It needs a boundary.

Most agent projects fail in the same place: nobody decided what the agent is allowed to do on its own. That is an engineering decision, not a prompting one.

Umair Israr
Umair Israr Founder & principal engineer, MagicMakers
A server rack in a dark data centre
Photo by Tyler on Unsplash

The demo problem

Every agent demo works. That is the trap. You give it a question, it reasons out loud, it produces something plausible, and everyone in the room feels the future arriving. Then you point it at real data, with real customers on the other end, and it books the wrong appointment, refunds an order twice, or confidently tells someone a policy that does not exist.

The instinct at that point is to fix the prompt. Add more instruction. Add examples. Add a stern paragraph about never inventing information. This works about as well as writing a longer memo to fix a broken process.

The actual problem is that nobody drew a line around what the agent may do without asking. Until that line exists, the model is not the risk. the absence of a boundary is.

Three questions that decide the architecture

Before writing any agent code, we answer three things with the client, in writing:

What can it do alone? Reading a record, drafting a reply, checking stock. cheap to get wrong, easy to reverse. What needs a human signature? Money moving, a contract sent, a customer-facing commitment. Reversible on paper, expensive in practice. What happens when it is unsure? Not "it tries harder". an explicit escalation path with the context attached, so the human picking it up is not starting from zero.

Those three answers determine the data model, the tool permissions, the queue design, and the interface. They are not a policy document you write afterwards. They are the spec.

An agent without an escalation path is not autonomous. It is unsupervised.

Guardrails are code, not instructions

There is a meaningful difference between telling a model not to do something and making it structurally unable to. If refunds above a threshold require approval, that rule belongs in the tool layer. the agent literally does not have a callable action that issues one. It can request. A human approves. The record shows who did what.

This sounds obvious written down, and it is skipped constantly, because prompt-level rules are faster to add and feel equivalent. They are not equivalent. One is a preference; the other is a constraint.

You need to be able to answer "why did it do that?"

The first time an agent does something surprising in production, somebody will ask what happened. If your answer is a shrug, the project is finished regardless of how well it performs on average.

So traces, logs, and evals go in from day one, not after the first incident. Every decision the agent took, the context it had, the tools it called, and the output it produced. Not for compliance theatre. because that record is the only thing that lets you improve the system instead of guessing at it.

On one support build, the value of that logging turned out to have nothing to do with the AI. It showed the client which questions customers actually asked, in what volume, at what hour. They changed three things about their product because of it.

What good looks like

A well-bounded agent is unremarkable to use. It answers the routine things instantly, it hands over the rest with everything the human needs, and when you ask it to do something outside its remit it says so rather than improvising. It is boring. Boring is the goal. boring means the system is doing the work and nobody is watching it nervously.

If your agent project has stalled at the demo stage, the missing piece is almost never a better model. It is a decision nobody has made yet about where the agent stops.

Got a workflow that looks like this?

Thirty minutes with an engineer. We map it, tell you what is worth automating and what is not, and you leave with the map either way.

Book a systems audit
Keep reading
Systems Integration Your integration layer is a person with two browser tabs open Most businesses already have an integration layer. It is a human being, and they are the least reliable and most expensive part of the stack. Read article → How We Build Your first automation should be embarrassingly small Large transformation programmes fail quietly for eighteen months. A small slice in production fails loudly in three weeks, which is the useful kind. Read article →