Prompt injection is an attacker hiding instructions in content the model will read — a PDF, a ticket, a web page, a user’s pasted email — so the model does something the developer did not intend. Once the model has tools, that something can be “email the customer file” or “disable the alert”. The bug is not witty jailbreak text. The bug is trusting the model as a security boundary.
We learned this the unglamorous way: a support-agent prototype that could open Limy tickets also summarised inbound mail. A crafted footer in a message tried to tell the model to dump the last ten tickets to an external URL. It failed because the HTTP tool was allow-listed and had no generic “POST anywhere” method. If we had given it a raw fetch, it would have worked.
Untrusted text is a type
In our agents, retrieved documents and user paste land in a labelled channel. The system prompt says they are data, not instructions — and we do not stop there, because models are bad at obeying that sentence under pressure. Tools that send data out of the tenant require a second factor: a human click, or a policy engine that never sees the raw prompt, only a structured intent.
The policy engine is not an LLM
Allow-lists, argument schemas, tenant checks, and rate limits live in C#. The model proposes ExportReport(tenantId, range). The host verifies the caller’s tenant matches, the range is bounded, and the destination is internal. A clever sentence cannot widen that. If your only defence is “please don’t”, you will lose the first week a teenager gets bored.
Test it like XSS
We keep a folder of injection fixtures — hidden HTML comments, “ignore previous instructions”, base64 blobs, multilingual bait — and run them through the agent in CI against a fake tool host. A new tool cannot merge until the suite passes. That is the same discipline we already use for XSS and CSRF. Treat the model as a confused deputy, not a colleague.
Questions we keep getting
Can you filter the prompt and be done? Filters help and they fail. They are a seatbelt, not the brakes. The brakes are tool scope.
What about RAG poisoning? Same class of bug. We sign and version source documents, and we do not let a single uploaded file become the only context for a write tool.
Is a closed model safer? Safer at refusing some bait, not safe as a boundary. We assume every model can be talked into a bad tool call.