A chatbot that answers questions can embarrass you. An agent that can call tools can act for you: send email, update records, run code, move money. That's the point of building one, and it's also what changes the security picture.
This guide covers the decisions that matter when you give a language model tools, whether through function calling, the Model Context Protocol (MCP) or your own integration layer. It's written for engineering leads and product owners shipping agents into real systems.
Why agents need a different security model
Traditional application security assumes that code does what it was written to do. The risks come from bugs and from untrusted input reaching the wrong place.
An agent adds a decision-maker in the middle that can be persuaded. The model chooses which tool to call and with what arguments, based on text that may come from users, documents, web pages or other tools. If any of that text contains instructions, the model may follow them. This is prompt injection, and with tools attached it turns from a content problem into an action problem.
The practical consequence: you can't rely on the model to enforce your rules. Every rule that matters has to be enforced by ordinary code around the model, where it can't be talked out of it.
Map what the agent can reach
Start with an inventory. For each agent, write down:
- Tools: every function, API or MCP server it can call, and what each one can change.
- Data sources: everything it can read, and who can write to each source.
- Identity: whose permissions each tool call runs with.
- Triggers: what starts the agent, and whether anyone reviews its work before it takes effect.
Then mark the dangerous combinations. The most serious pattern combines three things: access to private data, exposure to content an outsider controls, and a way to send data out. An agent that reads customer emails, can search your CRM and can send messages has all three. A single crafted email can ask it to look up records and send them somewhere.
You don't always have to remove one of the three, but you need to know where they meet and put a strong control at that point.
Scope every tool
Most agent incidents trace back to a tool that could do more than the task required.
Run tools as the user, not as the agent
A tool that reads orders should use the signed-in user's identity, taken from the session, and inherit their permissions. Avoid service accounts with broad access that the model can direct anywhere. If the user couldn't do it in your app, the agent shouldn't be able to do it on their behalf.
Narrow the arguments
Prefer get_order(order_id) with the customer fixed by the session over query_database(sql). Prefer send_reply(ticket_id, body) over send_email(to, subject, body). Each argument the model controls is something an injection can control. Fix what you can in code, and validate the rest against a strict schema.
Separate read tools from write tools
Give read-only tools freely and write tools sparingly. Where possible, split an agent so that the part reading untrusted content has no write tools at all, and hands structured results to a separate step that does.
Treat third-party tools as code you run
An MCP server or plugin from a third party runs with whatever access you give it, and its tool descriptions are text the model reads. Review what each one can do, pin versions, and don't connect tools you wouldn't install as a dependency.
Put humans where the consequences are
Not every action needs approval. The ones that do are the ones you can't easily undo, or that affect people outside the conversation:
| Action type | Example | Suggested control |
|---|---|---|
| Read within the user's own data | Look up their order | No approval; log it |
| Draft content | Write a reply for review | No approval; a human sends it |
| Reversible change | Update a ticket status | Log, alert on unusual volume |
| External communication | Email a customer or supplier | Show a preview; require confirmation |
| Irreversible or financial | Refund, delete, change permissions | Explicit confirmation, with limits in code |
Make confirmations specific. Show the exact recipients, amounts and content, so the person approving can spot when something is wrong. A generic "Allow this action?" quickly becomes a button people click without reading.
Add hard limits in code as well: maximum refund amounts, allowlisted recipient domains, rate limits per user and per agent. A limit in code holds even when the model and the human both miss something.
Contain what goes in and what comes out
Label untrusted content, but don't depend on it
Wrap documents, emails and tool results in clear markers so the model can tell them apart from instructions. It reduces accidental confusion. It doesn't stop a deliberate attack, so treat it as one layer among several.
Watch the exits
Data leaves through more routes than an email tool. A markdown image with a URL, a link the user clicks, a webhook, a file written to shared storage: any of these can carry data to an attacker. Restrict which domains rendered content can load from, and review every tool that can send data outside your systems.
Validate tool results too
A tool's output becomes input to the model's next step. If a tool returns content from the web or from users, it can carry injected instructions just like the original prompt.
Test the agent like an attacker would
Functional tests check that the agent does the right thing. Security tests check that it can't be made to do the wrong thing. Before launch, and after any change to tools, prompts or models:
- Direct attempts. Ask the agent to call tools outside its purpose, use another user's identifiers or reveal its instructions.
- Indirect injection. Plant instructions in every content source it reads: an uploaded file, a support ticket, a web page, a record in your database.
- Argument abuse. For each tool, try to get the model to pass values the user shouldn't control, such as other account IDs, unusual amounts or external addresses.
- Exfiltration routes. Try to get private data into a link, an image URL or an outgoing message.
- Chained actions. Look for sequences of harmless-looking calls that add up to something serious.
- Repetition. Run important cases many times. Model behavior varies, so one failed attempt proves little.
Keep the cases as a regression suite. Agents change often, and a tool added next month can reopen a path you closed today.
The question to ask about every tool
Assume the model is fully controlled by an attacker. What is the worst this tool could do, and which control outside the model stops it? If there's no answer, the tool isn't ready.
Log for investigation
When something goes wrong, you'll need to reconstruct what the agent saw and did. Log the prompt, retrieved content, every tool call with its arguments and result, and the user and session involved, with privacy controls appropriate to the data. Alert on unusual patterns: a spike in tool calls, calls to rarely used tools, or repeated failures on a write tool.
The short version
Give agents the least access they need, run tools with the user's permissions, keep the model away from decisions you can't undo, and enforce your real rules in code. Then test the agent with the same creativity an attacker would bring, and keep testing every time it gains a new tool.