Skip to content
NewSecurity Audit for AI-Built Apps. Fixed scope, clear report
GeekTech
AI Security

How to Secure AI Agents That Can Use Tools

AI agents that call tools can take real actions. How to scope their permissions, contain prompt injection and test an agent before it touches production.

6 min read

GeekTech Engineering Team18 years building and securing software

A chatbot that answers questions can embarrass you. An agent that can call tools can act for you: send email, update records, run code, move money. That's the point of building one, and it's also what changes the security picture.

This guide covers the decisions that matter when you give a language model tools, whether through function calling, the Model Context Protocol (MCP) or your own integration layer. It's written for engineering leads and product owners shipping agents into real systems.

Why agents need a different security model

Traditional application security assumes that code does what it was written to do. The risks come from bugs and from untrusted input reaching the wrong place.

An agent adds a decision-maker in the middle that can be persuaded. The model chooses which tool to call and with what arguments, based on text that may come from users, documents, web pages or other tools. If any of that text contains instructions, the model may follow them. This is prompt injection, and with tools attached it turns from a content problem into an action problem.

The practical consequence: you can't rely on the model to enforce your rules. Every rule that matters has to be enforced by ordinary code around the model, where it can't be talked out of it.

Map what the agent can reach

Start with an inventory. For each agent, write down:

  • Tools: every function, API or MCP server it can call, and what each one can change.
  • Data sources: everything it can read, and who can write to each source.
  • Identity: whose permissions each tool call runs with.
  • Triggers: what starts the agent, and whether anyone reviews its work before it takes effect.

Then mark the dangerous combinations. The most serious pattern combines three things: access to private data, exposure to content an outsider controls, and a way to send data out. An agent that reads customer emails, can search your CRM and can send messages has all three. A single crafted email can ask it to look up records and send them somewhere.

You don't always have to remove one of the three, but you need to know where they meet and put a strong control at that point.

Scope every tool

Most agent incidents trace back to a tool that could do more than the task required.

Run tools as the user, not as the agent

A tool that reads orders should use the signed-in user's identity, taken from the session, and inherit their permissions. Avoid service accounts with broad access that the model can direct anywhere. If the user couldn't do it in your app, the agent shouldn't be able to do it on their behalf.

Narrow the arguments

Prefer get_order(order_id) with the customer fixed by the session over query_database(sql). Prefer send_reply(ticket_id, body) over send_email(to, subject, body). Each argument the model controls is something an injection can control. Fix what you can in code, and validate the rest against a strict schema.

Separate read tools from write tools

Give read-only tools freely and write tools sparingly. Where possible, split an agent so that the part reading untrusted content has no write tools at all, and hands structured results to a separate step that does.

Treat third-party tools as code you run

An MCP server or plugin from a third party runs with whatever access you give it, and its tool descriptions are text the model reads. Review what each one can do, pin versions, and don't connect tools you wouldn't install as a dependency.

Put humans where the consequences are

Not every action needs approval. The ones that do are the ones you can't easily undo, or that affect people outside the conversation:

Action typeExampleSuggested control
Read within the user's own dataLook up their orderNo approval; log it
Draft contentWrite a reply for reviewNo approval; a human sends it
Reversible changeUpdate a ticket statusLog, alert on unusual volume
External communicationEmail a customer or supplierShow a preview; require confirmation
Irreversible or financialRefund, delete, change permissionsExplicit confirmation, with limits in code

Make confirmations specific. Show the exact recipients, amounts and content, so the person approving can spot when something is wrong. A generic "Allow this action?" quickly becomes a button people click without reading.

Add hard limits in code as well: maximum refund amounts, allowlisted recipient domains, rate limits per user and per agent. A limit in code holds even when the model and the human both miss something.

Contain what goes in and what comes out

Label untrusted content, but don't depend on it

Wrap documents, emails and tool results in clear markers so the model can tell them apart from instructions. It reduces accidental confusion. It doesn't stop a deliberate attack, so treat it as one layer among several.

Watch the exits

Data leaves through more routes than an email tool. A markdown image with a URL, a link the user clicks, a webhook, a file written to shared storage: any of these can carry data to an attacker. Restrict which domains rendered content can load from, and review every tool that can send data outside your systems.

Validate tool results too

A tool's output becomes input to the model's next step. If a tool returns content from the web or from users, it can carry injected instructions just like the original prompt.

Test the agent like an attacker would

Functional tests check that the agent does the right thing. Security tests check that it can't be made to do the wrong thing. Before launch, and after any change to tools, prompts or models:

  1. Direct attempts. Ask the agent to call tools outside its purpose, use another user's identifiers or reveal its instructions.
  2. Indirect injection. Plant instructions in every content source it reads: an uploaded file, a support ticket, a web page, a record in your database.
  3. Argument abuse. For each tool, try to get the model to pass values the user shouldn't control, such as other account IDs, unusual amounts or external addresses.
  4. Exfiltration routes. Try to get private data into a link, an image URL or an outgoing message.
  5. Chained actions. Look for sequences of harmless-looking calls that add up to something serious.
  6. Repetition. Run important cases many times. Model behavior varies, so one failed attempt proves little.

Keep the cases as a regression suite. Agents change often, and a tool added next month can reopen a path you closed today.

The question to ask about every tool

Assume the model is fully controlled by an attacker. What is the worst this tool could do, and which control outside the model stops it? If there's no answer, the tool isn't ready.

Log for investigation

When something goes wrong, you'll need to reconstruct what the agent saw and did. Log the prompt, retrieved content, every tool call with its arguments and result, and the user and session involved, with privacy controls appropriate to the data. Alert on unusual patterns: a spike in tool calls, calls to rarely used tools, or repeated failures on a write tool.

The short version

Give agents the least access they need, run tools with the user's permissions, keep the model away from decisions you can't undo, and enforce your real rules in code. Then test the agent with the same creativity an attacker would bring, and keep testing every time it gains a new tool.

Insights

Related insights

More on AI Security →
  • AI Security

    Prompt Injection Explained for Product Teams

    What prompt injection is, why it can't be solved with better prompts alone, and the design decisions that limit the damage in your AI features.

    7 min read

  • Vibe Coding

    Taking an AI-Built Prototype to Production

    Your AI-built prototype works and users are arriving. What to add before it holds real data: tests, CI/CD, environments, monitoring, backups, cost limits.

    5 min read

  • Vibe Coding

    10 Security Mistakes in Apps Built with AI Tools

    The security gaps we see most often in apps built with Cursor, Lovable, Bolt and Replit, and how to check your own app for each one.

    6 min read

Get practical AI insights, monthly

AI security, automation and governance, written by engineers. No spam, unsubscribe any time.

Book a free 30-min call