An AI agent that reads an email and can also search your customer database or send messages creates a security risk that a stronger prompt alone cannot solve. The key question is not only whether the agent can recognize suspicious instructions. It is what the agent could reach or do if content it reads changes its behavior.

That makes indirect prompt injection a workflow problem. To reduce the risk, map three things: the content the agent reads, the information it can access, and the actions its tools allow. Then look for unnecessary connections between them.

What is indirect prompt injection?

Indirect prompt injection occurs when instructions are hidden in content an agent processes—such as an email, webpage, shared document, or repository—instead of being given directly by the person using the agent.

For example, a customer email might include text telling an AI assistant to ignore its original task and send private information elsewhere. The email does not itself have permission to access the assistant’s tools. But the assistant may interpret the text while it is processing the email, and its connected tools determine what could happen next.

OWASP describes indirect prompt injection as a way untrusted external content can influence an AI system’s behavior. The important distinction is that the content is not a trusted instruction just because the agent can read it. Yet an agent may not reliably keep that boundary intact in every situation.

Why “ignore instructions in documents” is not enough

Telling an agent to treat emails and documents as untrusted is sensible, but it is not a complete security control. The agent still has to interpret those materials, and malicious or misleading instructions can be difficult to detect reliably. OWASP cautions against relying on prompt-based rules or detection systems as the only defense.

A useful way to think about this is to separate prevention from impact limitation. A filter or instruction attempts to prevent the agent from being influenced. Limiting permissions and available actions reduces what can happen if prevention fails.

This matters because the consequences depend on the whole workflow, not just the wording of the attack. An agent that summarizes a public webpage has a different risk from one that reads private email, searches customer records, and can send messages or modify business data.

NIST’s account of a large-scale red-teaming competition reports successful attacks against every frontier model targeted in that competition. That finding is not a prediction that every agent will be compromised. It is a reason not to treat a model’s ability to resist manipulation as a guarantee.

Trace the path from content to action

Before connecting an agent to business software, sketch its workflow in three parts:

| Part of the workflow | Questions to ask | |---|---| | Content | What can the agent read? Who can create or change that material? | | Access | What information can it retrieve beyond the content in front of it? | | Actions | What can its connected tools change, send, share, or delete? |

Then trace plausible paths between the three. For instance: an outside email contains an instruction → the agent can search the customer database → a messaging tool can send an email.

That path is more useful to examine than a list of suspicious phrases. The text might be obvious, subtle, or missed by a filter. The practical security question is whether the agent has a route from untrusted content to information or actions that should remain out of reach.

1. Map the content the agent reads

List every source the agent can process, including sources added indirectly. An agent that reads a forwarded email may also encounter attachments, quoted messages, or links. A document assistant may process content created by coworkers, customers, or outside collaborators.

For each source, ask who can put content there and whether the agent needs to follow links or retrieve more material. The goal is not to declare all external content unusable. It is to identify where instructions could arrive from people who should not control the agent.

2. Map the information the agent can access

Next, list the systems and records the agent can search or retrieve. Does an email assistant need access to every customer record, or only the information required to answer the current request? Does a document workflow need access to internal files unrelated to its task?

OWASP recommends applying least privilege: give a system only the access needed for its purpose. In practice, this means narrowing access by task, account, or data source where the tools allow it, rather than connecting the agent to a broad business account by default.

3. Map the actions its tools allow

Reading information and changing the outside world are different capabilities. Make a separate list of what each connected tool can do: view records, draft a message, send it, update a record, change a setting, or delete information.

Look for actions that are unnecessary for the agent’s job. If it only needs to prepare a response, it may not need a tool that sends messages. If it needs to find an invoice, it may not need the ability to change payment details. Where possible, use narrower permissions or separate tools that do not combine unrelated capabilities.

The aim is to reduce the number of paths from outside content to consequential action. A successful prompt injection should not automatically inherit every permission held by the account the agent uses.

Reduce connections before adding more filters

Consider an assistant that reviews incoming supplier emails. It can read messages, search the finance system, change supplier records, and send replies. A message containing a malicious instruction could potentially influence the agent while it has access to both sensitive information and tools that can act on it.

A safer design would examine which of those capabilities the task actually requires. Perhaps the agent needs to read the email and retrieve an invoice status, but not change supplier banking information. Perhaps it can prepare a reply without having the ability to send it. Those changes do not make prompt injection impossible; they reduce the consequences of a mistake or manipulation.

This is the central design principle: do not give one workflow more connected access and action capability than its job requires. Detection can be one layer, but the system should not depend on detection always working.

Test the workflow, not just the prompt

When evaluating an agent, test realistic routes through its workflow. For example, give it an ordinary task involving an email or document that contains an instruction to reveal unrelated information or take an out-of-scope action. Check not only whether the agent notices the instruction, but also whether its tools and permissions would allow the requested action.

Research such as AgentDojo evaluates prompt-injection attacks and defenses in tool-using agent tasks, including email-management scenarios. The practical lesson for a small business is modest but important: a prompt that behaves well in a clean example does not establish that the complete workflow is safe.

For each test, ask: What could this agent access? What could it change or send? Which connection made that possible? If a capability is not needed, removing it is generally a more dependable safeguard than hoping a future instruction or filter catches every attack.

The useful security question

To protect an AI agent from prompt injection, do not ask only, “Can it recognize malicious instructions?” Ask, “If content it reads influences it, what information and actions are reachable from that point?”

That question shifts attention from trying to make every email or document harmless to designing a workflow with fewer dangerous paths. Map the content, access, and actions; remove unnecessary connections; and treat filters and instructions as additional safeguards, not as the boundary holding the system together.

Sources