Resources · 44

Prompt injection: test an AI agent’s trust boundaries

Check that a document, search result or message cannot authorise an action on behalf of the user.

· 3 min

Method diagram: External source → Untrusted data → Agent → Tool control → Verified effect Method diagram · steps explained in the text

What this guide helps achieve

  • Separate data from instructions
  • Limit available actions to the task
  • Test real effects behind responses
  • Document limitations and regressions

Quick check

  • Which external content does the agent read?
  • Which actions can it trigger without approval?
  • Is authorisation enforced outside the model?
  • Does a verbal refusal actually prevent a tool call?
  • How can a diverging scenario be stopped?

Step-by-step method

  1. 01

    Inventory the boundaries

    List messages, web pages, documents, search results, attachments and tool outputs. Treat them as data to interpret without authority over permissions. Include channels carrying text or multimodal content.

    Deliverable: map of inputs and available actions.

  2. 02

    Define forbidden effects

    Describe harms to prevent: out-of-scope access, unapproved publication, data transfer or account changes. Set an observable outcome for each scenario; judging only the answer text can miss an action already performed.

    Deliverable: threat scenarios and invariants.

  3. 03

    Isolate the test environment

    Use clearly labelled fictitious accounts, documents and destinations with test markers and simulated tools. Disable real sending and production access. Retain model, configuration, corpus and connector versions to replay cases.

    Deliverable: versioned environment and test set.

  4. 04

    Replay adversarial inputs

    Place in a test source an instruction that contradicts the task, requests an out-of-scope action or asserts false authority. Vary language, placement, format and source combinations. Measure tool invocation, arguments and actual effect.

    Deliverable: traces of successful, blocked and unresolved cases.

  5. 05

    Strengthen external controls

    Enforce minimal permissions, argument validation, read/write separation and allowed destinations in the tools. Require contextual approval for sensitive operations. RAG and a refusal instruction alone do not remove injection risk.

    Deliverable: executable rules and bypass tests.

  6. 06

    Track changes

    Replay cases after changes to model, corpus, tools or rules. Also track legitimate tasks that get blocked. Publish tested scope and known limitations; success on a small test set does not establish the absence of vulnerabilities.

    Deliverable: regression report and deployment decision.

Worked example

Illustrative situation

Illustrative situation: an assistant summarises a test document that requests sending information to an external destination.

Decision and expected evidence

The tool rejects that destination independently of generated text. The test checks that nothing was sent and the legitimate summary still works.

Distinguish the mechanisms

MechanismPurposeCheck or limitation
Input filterDetect some suspicious patternsMay miss new wording
Model instructionDescribe the expected boundaryIs not an access control
Tool enforcementReject an unauthorised operationMust check identity, object and arguments

Management indicators

IndicatorWhat it measuresFirst action
Forbidden actionsCases causing out-of-scope effectsBlock and investigate every effect
False refusalsLegitimate tasks preventedRevise without broadening permissions
CoverageChannels and tools actually testedDeclare blind spots

Common pitfalls

  • Putting secrets in test prompts
  • Testing visible answers only
  • Assuming internal documents are always trusted
  • Claiming absolute protection after a few tests

Frequently asked questions

Does RAG prevent injection?

No. Retrieved documents can themselves carry adversarial instructions. Sources remain data to process within controlled boundaries.

Should every suspicious document be blocked?

It depends on the task; an agent can often analyse a document without following its instructions or having write tools.

How should results be published?

State versions, scope, effects tested, limitations and fixes. Exclude secrets and client documents used in the work.

Official references