Scope and fit
AI features add interfaces and trust boundaries but remain part of an application. A useful test plan begins with the data and actions the model can influence.
Map model and application components
Identify model provider or runtime, prompt construction, retrieval stores, plugins, tools, user roles, and output consumers. Mark where untrusted text can influence a privileged action.
Test data exposure and instruction handling
Use benign, controlled cases to probe cross-user retrieval, prompt injection, sensitive output, unsafe tool invocation, and policy bypass. Avoid real secrets or external targets unless explicitly authorized.
Evaluate controls around the model
Review connector permissions, output validation, rate limits, logging, and human approval. OWASP's GenAI resources can help structure test ideas; they do not establish model safety or regulatory approval.
Decisions and tradeoffs
Use this table as a working review record. Replace assumptions with evidence from the target environment.
| Decision area | Working guidance |
|---|---|
| Map model and application components | Identify model provider or runtime, prompt construction, retrieval stores, plugins, tools, user roles, and output consumers. Mark where untrusted text can influence a privileged action. |
| Test data exposure and instruction handling | Use benign, controlled cases to probe cross-user retrieval, prompt injection, sensitive output, unsafe tool invocation, and policy bypass. Avoid real secrets or external targets unless explicitly authorized. |
| Evaluate controls around the model | Review connector permissions, output validation, rate limits, logging, and human approval. OWASP's GenAI resources can help structure test ideas; they do not establish model safety or regulatory approval. |
Implementation questions
What should the team decide about map model and application components?
Identify model provider or runtime, prompt construction, retrieval stores, plugins, tools, user roles, and output consumers. Mark where untrusted text can influence a privileged action. Use a named owner and a written acceptance check so this decision can be reviewed after deployment.
What should the team decide about test data exposure and instruction handling?
Use benign, controlled cases to probe cross-user retrieval, prompt injection, sensitive output, unsafe tool invocation, and policy bypass. Avoid real secrets or external targets unless explicitly authorized. Use a named owner and a written acceptance check so this decision can be reviewed after deployment.
What should the team decide about evaluate controls around the model?
Review connector permissions, output validation, rate limits, logging, and human approval. OWASP's GenAI resources can help structure test ideas; they do not establish model safety or regulatory approval. Use a named owner and a written acceptance check so this decision can be reviewed after deployment.
Plan, build, verify, operate
Map model and application components: Identify model provider or runtime, prompt construction, retrieval stores, plugins, tools, user roles, and output consumers. Mark where untrusted text can influence a privileged action. Record the result and the next owner before changing the next boundary.
Test data exposure and instruction handling: Use benign, controlled cases to probe cross-user retrieval, prompt injection, sensitive output, unsafe tool invocation, and policy bypass. Avoid real secrets or external targets unless explicitly authorized. Record the result and the next owner before changing the next boundary.
Evaluate controls around the model: Review connector permissions, output validation, rate limits, logging, and human approval. OWASP's GenAI resources can help structure test ideas; they do not establish model safety or regulatory approval. Record the result and the next owner before changing the next boundary.
Deployment checks
Turn the page into a reviewable handover by assigning each check to a person and retaining its result.
Plan an AI Application Security Test Around Data and Tool Boundaries: decision 1
Write down the boundary, owner, dependency, and proof required for plan an ai application security test around data and tool boundaries before implementation begins.
Plan an AI Application Security Test Around Data and Tool Boundaries: decision 2
Write down the boundary, owner, dependency, and proof required for plan an ai application security test around data and tool boundaries before implementation begins.
Plan an AI Application Security Test Around Data and Tool Boundaries: decision 3
Write down the boundary, owner, dependency, and proof required for plan an ai application security test around data and tool boundaries before implementation begins.
Bound the AI red-team exercise
AI application red-team planning must scope the model, prompts, retrieval, tools, agents, identity roles, data sources, environments, rate limits, and external actions. A surprising model response is not enough; assess whether it reaches a protected record, bypasses an authorisation check, triggers a tool, or creates another concrete security condition.
Use test accounts and synthetic data in a segregated environment. State prohibited actions, handling for harmful content, tool permissions, observation, escalation, and stop authority. Never treat a model as an authorisation engine; enforce sensitive actions with deterministic controls outside the model.
DeployOpen can conduct the approved exercise. The customer owns model use policy, data permissions, production change, and acceptance of residual risk.
Handover and ownership
Before handover, name the system owner, support path, access boundary, backup or recovery responsibility, and the condition that pauses a change.
Keep a short record of what was tested, what remains outside scope, and when the review should happen again.

