Control Packs · AI-001 · v1.1.0
Untrusted Content and Tool Separation
Treat external content as data, keep instructions and data separate, and restrict tool use when an AI component processes untrusted input.
Status: review · Review: not independent. This is design guidance; review status does not establish independent verification or compliance.
What this safeguard addresses
Separate instructions from untrusted data, constrain tools, ground high-consequence output, and require human review where the consequence warrants it.
- Untrusted content changes authorityExternal content is treated as an instruction or is allowed to influence tool selection and authority without a bounded policy.
Versioned applicability rule
{
"all": [
{
"characteristic": "AI_CONTROLLED",
"equals": true
},
{
"characteristic": "UNTRUSTED_INPUT",
"equals": true
}
]
}Decisions for the project owner
Leave a decision open when its value is unknown. Suggested values become confirmed only through an explicit user decision.
- Which AI outputs require human review before they can affect an external system?The model must not silently decide where human control is required.AI-001-Q1 · confirmation
Requirements for the coding agent
- AI-001-R1Mark untrusted content as data, enforce tool and output allowlists, prevent content from changing authority, and route high-consequence actions through the confirmed review boundary.
Implementation recipe
Keep trusted instructions, untrusted content, tool authority, and output approval as separate typed inputs to one policy-enforced execution boundary.
- Label external content as untrusted data before it reaches the model context.
- Resolve tool access from a server-side allowlist that content and model output cannot modify.
- Validate tool arguments and model outputs against explicit schemas and policy rules.
- Require the confirmed human-review state before any high-consequence external effect.
- Record the model, tool, review state, and outcome without raw prompts, secrets, or customer content.
const context = { instructions: trustedPolicy, data: markUntrusted(input) };
const proposal = await model.propose(context);
if (!toolPolicy.allows(proposal.tool, proposal.args)) return deny();
if (proposal.highConsequence && !review.confirmed) return holdForReview();
return invokeValidated(proposal);- Prompt wording alone is not an authority boundary; enforcement must occur outside the model.
Tests and evidence to retain
- PROMPT_BOUNDARY_TESTVerify instruction/data separation, tool allowlists, injection resistance, and review escalation.
- AI-001-V1 · Instruction injection remains dataThe content does not change trusted instructions, available tools, or approval requirements.Evidence: Rejected or bounded proposal; Unchanged allowlist snapshot; MODEL_ACTION_EVENT record
- AI-001-V2 · Disallowed tool or argumentsThe boundary rejects the call before external execution and records a policy denial.Evidence: Schema or policy rejection; Absence of external side effect; MODEL_ACTION_EVENT record
- AI-001-V3 · High-consequence review boundaryThe proposal remains provisional until the confirmed reviewer approves it.Evidence: Held proposal record; Review-state transition; External side-effect check
- configuration_or_policy
- implementation_location
- test_result
- telemetry_definition
Passing a published example shows that example's behavior. A coding agent's implementation report remains a claim until its evidence is independently checked.
Related guidance
- SI-10 · Information Input ValidationNIST_SP_800_53_5_2_0 · partially addressesInstruction/data separation and output allowlisting operationalize a slice of input validation; they do not establish control compliance.
- AC-6 · Least PrivilegeNIST_SP_800_53_5_2_0 · partially addressesTool and authority boundaries constrain automated privilege; they do not cover an organization’s complete least-privilege program.
- MAP 3.5 · Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the Govern functionNIST_AI_RMF_1_0 · partially addressesThe confirmed review boundary operationalizes a human-oversight outcome for this use case; it is not an AI RMF assessment.
Mappings indicate contextual relevance or partial support. They do not establish equivalence, certification, government endorsement, or complete framework implementation.
- OWASP Cheat Sheet Series — LLM Prompt Injection Prevention Cheat SheetGUIDANCE_AI_CONTENT_BOUNDARY
- National Institute of Standards and Technology — Security and Privacy Controls for Information Systems and OrganizationsGUIDANCE_NIST_SP_800_53_5_2_0
- National Institute of Standards and Technology — Artificial Intelligence Risk Management FrameworkGUIDANCE_NIST_AI_RMF_1_0