Control Packs · AI-001 · v1.1.0

Untrusted Content and Tool Separation

Treat external content as data, keep instructions and data separate, and restrict tool use when an AI component processes untrusted input.

Status: review · Review: not independent. This is design guidance; review status does not establish independent verification or compliance.

What this safeguard addresses

Separate instructions from untrusted data, constrain tools, ground high-consequence output, and require human review where the consequence warrants it.

  • Untrusted content changes authorityExternal content is treated as an instruction or is allowed to influence tool selection and authority without a bounded policy.
Versioned applicability rule
{
  "all": [
    {
      "characteristic": "AI_CONTROLLED",
      "equals": true
    },
    {
      "characteristic": "UNTRUSTED_INPUT",
      "equals": true
    }
  ]
}

Decisions for the project owner

Leave a decision open when its value is unknown. Suggested values become confirmed only through an explicit user decision.

  • Which AI outputs require human review before they can affect an external system?The model must not silently decide where human control is required.AI-001-Q1 · confirmation

Requirements for the coding agent

  • AI-001-R1Mark untrusted content as data, enforce tool and output allowlists, prevent content from changing authority, and route high-consequence actions through the confirmed review boundary.

Implementation recipe

Keep trusted instructions, untrusted content, tool authority, and output approval as separate typed inputs to one policy-enforced execution boundary.

  1. Label external content as untrusted data before it reaches the model context.
  2. Resolve tool access from a server-side allowlist that content and model output cannot modify.
  3. Validate tool arguments and model outputs against explicit schemas and policy rules.
  4. Require the confirmed human-review state before any high-consequence external effect.
  5. Record the model, tool, review state, and outcome without raw prompts, secrets, or customer content.
const context = { instructions: trustedPolicy, data: markUntrusted(input) };
const proposal = await model.propose(context);
if (!toolPolicy.allows(proposal.tool, proposal.args)) return deny();
if (proposal.highConsequence && !review.confirmed) return holdForReview();
return invokeValidated(proposal);
  • Prompt wording alone is not an authority boundary; enforcement must occur outside the model.

Tests and evidence to retain

  • PROMPT_BOUNDARY_TESTVerify instruction/data separation, tool allowlists, injection resistance, and review escalation.
  • AI-001-V1 · Instruction injection remains dataThe content does not change trusted instructions, available tools, or approval requirements.Evidence: Rejected or bounded proposal; Unchanged allowlist snapshot; MODEL_ACTION_EVENT record
  • AI-001-V2 · Disallowed tool or argumentsThe boundary rejects the call before external execution and records a policy denial.Evidence: Schema or policy rejection; Absence of external side effect; MODEL_ACTION_EVENT record
  • AI-001-V3 · High-consequence review boundaryThe proposal remains provisional until the confirmed reviewer approves it.Evidence: Held proposal record; Review-state transition; External side-effect check
  • configuration_or_policy
  • implementation_location
  • test_result
  • telemetry_definition

Passing a published example shows that example's behavior. A coding agent's implementation report remains a claim until its evidence is independently checked.

Related guidance

  • SI-10 · Information Input ValidationNIST_SP_800_53_5_2_0 · partially addressesInstruction/data separation and output allowlisting operationalize a slice of input validation; they do not establish control compliance.
  • AC-6 · Least PrivilegeNIST_SP_800_53_5_2_0 · partially addressesTool and authority boundaries constrain automated privilege; they do not cover an organization’s complete least-privilege program.
  • MAP 3.5 · Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the Govern functionNIST_AI_RMF_1_0 · partially addressesThe confirmed review boundary operationalizes a human-oversight outcome for this use case; it is not an AI RMF assessment.

Mappings indicate contextual relevance or partial support. They do not establish equivalence, certification, government endorsement, or complete framework implementation.

Continue your review