01
The system prompt carries the whole policy
Sensitive rules and permissions live in instructions the model can misunderstand, expose, or ignore.
AI tooling · Prompt injection controls
Prompt injection can't be solved with one instruction or filter. The stronger approach limits the model's authority, separates trusted instructions from outside content, enforces policy in code, and requires approval before high-impact actions.
Visible signals
Manipulated content matters most when it can cross a trust boundary, reveal protected context, misuse a tool, or trigger action without informed approval.
01
Sensitive rules and permissions live in instructions the model can misunderstand, expose, or ignore.
02
Documents, websites, messages, and tool results enter the same context as authorized directions.
03
Broad data access, shared credentials, or powerful tools extend far beyond the task.
04
People review a summary after the workflow has already sent, changed, disclosed, or deleted something.
05
A filter catches known jailbreak language but misses hidden instructions, tool misuse, or data exposure.
06
The organization lacks the source, model context, policy decision, tool calls, approval, and outcome.
Control layers
Model instructions can help, but they shouldn't carry the full security burden. Layer controls so one bad judgment can't create an unauthorized result.
Classify sources and preserve provenance so outside content doesn't inherit system authority.
Define permitted inputs, expected outputs, prohibited behavior, refusal, and escalation.
Limit credentials, records, tools, actions, destinations, and access duration.
Use application logic for authorization, validation, allowlists, data rules, and transaction limits.
Require approval before high-impact actions and show the source, action, system, and uncertainty.
Test direct, indirect, encoded, multimodal, and tool-based attacks against real consequences.
Control path
Follow a real interaction from input through action. At each step, decide what the system trusts, what the model can do, and which rule runs outside it.
01
Ingest
Source, trust level, provenance, format, transformation, retrieval path, and handling rules.
02
Interpret
Application intent, user authority, task boundary, outside content, conflicts, and escalation.
03
Decide
Policy, authorization, schema, sensitive data, destination, limits, and validation.
04
Act
Tool, permission, credential, transaction limit, human approval, reversibility, and effect.
05
Review
Trace, blocked attempt, approval, outcome, incident response, test result, and updated control.
Evidence
The evidence includes every source of instruction, every system the model can reach, the code enforcing policy, the people approving action, and the final outcome.
01
System directions, user input, templates, memory, retrieved context, tool results, and their order.
02
Web pages, files, messages, knowledge sources, images, metadata, and other untrusted material.
03
Functions, records, API scopes, destinations, limits, secrets, boundaries, and authorization code.
04
Responses, structured data, links, code, routing, system updates, and the people or services that rely on them.
05
Human checkpoints, displayed context, cancellation, rollback, escalation, and incident ownership.
06
Attack simulations, blocked attempts, false alarms, tool calls, approvals, incidents, and customer effects.
No foolproof prompt layer
The goal is to reduce exposure, enforce authority, detect unsafe behavior, and limit impact. One prompt or filter can't remove prompt injection risk.
Common failure patterns
The weakest designs ask the model to enforce rules that belong in identity, policy, application logic, permissions, approval, and incident response.
01
Input filtering stands alone even though unsafe influence can arrive indirectly or through ordinary-looking output.
02
The application marks content as untrusted but gives the model the same context and permissions to act on it.
03
Natural-language rules replace identity checks, scoped permissions, and policy enforced in code.
04
Approvers see a summary without the source, affected records, uncertainty, or reason for review.
05
Tests miss compromised documents, retrieved content, tool output, chained actions, and data theft.
06
New models, tools, sources, and permissions expand risk without a new review.
Practical outputs
The result shows how the workflow limits authority, validates important behavior, involves people, records evidence, and changes as threats evolve.
01
Instructions, outside content, model context, tools, data, people, systems, and unclear authority.
02
Task limits, source separation, least privilege, policy, validation, approval, and recovery.
03
Scenarios tied to real consequences, including indirect injection, tool misuse, data exposure, and unsafe action chains.
04
Trace requirements, thresholds, owners, containment, rollback, evidence, and control updates.
Engagement fit
This work fits a defined AI use case that processes outside content and can reach sensitive data, tools, or important actions. It supports, but doesn't replace, application security, penetration testing, privacy, legal review, or incident response.
Start a fit checkRelated diagnostic paths
You need to define the AI role, oversight, data controls, measures, and responsible implementation path.
Review the diagnosticThe interaction crosses customer workflows, service rules, systems, ownership, and handoffs.
Review the diagnosticThe model reads or changes CRM data, recommends customer action, or triggers automation.
Review the diagnostic