Skip to content

GuardLayer threat model

What GuardLayer protects, what it assumes, what it can't stop, and how GuardLayer itself could be attacked. Read this before relying on it. GuardLayer lowers risk; it does not make prompt injection impossible.

1. What GuardLayer protects

Asset Threat GuardLayer control
The model's instructions Direct prompt injection and jailbreaks in user input scan_input: signature rules, de-obfuscated views, similarity, optional classifier / LLM judge
The agent's context Indirect injection in RAG chunks, web pages, emails, tool results scan_context, scan_tool_result; session marked untrusted or hostile
Secrets and personal data Leaks in prompts, outputs, tool arguments Secrets + PII scanners (redaction), secret_in_egress, sensitive_data_egress
The system prompt Extraction and leakage Canary tokens, prompt-leak overlap scanner
The host and its data Destructive or persistent commands, credential-file access Tool-call policy: capabilities, destructive / risky / persistence / credential rules
The network boundary Exfiltration to tunnels, request catchers, metadata endpoints, raw IPs Egress rules + optional egress allow-list
The user (in chat UIs) Markdown-image and link exfiltration Link scanner on output
Decisions after the fact Silent tampering with the record Hash-chained, optionally Ed25519-signed audit log; guardlayer audit verify

2. Trust boundaries and assumptions

  • Trusted: the application code that calls GuardLayer, GuardLayer's configuration, and the operator who approves REVIEW verdicts.
  • Untrusted: everything the model reads from outside (user input when not trusted, web, email, RAG, tool and MCP results) and everything the model decides (tool calls are judged before they run).
  • Assumptions GuardLayer depends on:
  • The application actually calls the scan functions on every edge, and passes a session where taint tracking is wanted.
  • Tools are tagged with correct capabilities. An untagged tool is assumed able to do anything, but a tool mis-tagged as read-only skips rules that would otherwise apply. The reverse costs utility: on AgentDojo's untagged banking tools, legitimate payments were blocked because the IBAN in them counted as personal data leaving the machine.
  • Human approvers read REVIEW requests. Rubber-stamping defeats the control ("review fatigue"). Measured in the agentic evaluation: with every review approved, 4 of 30 attacks succeeded (destructive and persistence actions), against 0 with reviews denied. Exfiltration stayed at 0 either way, because secrets are redacted and fingerprinted before egress.
  • The configuration file and the process environment are not attacker-writable.

3. What GuardLayer can't stop (residual risk)

Gap Why Mitigation
Paraphrased injections Signature and similarity layers catch known and near-known phrasing; held-out recall is ~0.23 rules-only, ~0.47 with the classifier; on real adaptive attacks (LLMail-Inject, held-out attacker teams) that hijacked a model, 44.5% rules-only and 56% with the classifier Architecture first: least-privilege tools, egress allow-list, REVIEW for consequential actions, decisions that don't read free text
Attack families seen only once AgentDojo's 0 / 10 attack success came after rules were added for its injection families; a new template can still get through (a plain TODO-style goal is undetectable by design) Treat the post-fix numbers as a closed gap, not a detection rate; rely on the session rules and REVIEW for consequential actions
Encoded or split secrets sensitive_data_egress matches verbatim copies (including embedded ones), not base64 or split values trifecta rule escalates untrusted + sensitive + outbound regardless of value matching
Semantic leaks Summarised or paraphrased sensitive information isn't fingerprintable Keep secrets out of the context; restrict what the agent can read
Cross-process taint Session state is per store; separate processes need a shared FileSessionStore Use the file store (or equivalent) when checks run in separate processes
Ordinary-domain egress Without egress_allowlist, only known-dangerous destinations are blocked Set tools.egress_allowlist for agents with sensitive access (or use airgap)
Fail-open by default In balanced, a scanner error lets text through fail_closed = true (set by strict and airgap)
Audit truncation A hash chain can't show lines removed from the end Store the reported head hash elsewhere; verify with --expected-head
Observe mode Nothing is enforced or redacted Use only while measuring false positives

Each preset states its own residual risk: guardlayer presets.

4. Attacks on GuardLayer itself

Threat Status
ReDoS (catastrophic regex backtracking from input) In scope for security reports. Inputs are size-limited by the limits scanner (50,000 chars by default); fingerprint matching is linear in argument length
Supply chain Core has zero runtime dependencies. Optional extras (ml, embeddings, api, signing, integrations) pull third-party packages, so install only what you use
Model tampering (optional classifier) The default model's upstream project is archived; GuardLayer pins it to an exact revision so a changed upstream can't silently alter verdicts. Each detection records model + revision. Pin custom models with revision
Leaking scanned text Audit log stores hashes by default (include_text opt-in); session state stores fingerprints (length, prefix check, truncated SHA-256), not values
REST API abuse Set GUARDLAYER_API_KEY beyond localhost and run behind TLS
Claude Code hook The hook can only tighten: it returns deny or ask, never allow, so it can't widen Claude Code's own permissions. It fails open unless fail_closed is set
Rule-pack poisoning Custom rule packs and configs are trusted input; protect them like code

5. Reporting

See SECURITY.md. Detection bypasses are expected for heuristic layers; report them as issues with the sample added to a dataset so the fix can be measured. Crashes, ReDoS, authentication bypasses and data leaks are security reports.