Changelog¶
All notable changes to this project are documented here. The format follows Keep a Changelog and the project uses Semantic Versioning.
[Unreleased]¶
Changed¶
- The destination check is on by default (
[session] untrusted_destination = "outbound"; set"off"to disable). An outbound action whose destination came from content an outsider can write is held when it carries something private (a body or message, or URL words that are in neither that content nor your prompts). It needs no injection detection. Measured with detectors removed: recorded attacks stopped by the defaults 5/27 -> 12/27; extra holds on real Claude Code sessions: 1 in 7,572 calls, 0 in 6,962 from other projects. Following links, opening files a page listed, and opening this machine's dev servers or files are not held. Hold messages name the recipient and where it came from. - Shell output and file runs are judged by what actually happens. A shell command's output is local content
unless the command can read from outside (network programs, fetching modules, unanalysable code); files a command
downloads keep their outsider origin;
untrusted_file_executedjudges files that are run, not merely mentioned, and judges them by what they would execute even after an injection was detected; a declared tool that can't read anything doesn't make the session untrusted. Real Claude Code sessions from other projects: 2.3% -> 1.6% held; OpenHands agents' real work: 58% -> under 1% out of the box, 0.2% with tools declared; recorded attacks unchanged. trifectajudges where data can actually go. Its outbound check now uses the destination check's places (real destination arguments, never this machine, nothing held for a link the agent merely follows); an irreversible action counts only if it can send data out (publish, push, share, send; not a local delete). Real Claude Code sessions from other projects: calls held 8.5% -> 2.3% (trifecta 499 -> 14 of 6,962); recorded attacks unchanged.- A tool's name no longer establishes trust (
[labels] default_integrity = "declared", the new default): a result counts as trusted only if you vouched for who writes it (output = "trusted",trusted_tools, a source'sintegrity = "trusted") or an integration did for its own tools (Claude Code's Read, Grep, Glob), and it is local. Declaring capabilities doesn't confer trust, so pasting apolicy draftdoesn't quietly restore it."trusted"restores the old name-based behaviour. On AgentDojo's published frontier-model runs, undetected attacks stopped: 423 -> 648 of 707; extra holds on real Claude Code sessions: 0. Declare your own tools (guardlayer policy draftproposes lines). - Irreversible actions: every argument counts (the outsider choosing the file to delete or hotel to book), and invite participants are recipients.
- An outsider's address can't be laundered through a file: trusted content that repeats an address an outsider already supplied no longer makes it "known"; only your own messages can.
confidentiality_exceeds_sinkjudges what a call carries: after reading a declared-private source, a call to a public sink is held when it carries that source's identifiers, names or a verbatim run, not for every call ("Done." now runs). Paraphrase is not detected.- New docs page What's stable: which features are core, which are add-ons, and which are experimental (task profiles, file labels, split-instruction detection, extraction, behavioural check). Experimental modules say so in their first lines. A dead-code scan (vulture, counting tests and examples as users) found nothing to remove.
[0.8.2] - 2026-10-02¶
Draft your tool declarations from what your agent actually did.
Added¶
guardlayer policy draft AUDIT.jsonl: drafts[tool.NAME]declarations from what the agent actually used. Every MCP server getsoutput = "untrusted"; each tool gets its guessed capabilities, andoutput_datawhen secrets or personal data were seen in its output, with lines markedCHECKwhere judgement is needed. Tools GuardLayer already knows (Claude Code built-ins, your config) aren't redeclared. Adopting the draft never weakens a tool: a remote tool given explicit read-only capabilities is kept remote (remote = true). Prompted by ADR-Bench, where undeclared, malicious tool servers looked exactly like normal ones.
[0.8.1] - 2026-10-02¶
Fewer false alarms on real agent tool output. Replaying ADR-Bench (Uber: 303 recorded sessions with 134 MCP servers) showed 0.8.0 interrupting about one normal session in six. Fixes were studied on half of the sessions and checked on the other half: 18% → 4% of normal sessions interrupted on the held-out half. Detection on the public datasets, LLMail-Inject and the Indian-language set is unchanged.
Changed¶
- Personal data in tool output makes a session private, not sensitive. Exact copies of those values leaving the
machine are still blocked (
sensitive_data_egress) and declared sinks still enforce their limits, but personal data alone no longer triggerstrifecta. In recorded AgentDojo runs that rule, on personal data alone, interrupted 12 of 130 normal tasks and was the only thing that stopped 1 of 130 attacks. Secrets still make a session sensitive.strictandairgapkeep the old behaviour ([session] trifecta_on_pii = true).
Fixed¶
- Decimal numbers were read as card, phone or Aadhaar numbers (
31.41592653589793,9724.1633798299). - Leetspeak decoding turned numbers into letters, so IP addresses such as
1.1.1.1looked like split letters (i.i.i.i). Tokens without letters are no longer decoded. dangerous_schemefired on "Audio File: x.wav", "Created file: ..." andfile://in error messages.file:now counts only as a link target (markdown orhref/src);javascript:/vbscript:only when code follows.
[0.8.0] - 2026-10-02¶
Simpler and faster. One place to describe each tool, a Claude Code hook that answers in about 10 ms, Indian-language detection in the core, and a README that says what is proven and what isn't. Several new features ship as experimental: they work and are tested, but haven't been used outside tests and benchmarks yet.
Added¶
- Hook server for Claude Code (
guardlayer hook claude-code --server). The command hook starts Python for every event (about 0.6 s on Windows, before and after each tool call). The hook server is a long-running local GuardLayer that Claude Code calls through its HTTP hooks, with the same checks. Before a tool call: 608 ms → 10 ms on Windows.--server --print-configprints the settings, including aSessionStarthook (--ensure-server) that starts the server in the background and warns you if it can't. Loopback only; refuses browser-style requests (anOriginheader, a non-loopbackHost, a non-JSON body); optional bearer token (--token-env); an edited config is picked up on the next event. If the server isn't running, Claude Code lets tool calls through: the command hook remains for where that matters more than speed. [tool.NAME]: everything about one tool in one place:capabilities,output(trusted or untrusted),output_data,accepts_untrusted,max_data,may_send,remote,arguments,destinations. Until now whether a tool's output is trusted could be said in four places, and where it may send data in six. It expands into the existing[tools],[session]and[labels]settings, so both forms work together. Unknown keys fail with the list of known ones.- Indian-language detection in the core (
ignore_instructions_indic): "ignore (all) previous instructions" in Hindi, Marathi, Bengali, Gujarati, Punjabi, Tamil, Telugu, Kannada, Malayalam and romanised Hindi, in either word order. No false positives on 2,056 Wikipedia introductions in these languages (256 chosen because they use the rule's own words) or on 4,913 other benign texts. - ONNX runtime for the classifier (
runtime = "onnx",multilingualextra: onnxruntime and tokenizers, no PyTorch). It runs a model directory you downloaded and checked; nothing is fetched for you. Measured with the newbenchmarks/classifier_eval.pyon Horizon Labs' 30-languageprompt-injection-guard-small: much higher recall on unseen public attacks (deepset 14 → 32 of 60), but new false alarms on normal agent data (11 of 404 AgentDojo tool outputs, 92 of 238 LLMail emails), so it is documented, not recommended for blocking (Recipes → Other languages).
Added, experimental¶
- Task profiles (
[tasks.NAME],guard.session(id, task=..., task_args=...)): the tools a task may use, and argument values bound to the trusted request ({task.customer_email}). Outside the profile:out_of_taskandtask_argument_not_allowed(review). A session's task can only narrow without an approval. - Instructions split across two pieces of content (
split_injection): the end of the previous untrusted content is scanned together with the start of the next. Only a redacted 500-character tail is kept in the session. - PDFs and images in tool results (
guardlayer.extract;extractandocrextras): text is extracted and scanned; content that can't be read (unreadable_content) makes the session untrusted. - Behavioural check (
check_intent,acheck_intent,needs_intent_check): replays the conversation through your own model with the user's request hidden; the same action proposed anyway isinjection_driven_action(review). On AgentDojo banking with a local 7B model it flagged nothing (0 of 30 replays) while 4 of 10 attacks succeeded: without a task, the model only summarised. Measure it on your own agent before relying on it.
Changed¶
strictandairgaptreat every tool result as untrusted unless you declare the tool trusted ([tool.NAME] output = "trusted").balancedis unchanged. On AgentDojo banking and Slack this gave the same task success and attack success, with one extra review on normal banking tasks.- A tool you list in
remote_toolsyourself is now remote even when you also give it explicit capabilities. Before, the listing was silently ignored; the built-in patterns (*search*,mcp__*, ...) still give way to explicit capabilities. import guardlayerno longer loads the compliance module, and the similarity scanner loads its attack corpus on first use, so a process that checks one tool call starts faster.from guardlayer import build_evidenceworks as before.- The README is rewritten around what GuardLayer is for, with an evidence table that says how much to trust each number; the full evaluation moved to the documentation.
Fixed¶
- Zero-width false positives on Indian-language, Persian and emoji text. The zero-width joiner and non-joiner are part of correct spelling in these scripts, but were counted as hidden characters: 105 of 2,056 normal Indian-language texts were flagged. A joiner after a letter of such a script no longer counts; after a Latin letter it still does. Now 3 of 2,056.
What changes on upgrade¶
- With
strictorairgap, expect more reviews in sessions that read local files: their output now counts as untrusted. Declare your own content trusted with[tool.Read] output = "trusted"(or[labels] default_integrity = "trusted"to restore 0.7 behaviour). - Indian-language content with an instruction override is now flagged; Indian-language text with zero-width joiners is flagged far less.
- Session files written by 0.8 carry new fields; 0.7 ignores them, so downgrading keeps working.
- New rules (
out_of_task,task_argument_not_allowed,split_injection,unreadable_content,injection_driven_action) only fire when you use the feature behind them, exceptsplit_injectionandunreadable_content, which apply to sessions that read untrusted content.
[0.7.0] - 2026-09-30¶
Labels. The same gap tests before and after (fake data, harmless instructions): with 0.6.3 all five containment
gaps were allowed; with 0.7.0, disguised secrets are blocked and write-then-run needs review out of the box, and the
other three are closed by a [labels] / [[tools.arguments]] configuration (guardlayer policy check shows where).
The scripted agentic suite is unchanged (0/30 attacks, 7/8 benign tasks, 1 benign review), also with
default_integrity = "untrusted".
Added: labels (information-flow control)¶
- Labels on everything the agent reads (
guardlayer.labels): integrity (trusted<untrusted<hostile) and confidentiality (public<private<restricted), combined most-restrictive-wins; the session's context label is on every tool-call result and audit entry. Docs: Concepts, "Labels and information flow". [labels]sources, sinks, destinations, default_integrity: declare what a tool returns (get_customeris private,read_issueis untrusted) and what a tool accepts (max_confidentiality,accepts_untrusted). New rulesconfidentiality_exceeds_sinkanduntrusted_to_protected_sink(review by default) fire only for declared sinks. Destinations let matching argument values receive more (internal recipients may get private data).default_integrity = "untrusted"makes every undeclared tool result untrusted.[[tools.arguments]]: allow/deny globs for one argument (recipients, URL paths, repos); recipient lists are split and display names dropped.- File labels: a file written while the session's label is above trusted/public keeps that label; a later call that
mentions it, in any session, inherits the label, and running it needs review (
untrusted_file_executed). Reviewed writes are recorded only after they ran (GuardLayer.record_written, called by the Claude Code hook andguard_tool). - Normalised fingerprints: remembered secrets also match with separators removed and in base64, base64url, hex and URL-encoded form.
guardlayer policy check: per tool, capabilities (declared, inferred or unknown), output label, sink limits and egress limits, with plain-language warnings;--strictfor CI,--json,--claude-code.- Docs tests now parse every TOML example and load its session, labels and tools sections.
These close the five containment gaps found in the 2026-09-29 gap analysis (poisoned local file, business data not recognised as sensitive, leaks through allowed channels, disguised copies of secrets, write-then-run), each reproduced as a test.
What changes on upgrade¶
Without a [labels] section most behaviour is unchanged, but three protections apply by default:
- a disguised copy of a remembered secret (spelled out, separators removed, base64, hex or URL-encoded) in an outgoing
call is now blocked (sensitive_data_egress); before, it was allowed, or held for review only if untrusted content
had been read;
- running a file that was written after the session read untrusted content now needs review (untrusted_file_executed);
- in the Claude Code hook, shell output (BashOutput) and sub-agent reports (Task, Agent) now count as untrusted,
so the usual session rules apply after them. Override with [labels] sources.
Set any of these rules to "log" in [session] actions to observe instead of enforce while you evaluate.
Added¶
- AgentDojo results for strip mode (banking and Slack, today's rules): attacks 0 / 10 in both modes, but attacked tasks
didn't recover (banking 5 / 10 either way; Slack 0 / 10, where 32 of 33 poisoned results still fell back to withholding).
Strip stays opt-in. Results in
benchmarks/results/agentdojo-qwen2.5-coder-7b-strip-2026-09-29.jsonl.
[0.6.3] - 2026-09-29¶
Added¶
guardlayer audit report: what GuardLayer decided, or in observe mode would have decided, by rule and by tool, with the latest notable entries and a count of redactions (--since-days,--min,--json). Built for pilots: run it daily and sort each entry into correct, false alarm or unsure.- Docs: "One-week pilot on your own work": install from PyPI into its own environment, observe-mode config, hook on one project, a five-minute daily review, when to start enforcing, and what to report back.
[0.6.2] - 2026-09-29¶
Security¶
- Fixed five ReDoS (catastrophic backtracking) paths that let a single crafted input up to the 50,000-character limit
cost minutes of CPU: rules
fake_role_header(blank-line runs) andfake_system_marker(runs of#), thelimitsscanner's dialogue-turn counter (blank-line runs), thesecretsscanner's genericNAME_PASSWORD=pattern and URL-credential pattern (a-a-a-...runs), and thepiiscanner's email pattern (long runs without@). All now linear: the worst of 40 adversarial 50,000-character inputs takes 1.5 s through the whole guard (seven previously exceeded two minutes). Detection is unchanged except that email local parts over 64 characters (invalid per RFC 5321) no longer match. Newtests/test_redos.pytimes every rule and the whole guard at two input sizes and fails on super-linear growth. - Redaction now removes every copy of a detected secret or personal-data value, not only the matched span. A key
glued to other text (
0AKIA...) didn't match its pattern and survived next to a redacted copy of the same key (found by the new property tests). Linear single pass; variable names such asDB_PASSWORDstay readable. tests/test_properties.py: property-based tests (Hypothesis) run random and attack-fragment inputs through every scan method, tool calls, tool results and strip mode, checking no crash, valid scores and spans, serialisable results, and that redacted secrets don't survive.
Added¶
on_injection="strip"forguard_tool, LangGraphguard_toolsand the OpenAI Agents SDKguardrails: cut the injected part out of a tool result instead of withholding all of it (strip_injections: from the first to the last flagged line, widened to an enclosing tag pair such as<INFORMATION>...</INFORMATION>). Only with a session, only for string results, and only when the cut is under 80% of the text, every detection has a location and the remainder scans clean; otherwise withheld as before. Default unchanged ("withhold"). AgentDojo harness:guardlayer-stripdefense. Not a sanitizer: on 765 detected held-out LLMail-Inject attacks, 496 were still withheld and 269 stripped, and in 117 of those 269 the attacker's target address survived the cut (llmail_eval.py --strip-check); the session's review of the next action is what keeps strip safe.- Detection-rule referee (
benchmarks/referee.py): a candidate rule pack ships only if it (1) detects more held-out LLMail-Inject attacks, (2) adds no hits on any benign set (4,509 public prompts, AgentDojo environment texts, LLMail benign emails), (3) fires on the dev attacks it was written from, and (4) names nothing specific to the challenge's goal. Every rule is also judged alone, and every run, accepted or rejected, is appended tobenchmarks/results/referee.jsonl. - Held-out split for LLMail-Inject by attacker team (60% of 99 teams held out, fixed on 2026-09-29 before any miss was
inspected; an email sent by any held-out team is held out).
llmail_eval.pynow reports held-out numbers separately. benchmarks/agentdojo_benign_texts.py: exports AgentDojo's benign environment texts (404) for false-positive tests.- Five context rules accepted by the referee (round 1, written from 60 sampled dev-team misses):
forged_chat_turn,forged_safety_verdict,summary_anchored_action,split_letter_obfuscation,agent_goal_statement. On held-out teams, rules-only detection of attacks that hijacked the model went from 10.4% to 44.5% (1,719 emails), and of those that also evaded the challenge's defenses from 9.3% to 38.5% (161); still no hits on 5,151 benign texts. A sixth candidate,forged_tool_call, was rejected (dev +2, held-out +0). Typed input is unaffected (public benchmark unchanged).
[0.6.1] - 2026-09-29¶
First release published to PyPI (0.5.0 and 0.6.0 were tagged on GitHub only; their publish step failed before trusted publishing was set up).
Security¶
- Hardened the project's own build pipeline: every GitHub Action pinned to a full commit SHA; read-only default token
in all workflows;
persist-credentials: falseon checkouts; release tooling pinned (build,twine); the third-party release action replaced with the runner'sgh; concurrency limits; Dependabot for actions and pip with a 7-day cooldown; aworkflow-securityCI job running zizmor. Documented in SECURITY.md ("How releases are built").
Added¶
benchmarks/llmail_eval.py: a held-out test on Microsoft's LLMail-Inject (phase 2, 38,014 unique real attacker emails, MIT). GuardLayer detects 17.4% of the attacks that hijacked the model with rules only and 47.0% with the classifier (50.4% of those that also evaded the challenge's defenses); no false positives on 238 benign emails. Never used for tuning.- AgentDojo banking with
allow_egressfor the payment tools: benign utility 4 -> 5 / 10, no benign blocks, attack success still 0 / 10. [session] allow_egress: tool-name glob -> data types that tool may send out ({ send_money = ["iban"] }). Those types, for that tool only, no longer triggersensitive_data_egress, nortrifectawhen they are the only sensitive data in the session;after_injectionstill applies. Fingerprints now carry the rule that found the value (<len>:<prefix>:<sha>:<kind>) and sessions recordsensitive_kinds; untyped fingerprints from older session files are never exempt. Found by AgentDojo: legitimate payments to an IBAN from a bill were blocked.
[0.6.0] - 2026-09-27¶
Added¶
- Control-mapped compliance evidence (
guardlayer.compliance,guardlayer evidence export | controls). Verifies a hash-chained audit log, then maps every entry to the controls it evidences: OWASP Top 10 for LLM Applications 2026 (and 2025 IDs), OWASP Top 10 for Agentic Applications 2026, MITRE ATLAS, ISO/IEC 42001 Annex A (A.6.2.6, A.6.2.8), NIST AI RMF (MEASURE 2.4, 2.7, MANAGE 4.1) and EU AI Act Art. 12, 14 and 15. Exports JSONL (header with verification result, source SHA-256 and head hash; one record per entry; per-control summary), CSV (one row per entry x control) or a text summary. Refuses unverified logs unless--allow-unverified. Mapping version2026.09. 13 tests intests/test_compliance.py. benchmarks/perf.py: latency (p50/p95/p99) for every guard edge at 200 to 48,000 characters, multi-process throughput, memory, and a REST API load test (--api). Standard library only; deterministic payloads. Results inbenchmarks/results/.DEPLOYMENT.mdanddeploy/: deployment shapes, a hardened Compose file, Kubernetes manifests (ConfigMap, Deployment with non-root/read-only/no-capabilities, Service, deny-egress NetworkPolicy, HPA, PodDisruptionBudget; strict-validated against Kubernetes 1.31), sizing, sessions across replicas, audit-log storage, rollout, and measured performance.[audit] pathaccepts{hostname}and{pid}, so each worker or pod owns its own hash-chained file.- CSA AI Controls Matrix v1.1.1 in the evidence export (
--framework csa-aicm, mapping version2026.09.2), verified against CSA's official spreadsheet: log records (LOG-09), input and output monitoring (LOG-15/16), sanitized logs (LOG-08, when the entry holds hashes only), guardrails (TVM-13), input/output validation (AIS-09/10), prompt differentiation (AIS-15), agent boundaries and access (AIS-11, IAM-18), sensitive data (DSP-10/17), credentials (IAM-14), human supervision (GRC-15); audit log protection (LOG-02) only when the log verifies. IDs and titles are referenced with attribution; no control text is included. - NYDFS 23 NYCRR Part 500 in the evidence export (mapping version
2026.09.17), from DFS's published amended text: 500.6(a)(2) on every entry, 500.14(a)(2) for injections in content the agent reads (web, email, tool results), 500.14(a)(1) and 500.7(a)(1) for tool policy and session taint. - DORA in the evidence export (mapping version
2026.09.16), from the Commission's adopted RTS text: RTS 2024/1774 Art. 12(1) on every entry, Art. 12(2)(d) when the log verifies, DORA Art. 10(1) on detections, RTS Art. 21(a)/(d) for tool policy, RTS Art. 11(2)(i) wherever data leaving is blocked. - NIS2 in the evidence export (mapping version
2026.09.15): Implementing Regulation 2024/2690 annex 3.2.1 on every entry and 3.2.5 when the log verifies; Directive Art. 21(2)(b) on detections; Art. 21(2)(i) and annex 11.1.1 for tool policy. Numbers checked against ENISA's Technical Implementation Guidance. - FedRAMP 20x Key Security Indicators in the evidence export (mapping version
2026.09.14), from FedRAMP's Consolidated Rules 2026.09.13.02: KSI-MLA-LET on every entry, KSI-IAM-ELP for tool policy, KSI-CNA-RNT for egress rules. Process KSIs (reviews, SIEM operation, incident response) aren't claimed. Rev. 5 authorisations use the SP 800-53 evidence. - CMMC 2.0 Level 2 in the evidence export (mapping version
2026.09.13), practice IDs and titles from the DoD CMMC Assessment Guide Level 2 v2.13: AU.L2-3.3.1 on every entry, AU.L2-3.3.8 when the log verifies, SI.L2-3.14.6 on detections, AC.L2-3.1.1/3.1.2/3.1.5 for tool policy, AC.L2-3.1.3 for session taint and data leaving, SC.L2-3.13.1 for egress rules and SC.L2-3.13.6 for allow-list blocks, SI.L2-3.14.2 for blocked persistence. - PCI DSS v4.0.1 in the evidence export (mapping version
2026.09.12): 10.2.1 on every entry, 10.3.4 when the log verifies, 3.4.1 for card numbers masked in output, 7.2.5 for tool policy, 1.3.2 for allow-list egress blocks. 3.5.1 is deliberately not claimed (the log's text hash is unkeyed). - GDPR in the evidence export (mapping version
2026.09.11), only where personal data is involved: Art. 5(1)(f), 25(1) and 32(1)(b) for detected and redacted personal data; Art. 5(1)(c) and 25(2) for audit entries that keep hashes rather than text (not claimed withinclude_text=True). Art. 32(1)(a) is not claimed: a hash isn't pseudonymisation. - HIPAA Security Rule in the evidence export (mapping version
2026.09.10): 164.312(b) audit controls on every entry, 164.312(a)(1) access control for tool policy (the rule covers software programs), 164.312(e)(1) transmission security for blocked data leaving, 164.308(a)(6)(ii) and 164.308(a)(1)(ii)(D) on detections, 164.308(a)(5)(ii)(B) for blocked persistence. Checked against the eCFR (2026-09-24). Relevant only where ePHI is handled. - SOC 2 (AICPA Trust Services Criteria 2017) in the evidence export (mapping version
2026.09.9): CC7.2 on every entry, CC7.3 on detections, CC6.1/CC6.3 for tool policy, CC6.6 for injections in outside content, CC6.7 for blocked data movement, CC6.8 for blocked persistence, C1.1 for secrets and personal data. Criterion IDs checked against the AICPA's 2022 revised edition; descriptions are GuardLayer's own. - ISO/IEC 27001:2022 Annex A in the evidence export (mapping version
2026.09.8): A.8.15 Logging and A.8.16 Monitoring activities on every entry, A.5.33 Protection of records when the log verifies, A.8.11 Data masking and A.5.34 for redacted PII, A.8.12 Data leakage prevention for blocked egress and output leaks, A.5.15 Access control for tool policy, A.8.3 for credential files, A.8.23 Web filtering for egress rules. No input-validation or human-oversight claim: the 2022 Annex A has no such control. - NIST CSF 2.0 in the evidence export (mapping version
2026.09.7): PR.PS-04 and DE.CM-09 on every entry, PR.DS-01 when the log verifies, PR.AA-05 for tool policy, PR.DS-02 for data leaving (egress rules, session taint), PR.DS-10 for secrets and personal data redacted before the model, PR.PS-05 for blocked persistence, DE.AE-06 for reviews. Subcategory text from NIST's CSF 2.0 export, matched exactly. - NIST SP 800-53 Rev. 5.2.0 in the evidence export (mapping version
2026.09.6), titles from NIST's OSCAL catalog: AU-2/AU-3/AU-12 on every entry; AU-9 and AU-9(3) only when the hash chain verifies and AU-10 (non-repudiation) only when Ed25519 signatures verify; SI-4 on detections; SI-10 input validation; SI-15 output filtering; AC-3/AC-6 for tool policy; AC-4 for session taint and secrets leaving; SC-7 for egress rules and SC-7(5) when an allow-list blocks; SC-5 for size limits. AC-3(2) dual authorization is deliberately not claimed for single-approver reviews. - ETSI EN 304 223 V2.1.1 in the evidence export (mapping version
2026.09.5): the European Standard (2025-12) that supersedes ETSI TS 104 223. It renumbers the UK Code's provisions (5.4.2-1/-2 logging and analysis, 5.1.4-1/-3 human oversight, 5.1.2-6 permissions, 5.2.1-4 and 5.2.1-4.1 sensitive data and input checks) and adds 5.1.2-2 (withstanding adversarial attacks), mapped to blocked injections and jailbreaks. Provision numbers only, checked against ETSI's PDF. - UK Code of Practice for the Cyber Security of AI (2025) in the evidence export (mapping version
2026.09.4, Open Government Licence v3.0): 12.1 logging, 12.2 behaviour analysis, 4.1/4.3 human oversight, 2.6 least-privilege permissions for the AI system, 5.4 sensitive data, 5.4.1 input checks and sanitisation. - MITRE ATLAS mitigations and OWASP AISVS 1.0 in the evidence export (mapping version
2026.09.3). ATLAS mitigations (v2026.09): M0020, M0024, M0028, M0029, M0030, M0033, M0036. AISVS 1.0: C2.1.2–C2.1.8, C7.3.2–C7.3.4, C9.2.1, C9.3.5, C9.5.1/C9.5.3/C9.5.4, C12.1.2, C12.2.1, C12.2.3. Human-in-the-loop mappings apply only to reviews of agent tool calls. IDs verified against the official sources; AISVS descriptions are GuardLayer's own. - Two more held-out datasets in
benchmarks/public_eval.py: Lakera's Gandalf injections (1,000, recall 0.57) and SPML (16,011 prompts: precision 1.00, recall 0.21, no false positives on 3,470 benign prompts). Downloads are now atomic, retry with back-off, and skip empty rows. benchmarks/agentdojo_eval.py: GuardLayer as a defense on AgentDojo (ETH Zurich's third-party agent benchmark), with a local model;guardlayer-untrusteddefense (every tool result untrusted),--max-iters, and the GuardLayer commit recorded with each result. Results (qwen2.5-coder 7B, banking and Slack, 10 attacks each): attacks succeeded 7 → 6 (banking) and 4 → 3 (Slack) with 0.5.0, and 0 / 10 in both after the detection rules below, measured after seeing the attacks. Benign utility drops by 2 of 10 tasks per suite; in banking, legitimate payments to an IBAN were blocked as personal data leaving the machine (the tools are untagged).- Detection rules from AgentDojo's attack families: typo-tolerant "ignore previous instructions", content addressed to "the AI", instructions posed as a precondition of the user's task, fake system markers in content. 4 of AgentDojo's 5 families detected at text level (was 1); no false positives on 4,509 benign prompts and 1,006 benign AgentDojo environment texts.
- Agentic evaluation (
benchmarks/agentic_eval.py): 30 injection attacks (5 attacker goals x 3 injection styles) and 8 benign tasks in a simulated workspace, run by a real model through Ollama or by a scripted worst-case agent that obeys every injection. Scores executed actions (hijacked / succeeded / utility / approvals asked), with and without GuardLayer, with reviews denied or rubber-stamped. The scripted suite runs in CI (tests/test_agentic_scripted.py).
Fixed¶
- Prefixed secret names weren't redacted. The generic assignment rule needed a word boundary before the name, so the usual
.envforms (DB_PASSWORD=,POSTGRES_PASSWORD=,AWS_SECRET_ACCESS_KEY=,GITHUB_TOKEN=) were missed and could be sent out. Found by the agentic evaluation: it was the only way a secret leaked when every review was rubber-stamped. - GuardLayer's own redaction marker was re-flagged as a secret (
DB_PASSWORD=[REDACTED:GENERIC_SECRET]), so an agent forwarding redacted text raised a redundantsecret_in_egressreview. read_email-style tools were inferred as network-capable, so reading a mailbox after PII had entered the session raised a falsetrifectareview (5 approval requests on 8 benign tasks in the agentic evaluation, down to 1). Read verbs on messaging nouns (email, mail, inbox, slack, sms, message) now inferreadonly; their results still count as untrusted. URL, web and API tools keepnetwork, andwebpage/website/urinames now infernetworktoo (get_webpage(url)can carry data out).- Similarity scanner coverage of long texts. Its window budget stopped at the first 64 windows (about 4,000 characters), so a known attack in the middle or at the end of a long page or document was never compared. Windows are now spread across the whole text (overlapping by one sentence), and the default budget is 256. On known attacks inserted at five positions in benign documents, similarity-layer recall went from 45/75 to 75/75 at 8,000 characters, 27/75 to 72/75 at 20,000 and 15/75 to 58/75 at 48,000. Costs up to ~80 ms more on the longest inputs. Public benchmark results are unchanged.
Changed¶
- README speed claim corrected from "~1 ms per scan" to measured figures: ~1.4 ms for a 200-character prompt, ~0.2 ms for a shell tool call, and roughly 10–13 ms per 1,000 characters for longer inputs.
- Faster heuristics on text without leetspeak or encoding (identical de-obfuscated views are skipped) and a small speed-up in similarity search (cached feature hashes). Results are identical.
- Dockerfile:
/var/log/guardlayerowned by the service user; documented--read-onlyrun.
Changed¶
- README threat-coverage table now uses the OWASP Top 10 for LLM Applications 2026 numbering, with 2025 IDs alongside.
[0.5.0] - 2026-09-27¶
First release published to PyPI (pip install guardlayer).
Security¶
- The default classifier model is pinned to an exact revision (
90c9989b1a342275dd0d1a95aad283c04e075671). Its upstream project was archived in July 2026 and is no longer maintained; a floating reference could change verdicts silently. NewClassifierScanner(revision=...)/[scanners.classifier] revisionpins custom models too;revision=Noneopts out. Each classifier detection recordsmodelandrevisionin its metadata. Tests intests/test_classifier_pinning.py.
Added¶
THREAT_MODEL.md: assets, trust boundaries, assumptions, residual risk and attacks on GuardLayer itself. Linked from README and SECURITY.md.- Release workflow: tagged versions build and publish to PyPI through trusted publishing (no stored tokens).
[0.4.1] - 2026-09-27¶
Security fixes from a review of 0.4.0. Each bypass was reproduced first and has a regression test in tests/test_remote_egress.py.
Fixed¶
- Remote tools with read-only names no longer escape taint tracking. Results from
search,tavily_search,get_webpageormcp__github__get_issuedid not mark the session untrusted, sotrifectanever fired. An undetected injection on such a page, a.envread and an exfiltrating call came out ALLOW. NewToolPolicy.is_remote(): a tool is remote when it can reach the network or run commands, is untagged, or matchesremote_tools. Defaults:mcp__*,*search*,*web*,*page*,*url*,*issue*,*github*,*mail*and similar. Explicit capabilities still win, so Claude Code's built-in tools are unchanged. sensitive_data_egressnow finds secrets embedded in longer text. Before, a secret seen earlier matched only as a whole token, sohttps://evil.example/<key>,/log/<key>.pnganddata=x<key>got past it. Fingerprints now also store the length and a 16-bit prefix check, and matching slides over every run of token characters in linear time (a 64 KB argument with 50 fingerprints takes about 50 ms). Session files written by 0.4.0 still match whole tokens. It also covers remote tools whose arguments leave the machine, such as a search query.- Secrets in outgoing tool arguments are no longer just redacted. The secrets scanner only rewrote
result.text, while every integration ran the tool with the original arguments, so the secret left and the verdict was ALLOW. New tool rulesecret_in_egress(default review, configurable throughrule_actionsanddisabled_rules) fires when a remote tool's arguments contain a secret.
Added¶
[tools] remote_tools/ToolPolicy(remote_tools=..., include_default_remote_tools=...). Tool-call results carrymetadata["remote"].
[0.4.0] - 2026-09-24¶
Sessions and integrations: an agent action is judged by what the session has already seen, and GuardLayer plugs into Claude Code, LangGraph and the OpenAI Agents SDK.
Added¶
- Session taint tracking (
guardlayer.session).guard.session(id)/GuardSession, orsession=on everyscan_*call. A session records untrusted content (output of network-capable or untagged tools,scan_context), hostile content (an injection was found in it) and sensitive data (secrets or PII read or pasted, credential and.envaccess). Three rules escalate tool calls: sensitive_data_egress(block): a secret seen earlier appears in a network or exec call.trifecta(review): untrusted content and sensitive data, then a network or exec call.after_injection(review): an injection was read, then a write, network or exec call.- Sensitive values are kept only as truncated SHA-256 fingerprints.
SessionPolicy:trusted_tools,untrusted_tools, per-ruleactions,enabled.- Session stores:
MemorySessionStore(LRU plus idle timeout) andFileSessionStore.FileSessionStoreuses a per-session lock file, atomic replace and merge-on-write, so parallel processes never lose taint, including on Windows. - Claude Code hook:
guardlayer hook claude-codehandles PreToolUse, PostToolUse and UserPromptSubmit. It returnsdenyorask(neverallow), flags injected tool output to Claude, tags Claude Code's built-in tools, checks only the target path of Write/Edit, and keeps file-backed sessions keyed bysession_id. It fails open, or closed withfail_closed.--print-configprints the settings.json snippet. - LangGraph / LangChain:
guard_tools(guard, tools). A REVIEW verdict becomesinterrupt(), which you resume withCommand(resume=True). The graph'sthread_idbecomes the session. - OpenAI Agents SDK:
guardrails(guard)returns input, output, tool-input and tool-output guardrails. The session comes from the run context'ssession_id. guard_tool: wraps any sync or async tool function with a pre-call check and a post-call result scan, a refusal orToolBlocked, anapprovecallback for REVIEW, and output withholding.- Config: a
[session]section (store,dir,ttl_seconds,max_sessions,trusted_tools,untrusted_tools,actions,enabled) and aGUARDLAYER_STATE_DIRenv var.strictblocksafter_injection;airgapalso blockstrifecta. - REST:
session_idon the scan endpoints,POST /v1/scan/tool-result,GET/DELETE /v1/sessions/{id}. - Capabilities:
ToolPolicy.resolve()andcan_act(). An explicit empty capability list marks a tool as harmless. - New extras:
langgraphandopenai-agents. A new CI job runs the integration tests with both frameworks installed.
Changed¶
scan_tool_callskips the content scanners for tools tagged read-only, whose arguments cannot cause harm (for example, a search for "rm -rf"). Passscan_content=Trueto force the scan.asynciois imported lazily, cutting import time by about 35%.
[0.3.0] - 2026-09-24¶
"Agent Guard": GuardLayer now governs what an agent is about to do, not just what text says.
Added¶
- Tool-call policy (
guardlayer.tools.ToolPolicy, used byscan_tool_call). Tools carryread/write/network/execcapabilities, set explicitly (glob patterns allowed) or inferred from the tool name. Adds allow- and deny-lists with globs, per-capability actions (capability_actions={"exec": "review"}), customToolRules,rule_actionsoverrides anddisabled_rules. About 0.1 ms per call. - Built-in tool rules:
destructive_command(block),risky_command(review),persistence(review),credential_file(block) anddotenv_file(review). - Egress control for network- and exec-capable tools:
egress_metadata_endpoint(block),egress_exfil_servicefor tunnels, request catchers, OAST and file drops (block),egress_not_allowedagainstegress_allowlist(block) andegress_raw_ipfor public IPs (flag). Private and loopback addresses are not egress. REVIEWverdict and action: hold an action until a human approves it. Verdicts are ordered ALLOW < FLAG < REVIEW < BLOCK, andScanResult.needs_reviewis new.Detection.action: a rule can carry its own action, which takes precedence over the category action.- Observe (shadow) mode:
Policy(mode="observe"), plus per-ruleobserve/enforceglob lists. Results gainshadow_verdict,observed_rulesandeffective_verdict. Observed detections are neither enforced nor redacted, but they still count towardscore. - Presets (
guardlayer.presets):observe,balanced,strictandairgap, each with a stated residual risk. Available throughGuardLayer.from_preset(),preset = "..."in config,--preseton the CLI orGUARDLAYER_PRESET. - Tamper-evident audit log:
AuditLoggerhash-chains entries (seq,prev_hash,entry_hash), continues an existing chain on restart and can sign entries with Ed25519 (signer=, newsigningextra).verify_audit_log()reports the first bad line and the head hash. It takes anexpected_headto catch truncation. - Config:
[tools]section (allowlist,denylist,capabilities,capability_actions,rules,rules_file,rule_actions,disabled_rules,egress_allowlist),[audit]section,[guard] mode/observe/enforce, and aGUARDLAYER_MODEenv var. - CLI:
tool-call,presets,audit verify,audit keygen,--preset, andreviewas a--fail-onlevel.rulesnow lists the tool rules too. - REST:
/v1/scan/tool-callacceptsmetadataand returns capabilities./v1/settingsreports the preset, mode and tool policy.
Changed¶
ScanResult.allowedis now false for REVIEW as well as BLOCK, and@guard.protectstops on either.AuditLoggerfilters oneffective_verdict, so observe-mode results that would have been blocked are still logged. Chaining is on by default. An existing unchained log file must be replaced with a new file.[guard] tool_allowlistmoved to[tools] allowlist. The old key still works.examples/agent_tools.pyshows block, review and allow, and writes a verifiable audit log. It also no longer crashes on Windows consoles.
Also in this release (previously unreleased)¶
benchmarks/public_eval.py: a reproducible benchmark on deepset/prompt-injections and jackhhao/jailbreak-classification. Results are in the README.- 10 more heuristic rules (47 total):
forget_everything,change_instructions,new_instructions_follow,prompt_beginning, instruction override / new-instruction / prompt-extraction rules for German, Spanish, French, Portuguese, Italian, Dutch, Russian and Croatian/Serbian,unethical_ai_persona,has_no_rulesandjailbreak_marker.
Changed¶
- The override rule now also covers orders, tasks, assignments and information.
disable_safetyalso covers "OpenAI/Anthropic/company policy". - The similarity search uses an inverted index for sparse vectors, which is about 2x faster on long prompts with identical results.
- Held-out recall with zero false positives: deepset 0.08 → 0.23, jailbreak-classification 0.66 → 0.72.
ClassifierScannerclassifies long texts in overlapping chunks (head and tail kept, up tomax_chunks) instead of truncating at 512 tokens, so an injection at the end of a long document is still seen.benchmarks/public_eval.pygains--classifier,--classifier-only,--thresholdand--splits, and scans each sample only once.- The README benchmark table covers the classifier: deepset held-out recall 0.23 (rules) → 0.47 (rules + classifier), with precision still 1.00.
Fixed¶
load_samplesandguardlayer batchno longer split JSONL records on Unicode line separators (U+2028, U+0085) inside strings.
[0.2.0] - 2026-09-23¶
A rebuild from a single-detector prototype into a complete input/output guard layer.
Added¶
- Three directions:
scan_input,scan_outputandscan_context(indirect injection through RAG chunks, web pages and tool results). - Policy engine (
Policy,Action): per-category and per-direction actions (score/block/flag/redact/log), fail-open or fail-closed on scanner errors, a configurable redaction format. - Sanitized output:
ScanResult.texthas redactions applied, andScanResult.modifiedmarks when that happened. - Scanners:
ObfuscationScanner,SimilarityScanner(with a dependency-free vector store and auto-learning),SecretsScanner,PIIScanner(Luhn / mod-97 / Verhoeff validation),CanaryScanner,PromptLeakScanner,LinkScanner,LimitsScanner,DenyListScanner,ClassifierScanner(optional),LLMJudgeScanner,RelevanceScanner(optional). - Heuristics: expanded to 37 rules with categories and direction scoping. Every rule also runs against de-obfuscated views (homoglyphs, leetspeak, spaced letters, zero-width characters, Unicode-tag smuggling, base64/hex/percent/rot13 payloads). Custom rule packs load from JSON or TOML.
- Canary tokens for leak detection and goal-hijack (echo) detection.
- Agent support:
scan_tool_callwith a tool allow-list, andscan_tool_result. @guard.protectdecorator for sync and async LLM calls, plusGuardBlocked.- Async API (
ascan,ascan_input,ascan_output,ascan_context), hooks, andAuditLogger(JSONL that stores text hashes by default). - Configuration from TOML/JSON with environment-variable overrides (
GuardLayer.from_config). - Evaluation harness (
guardlayer eval) and a bundled labelled sample. - CLI:
scan,batch,eval,canary,rules,serve,--configand--fail-on. - REST API v1: input, output, context, batch, tool-call, canary, corpus and settings endpoints, with optional API-key auth.
- Dockerfile, a CI matrix covering Python 3.10–3.13 and Windows, a core-only install smoke test, a package build, and examples.
Changed¶
- The
Scannerprotocol is nowscan(text, context: ScanContext)and scanners declaredirections. Detectiongainscategoryandmetadata.ScanResultgainsid,timestamp,text,modified,latency_ms,timings_ms,errorsandmetadata.- Aggregation counts each rule once, so repeated hits no longer inflate the score.
[0.1.0] - 2026-09-12¶
- Initial release: heuristic scanner, noisy-or pipeline, CLI and a minimal REST API.