Labels: where data came from, and where it may go¶
Detection tries to recognise an attack in text, and a determined attacker can rephrase until it doesn't. Labels don't depend on recognising anything. Every piece of content the agent reads gets a label, the session keeps the most restrictive label of everything it has read, and each tool call is checked against it: may content from there drive this tool, and may data this sensitive reach it? The answer is decided in code, whatever the model was told.
The two axes¶
| Axis | Levels (low → high) | Meaning |
|---|---|---|
| Integrity | trusted → untrusted → hostile |
who could have written it; hostile means an injection was detected in it |
| Confidentiality | public → private → restricted |
how bad a leak would be; restricted means credentials, secrets or personal data |
Labels combine most-restrictive-wins: a session that has read one untrusted web page and one private customer record
is untrusted + private. Check it with session.state.label; every tool-call result and audit entry records it.
Where labels come from¶
- Capabilities (the default): results of network, exec, remote or unknown tools are
untrusted. - Detections: a secret or personal data raises confidentiality to
restricted; a detected injection makes integrityhostile. - Your declarations, for what capabilities can't know:
[labels.sources]
get_customer = { confidentiality = "private" } # business data, not a "secret" pattern
read_issue = { integrity = "untrusted" } # outsiders write issues
internal_kb = { integrity = "trusted" } # a vetted internal service
Declarations only raise a label; detections can raise it further.
Local data is trusted by default
Results of local, read-only tools (files, databases) count as trusted unless you say otherwise. If outsiders can
write that data (repositories, shared drives, tickets, uploads), set default_integrity = "untrusted" in [labels]:
every undeclared tool result then counts as untrusted. This default may change in a future release; run
guardlayer policy check to see what is assumed today.
What tools accept¶
Tools that act can declare what they are willing to run with:
[labels.sinks]
post_comment = { max_confidentiality = "public" } # a public channel: nothing private may reach it
write_file = { accepts_untrusted = false } # untrusted content must not drive it
send_money = { accepts_untrusted = false }
| Rule | Fires when | Default action |
|---|---|---|
confidentiality_exceeds_sink |
the call carries data from a source more sensitive than the tool's max_confidentiality |
review |
untrusted_to_protected_sink |
the session has read untrusted (or hostile) content and the tool has accepts_untrusted = false |
review |
"Carries" means the call's arguments contain the private source's identifiers (IDs, numbers, e-mail addresses), names (capitalised words), or a run of six words copied from it. A status note that shares none of these runs ("Done."); a summary in new words also runs (not detected). Private data that reaches the session without text to compare (a labelled file) keeps the stricter judgement: every call to the sink is held.
These add to the session rules you already have (sensitive_data_egress, trifecta, after_injection); change any
action in [session] actions.
Exact destinations and argument values¶
Allowing a domain allows everything on it. Argument rules say which values a tool may take:
[[tools.arguments]]
tool = "http_post"
argument = "url"
allow = ["https://api.github.com/repos/myorg/*"] # our repositories, not someone's gist
[[tools.arguments]]
tool = "send_email"
argument = "to"
allow = ["*@mycompany.com", "*@partner.example"]
action = "review" # or "block"
Values are matched case-insensitively as globs; recipient lists are checked one address at a time
("Asha <asha@mycompany.com>, b@outside.example"), and nested arguments are found. deny = [...] works the same way.
Destinations let some values receive more sensitive data than the tool's cap:
[labels]
sinks = { send_email = { max_confidentiality = "public" } }
destinations = [{ tool = "send_email", argument = "to", match = "*@mycompany.com", max_confidentiality = "private" }]
Internal recipients may get private data; one outside address in the same email caps the whole call at public.
Files keep their label¶
Experimental
File labels are new in 0.7 and haven't been tried outside tests and benchmarks. It may change or be removed in a later release, depending on how it works in real use.
An agent could write a script while reading an untrusted page, then run it with a command that looks harmless. So a file
written while the session's label is above trusted/public keeps that label:
- a later call that mentions the file (reading, uploading or running it) raises the caller's label to it, in any session, including a new one the next day;
- running a file written in an untrusted context needs review (
untrusted_file_executed).
A write that needed approval is recorded only after it actually ran (the Claude Code hook and guard_tool report it;
elsewhere call guard.record_written(tool, arguments, session=...)). File labels live next to the session state: a
file-labels.json beside on-disk sessions, in memory otherwise.
Disguised copies of secrets¶
Secrets the session has seen are remembered as fingerprints, never in the clear. An outgoing call is checked for the
secret as-is, with separators removed (s k - p r o j ...), and in base64, base64url, hex and URL-encoded form. Any of
those blocks the call (sensitive_data_egress), even when nothing untrusted was read. Paraphrased or summarised
information can't be fingerprinted; confidentiality labels cover that case instead.
Images, PDFs and other non-text content¶
Experimental
Extraction is new in 0.8. It may change or be removed in a later release, depending on how it works in real use.
Tool results that are bytes, or MCP content blocks with images, audio or embedded files, aren't scanned as if they were
text. Readable formats are extracted and scanned like any other content: PDFs with pip install "guardlayer[extract]",
images by OCR with pip install "guardlayer[ocr]" plus the Tesseract program. Add your own extractor with
guard.extractors.append(fn), where fn(data: bytes, mime: str) -> str | None.
Whatever can't be read (audio, video, binaries, images without OCR) is recorded as unreadable_content and makes the
session untrusted, unless the tool is declared trusted. An instruction hidden in an image nobody could read still
can't drive a protected action or carry private data out.
Check your configuration¶
guardlayer --config guardlayer.toml policy check # every tool named in the config
guardlayer --config guardlayer.toml policy check --tools send_email read_file
guardlayer policy check --claude-code # Claude Code's built-in tools
For each tool it shows its capabilities (declared, inferred or unknown), what its output counts as, what it accepts as
a sink, and whether its egress is limited, then warns about gaps: unknown capabilities, unlimited egress, output assumed
trusted, private data declared but sinks uncapped, network tools marked trusted, failing open. --strict exits 1 on any
warning, for CI; --json is for tooling.
Claude Code¶
The hook labels shell output (BashOutput) and sub-agent reports (Task, Agent) as untrusted, scans them, and
reports writes so file labels work across sessions. Your [labels] sources win over these defaults.