Python API¶
Generated from the source docstrings. Everything here is importable from guardlayer unless noted.
The guard¶
guardlayer.GuardLayer ¶
GuardLayer(
scanners: Sequence[Scanner] | None = None,
*,
policy: Policy | None = None,
flag_threshold: float | None = None,
block_threshold: float | None = None,
canaries: CanaryManager | None = None,
auto_learn: bool = False,
tool_allowlist: Iterable[str] | None = None,
tool_policy: ToolPolicy | None = None,
session_policy: SessionPolicy | None = None,
sessions: SessionStore | None = None,
hooks: Iterable[Callable[[ScanResult], None]] = (),
)
Filter LLM inputs, outputs and third-party context through a configurable scanner ensemble.
scan_input ¶
scan_input(
prompt: str,
*,
session: str | GuardSession | None = None,
**context_fields: Any,
) -> ScanResult
Scan a user prompt before it reaches the model.
scan_output ¶
scan_output(
response: str,
*,
prompt: str | None = None,
system_prompt: str | None = None,
canary: Canary | str | None = None,
session: str | GuardSession | None = None,
**context_fields: Any,
) -> ScanResult
Scan a model response before it reaches the user (or a tool).
scan_context ¶
scan_context(
content: str,
*,
source: str | None = None,
session: str | GuardSession | None = None,
**context_fields: Any,
) -> ScanResult
Scan third-party content (RAG chunk, web page, email, tool result) for indirect injection.
With a session, the content marks the session as having read untrusted content
(and hostile / sensitive content, if found).
scan_tool_call ¶
scan_tool_call(
tool_name: str,
arguments: Mapping[str, Any] | str | None = None,
*,
session: str | GuardSession | None = None,
scan_content: bool | None = None,
**context_fields: Any,
) -> ScanResult
Scan a model-proposed tool call (name + arguments) before executing it.
Runs the tool policy (allow/deny lists, capability actions, argument and egress rules)
and, with a session, the taint rules (what the session has already read decides what
it may do next). The content scanners also run over the arguments, except for tools
tagged read-only (scan_content=None, the default), whose arguments cannot cause harm.
A REVIEW verdict means: ask a human first.
scan_tool_result ¶
scan_tool_result(
tool_name: str,
result: Any,
*,
session: str | GuardSession | None = None,
arguments: Mapping[str, Any] | str | None = None,
**context_fields: Any,
) -> ScanResult
Scan what a tool returned before the model reads it (indirect injection channel).
With a session, results from network-capable (or untagged) tools mark the session
untrusted, injections mark it hostile, and secrets/PII mark it sensitive. Pass the call's arguments so a
file the agent wrote after reading untrusted content is read back as untrusted, whatever tool reads it.
needs_intent_check ¶
Is check_intent worth a model call for this tool call?
Only for tools that can act (network, exec, write, or untagged), and, with a session, only once the session has read untrusted content: before that, nothing but the user can be driving the agent.
check_intent ¶
check_intent(
tool_name: str,
arguments: Mapping[str, Any] | str | None,
*,
messages: Sequence[Mapping[str, Any]],
replay: Replay,
session: str | GuardSession | None = None,
task: str = NEUTRAL_TASK,
**context_fields: Any,
) -> ScanResult
Behavioural hijack check (see guardlayer.intent): replay the conversation with the user's request hidden.
replay(masked_messages) calls your model and returns the tool calls it proposes. If it proposes this same
action anyway, the action is driven by content the agent read: injection_driven_action (review by default).
Language- and wording-independent; one extra model call, so use needs_intent_check to pick risky calls.
acheck_intent
async
¶
acheck_intent(
tool_name: str,
arguments: Mapping[str, Any] | str | None,
*,
messages: Sequence[Mapping[str, Any]],
replay: Callable[..., Any],
session: str | GuardSession | None = None,
task: str = NEUTRAL_TASK,
**context_fields: Any,
) -> ScanResult
check_intent with an async replay.
scan ¶
scan(
text: str,
direction: Direction = "input",
*,
context: ScanContext | None = None,
**context_fields: Any,
) -> ScanResult
Scan one text. Extra keyword args populate the ScanContext.
scan_batch ¶
scan_batch(
texts: Iterable[str],
direction: Direction = "input",
**context_fields: Any,
) -> list[ScanResult]
session ¶
session(
session_id: str | None = None,
*,
task: str | None = None,
task_args: Mapping[str, Any] | None = None,
) -> GuardSession
A view of this guard bound to one session, so tool calls are judged by what came before.
task puts the session under a task profile (see guardlayer.tasks), with task_args from the trusted request.
protect ¶
protect(
func: F | None = None,
*,
system_prompt: str | None = None,
on_block: str = "raise",
blocked_message: str = "Sorry, I can't help with that request.",
) -> Any
Decorate fn(prompt, *args, **kwargs) -> str (sync or async) with input and output filtering.
The wrapped function receives the (possibly redacted) prompt, and the caller gets the
(possibly redacted) response. On a BLOCK it raises GuardBlocked, or returns
blocked_message when on_block="message". Non-string responses pass through unscanned.
add_canary ¶
Embed a canary token in a (system) prompt; pass the returned Canary to scan_output.
from_preset
classmethod
¶
Build a guard from a named preset (see guardlayer.presets), with optional config overrides.
from_config
classmethod
¶
Build a guard from a TOML/JSON file path or a config dict (see guardlayer.config).
guardlayer.GuardBlocked ¶
Bases: Exception
Raised by protect-wrapped calls when a prompt or response is blocked (or held for review).
Results¶
guardlayer.ScanResult
dataclass
¶
ScanResult(
verdict: Verdict,
score: float,
direction: Direction,
detections: list[Detection] = list(),
text: str = "",
modified: bool = False,
id: str = (lambda: uuid.uuid4().hex)(),
timestamp: float = time.time(),
latency_ms: float = 0.0,
timings_ms: dict[str, float] = dict(),
errors: list[str] = list(),
metadata: dict[str, Any] = dict(),
shadow_verdict: Verdict | None = None,
observed_rules: list[str] = list(),
)
guardlayer.Detection
dataclass
¶
Detection(
scanner: str,
rule: str,
category: str,
severity: float,
message: str,
span: tuple[int, int] | None = None,
metadata: dict[str, Any] = dict(),
action: str | None = None,
)
A single signal raised by one scanner.
guardlayer.Verdict ¶
Bases: str, Enum
Overall decision for a scanned piece of text.
guardlayer.Action ¶
Bases: str, Enum
What the pipeline does with a detection of a given category.
guardlayer.Category ¶
Bases: str, Enum
Threat categories. Detections carry the string value, so custom scanners may add their own.
Policies¶
guardlayer.Policy
dataclass
¶
Policy(
flag_threshold: float = 0.4,
block_threshold: float = 0.8,
actions: dict[str, Action] = (
lambda: dict(DEFAULT_ACTIONS)
)(),
fail_closed: bool = False,
redaction_format: str = "[REDACTED:{rule}]",
mode: Literal["enforce", "observe"] = "enforce",
observe: list[str] = list(),
enforce: list[str] = list(),
)
How detections become a verdict.
actions maps a category — or a "direction:category" pair, which takes
precedence — to an Action. Unlisted categories are scored. A detection that carries
its own action (tool-policy rules do) uses that instead.
mode = "observe" records everything but enforces nothing (shadow mode). observe
lists detections to only observe even in enforce mode; enforce lists detections to
keep enforcing in observe mode. Entries are globs matched against the rule name,
scanner:rule and the category, e.g. "heuristics:*", "egress_raw_ip", "pii".
is_observed ¶
True when this detection is recorded but not enforced.
guardlayer.ToolPolicy ¶
ToolPolicy(
*,
allowlist: Iterable[str] | None = None,
denylist: Iterable[str] = (),
capabilities: Mapping[str, Iterable[str]] | None = None,
capability_actions: Mapping[str, Action | str]
| None = None,
rules: Iterable[ToolRule | Mapping[str, Any]] = (),
include_default_rules: bool = True,
disabled_rules: Iterable[str] = (),
rule_actions: Mapping[str, Action | str] | None = None,
egress_allowlist: Iterable[str] | None = None,
block_exfil_services: bool = True,
flag_raw_ips: bool = True,
infer: bool = True,
remote_tools: Iterable[str] = (),
include_default_remote_tools: bool = True,
arguments: Iterable[
ArgumentRule | Mapping[str, Any]
] = (),
)
Evaluate a proposed tool call against allow/deny lists, capability actions, argument rules and egress rules.
is_remote ¶
True when the tool talks to something outside this machine: its results are untrusted
content and its arguments leave the machine. Network/exec-capable and untagged tools are
remote; so are inferred or untagged tools matching remote_tools (e.g. mcp__*,
*search*), even when their names sound read-only. Explicit capabilities win over the default patterns, not
over tools you list in remote_tools yourself.
is_declared ¶
Whether the tool's capabilities were declared (config, or an integration's own tools), not guessed.
reads_nothing ¶
Declared with no way to read anything (no read, network or exec capability), e.g. a "think" or "finish" tool: its output can only echo the agent, so nobody outside can have written it. Inferred capabilities don't count: an unknown tool may read anything.
vouches ¶
Whether an integration vouches for this tool as its own (see vouched).
resolve ¶
(capabilities, tagged). An explicit empty list tags a tool as harmless; untagged tools match every rule.
can_act ¶
True if the tool may have side effects or reach the network (anything beyond reading).
guardlayer.ToolRule
dataclass
¶
ToolRule(
name: str,
action: Action | str = Action.BLOCK,
pattern: str | None = None,
tools: tuple[str, ...] | None = None,
capabilities: frozenset[str] | None = None,
category: str = Category.TOOL_MISUSE.value,
severity: float = 0.9,
message: str = "",
flags: int = re.IGNORECASE,
)
A rule over tool calls.
Fires when the tool matches tools (glob patterns; None = any tool), the tool has one of
capabilities (None = any; untagged tools always match), and pattern (a regex over the
call's argument text; None = always) is found.
guardlayer.SessionPolicy
dataclass
¶
SessionPolicy(
enabled: bool = True,
actions: dict[str, Action] = (
lambda: dict(DEFAULT_SESSION_ACTIONS)
)(),
untrusted_tools: list[str] = list(),
trusted_tools: list[str] = list(),
hostile_min_verdict: Verdict = Verdict.FLAG,
allow_egress: dict[str, list[str]] = dict(),
sources: dict[str, dict[str, Any]] = dict(),
sinks: dict[str, dict[str, Any]] = dict(),
default_integrity: str = "declared",
destinations: list[dict[str, Any]] = list(),
tasks: dict[str, Any] = dict(),
trifecta_on_pii: bool = False,
after_injection_scope: str = "consequence",
trifecta_scope: str = "destination",
consequences: dict[str, str] = dict(),
destination_args: dict[str, list[str]] = dict(),
untrusted_destination: str = "outbound",
)
What counts as untrusted, and what to do when a tainted session tries to act.
untrusted_tools: tool-name globs whose results count as untrusted. By default a tool result is untrusted when the tool can reach the network or is untagged;scan_contextinput is always untrusted.trusted_toolsexcludes tools from both untrusted and hostile.actions: action per session rule (sensitive_data_egress,trifecta,after_injection); set one to"log"to switch it off.allow_egress: tool-name glob -> data types (detection rule names such asibanoremail) that tool may send out. Those types don't triggersensitive_data_egress, nortrifectawhen they are the only sensitive data in the session. It also means an undetected injection could direct that tool to send that data type; keep it narrow.sources: tool-name glob -> the label of what that tool returns, e.g.{"get_customer": {"confidentiality": "private"}, "read_issue": {"integrity": "untrusted"}}. Declarations only raise the session label; content detections can raise it further.sinks: tool-name glob -> what that tool accepts:accepts_untrusted = false(untrusted content must not drive it) and/ormax_confidentiality(the most sensitive data it may receive).default_integrity:"declared"(default: a tool's result is trusted only if its capabilities were declared, by you or by an integration for its own tools, or you named it intrusted_tools/sources, and it is local; a name can't establish trust),"trusted"(tools inferred local from their names are trusted too; the behaviour before 0.9), or"untrusted"(every tool result is untrusted unless listed intrusted_tools).trifecta_on_pii: personal data found in tool output counts as sensitive fortrifecta(as secrets always do). Off by default: it makes the session private instead; exact copies leaving are still caught.
declared_consequence ¶
The consequence declared for tool (the most severe matching glob), or None.
declared_trusted ¶
Listed in trusted_tools, or declared integrity = "trusted" in sources (and nowhere untrusted).
allowed_kinds ¶
Data types tool may send out (union over every matching allow_egress pattern).
source_label ¶
The declared label of what tool returns (most restrictive over matching patterns), or None.
sink ¶
(accepts_untrusted, max_confidentiality) for tool; the strictest over matching patterns.
call_cap ¶
The most sensitive data this particular call may carry.
Destinations can allow more for matching values (internal recipients may receive private data); every other value gets the tool's sink cap. The call's cap is the lowest across its values.
declares ¶
Whether the user stated who writes this tool's results: trusted_tools, untrusted_tools, or a source's
integrity. Its capabilities or how confidential its data is don't say that.
guardlayer.GuardSession ¶
A guard bound to one session: every scan reads and updates the session's taint.
set_task ¶
set_task(
name: str,
task_args: Mapping[str, Any] | None = None,
*,
approved: bool = False,
) -> None
Put the session under the task profile name. Call it from trusted code with the user's request.
Setting a first task, or a narrower one, needs nothing. Switching to a task that allows a tool the current one
doesn't (widening) raises PermissionError unless approved=True (a human agreed), and the approval is logged.
narrow ¶
Keep only these of the current task's tools (never adds any). No approval needed.
clear_task ¶
Remove the task restriction. That widens what the agent may do, so it needs approved=True.
Audit and evidence¶
guardlayer.AuditLogger ¶
AuditLogger(
path: str | Path | None = None,
*,
stream: IO[str] | None = None,
min_verdict: Verdict = Verdict.ALLOW,
include_text: bool = False,
use_logging: bool = False,
chain: bool = True,
signer: AuditSigner | str | Path | None = None,
)
guardlayer.AuditSigner ¶
Signs audit entry hashes with an Ed25519 private key.
guardlayer.verify_audit_log ¶
verify_audit_log(
path: str | Path,
*,
public_key: Any | str | Path | bytes | None = None,
expected_head: str | None = None,
) -> AuditVerification
Check a chained audit log: sequence, hash links, entry hashes and (with public_key) signatures.
expected_head, when given, must be the hash of the last entry — this detects truncation.
guardlayer.compliance.build_evidence ¶
build_evidence(
path: str | Path,
*,
public_key: Any | str | Path | bytes | None = None,
expected_head: str | None = None,
frameworks: Iterable[str] | None = None,
) -> EvidencePack
Verify an audit log and map each entry to framework controls.
Always returns a pack; check pack.verification.ok before relying on it (the CLI refuses
unverified logs unless told otherwise). frameworks limits the controls to those frameworks.
guardlayer.compliance.EvidencePack
dataclass
¶
EvidencePack(
source: str,
source_sha256: str,
verification: AuditVerification,
frameworks: list[str],
records: list[dict[str, Any]] = list(),
generated_at: str = (
lambda: (
datetime.now(timezone.utc)
.isoformat()
.replace("+00:00", "Z")
)
)(),
)
control_summary ¶
Per control: how many entries evidence it, by verdict, and the time range covered.
Integrations¶
guardlayer.integrations.tools.guard_tool ¶
guard_tool(
guard: GuardLayer,
fn: F | None = None,
*,
name: str | None = None,
session: Any = None,
approve: Callable[[ScanResult], bool] | None = None,
on_block: str = "message",
withhold_at: Verdict | str = Verdict.BLOCK,
on_injection: str = "withhold",
) -> Any
Wrap fn (sync or async). Usable as guard_tool(guard, fn) or @guard_tool(guard, ...).
session is a session id, a GuardSession, or a zero-argument callable that returns one
(for per-request sessions). name defaults to the function name.
guardlayer.integrations.langgraph.guard_tools ¶
guard_tools(
guard: GuardLayer,
tools: Iterable[Any],
*,
session: Any = None,
on_review: str = "interrupt",
withhold_at: Verdict | str = Verdict.BLOCK,
on_injection: str = "withhold",
) -> list[Any]
Return guarded copies of LangChain tools (anything with .name, .invoke and .ainvoke).
session defaults to the graph's thread_id. Pass a string, a GuardSession or a callable to override it.
guardlayer.integrations.openai_agents.guardrails ¶
guardrails(
guard: GuardLayer,
*,
session: str
| Callable[[Any], str | None]
| None = None,
withhold_at: Verdict | str = Verdict.BLOCK,
on_injection: str = "withhold",
) -> GuardrailSet
Build the four guardrails. Needs pip install openai-agents.
Extending¶
guardlayer.BaseScanner ¶
Convenience base class: holds name/directions and builds Detections.
guardlayer.Rule
dataclass
¶
Rule(
name: str,
pattern: str,
category: str,
severity: float,
message: str,
directions: frozenset[str] = _IN,
ignore_case: bool = True,
multiline: bool = False,
)
guardlayer.LLMJudgeScanner ¶
LLMJudgeScanner(
judge: JudgeFn,
*,
threshold: float = 0.5,
category: str = Category.PROMPT_INJECTION.value,
directions: Iterable[str] | None = None,
)
Bases: BaseScanner
Delegate to an LLM (or any scoring function) you supply.
judge(text, context) returns a probability, or (probability, reason). A typical
implementation calls your model with build_judge_prompt(text) and parses the reply
with parse_judge_score. It is the slowest layer, so enable it selectively.
guardlayer.default_scanners ¶
The zero-dependency default ensemble.