Skip to content

Python API

Generated from the source docstrings. Everything here is importable from guardlayer unless noted.

The guard

guardlayer.GuardLayer

GuardLayer(
    scanners: Sequence[Scanner] | None = None,
    *,
    policy: Policy | None = None,
    flag_threshold: float | None = None,
    block_threshold: float | None = None,
    canaries: CanaryManager | None = None,
    auto_learn: bool = False,
    tool_allowlist: Iterable[str] | None = None,
    tool_policy: ToolPolicy | None = None,
    session_policy: SessionPolicy | None = None,
    sessions: SessionStore | None = None,
    hooks: Iterable[Callable[[ScanResult], None]] = (),
)

Filter LLM inputs, outputs and third-party context through a configurable scanner ensemble.

scan_input

scan_input(
    prompt: str,
    *,
    session: str | GuardSession | None = None,
    **context_fields: Any,
) -> ScanResult

Scan a user prompt before it reaches the model.

scan_output

scan_output(
    response: str,
    *,
    prompt: str | None = None,
    system_prompt: str | None = None,
    canary: Canary | str | None = None,
    session: str | GuardSession | None = None,
    **context_fields: Any,
) -> ScanResult

Scan a model response before it reaches the user (or a tool).

scan_context

scan_context(
    content: str,
    *,
    source: str | None = None,
    session: str | GuardSession | None = None,
    **context_fields: Any,
) -> ScanResult

Scan third-party content (RAG chunk, web page, email, tool result) for indirect injection.

With a session, the content marks the session as having read untrusted content (and hostile / sensitive content, if found).

scan_tool_call

scan_tool_call(
    tool_name: str,
    arguments: Mapping[str, Any] | str | None = None,
    *,
    session: str | GuardSession | None = None,
    scan_content: bool | None = None,
    **context_fields: Any,
) -> ScanResult

Scan a model-proposed tool call (name + arguments) before executing it.

Runs the tool policy (allow/deny lists, capability actions, argument and egress rules) and, with a session, the taint rules (what the session has already read decides what it may do next). The content scanners also run over the arguments, except for tools tagged read-only (scan_content=None, the default), whose arguments cannot cause harm. A REVIEW verdict means: ask a human first.

scan_tool_result

scan_tool_result(
    tool_name: str,
    result: Any,
    *,
    session: str | GuardSession | None = None,
    arguments: Mapping[str, Any] | str | None = None,
    **context_fields: Any,
) -> ScanResult

Scan what a tool returned before the model reads it (indirect injection channel).

With a session, results from network-capable (or untagged) tools mark the session untrusted, injections mark it hostile, and secrets/PII mark it sensitive. Pass the call's arguments so a file the agent wrote after reading untrusted content is read back as untrusted, whatever tool reads it.

needs_intent_check

needs_intent_check(
    tool_name: str,
    *,
    session: str | GuardSession | None = None,
) -> bool

Is check_intent worth a model call for this tool call?

Only for tools that can act (network, exec, write, or untagged), and, with a session, only once the session has read untrusted content: before that, nothing but the user can be driving the agent.

check_intent

check_intent(
    tool_name: str,
    arguments: Mapping[str, Any] | str | None,
    *,
    messages: Sequence[Mapping[str, Any]],
    replay: Replay,
    session: str | GuardSession | None = None,
    task: str = NEUTRAL_TASK,
    **context_fields: Any,
) -> ScanResult

Behavioural hijack check (see guardlayer.intent): replay the conversation with the user's request hidden.

replay(masked_messages) calls your model and returns the tool calls it proposes. If it proposes this same action anyway, the action is driven by content the agent read: injection_driven_action (review by default). Language- and wording-independent; one extra model call, so use needs_intent_check to pick risky calls.

acheck_intent async

acheck_intent(
    tool_name: str,
    arguments: Mapping[str, Any] | str | None,
    *,
    messages: Sequence[Mapping[str, Any]],
    replay: Callable[..., Any],
    session: str | GuardSession | None = None,
    task: str = NEUTRAL_TASK,
    **context_fields: Any,
) -> ScanResult

check_intent with an async replay.

scan

scan(
    text: str,
    direction: Direction = "input",
    *,
    context: ScanContext | None = None,
    **context_fields: Any,
) -> ScanResult

Scan one text. Extra keyword args populate the ScanContext.

scan_batch

scan_batch(
    texts: Iterable[str],
    direction: Direction = "input",
    **context_fields: Any,
) -> list[ScanResult]

session

session(
    session_id: str | None = None,
    *,
    task: str | None = None,
    task_args: Mapping[str, Any] | None = None,
) -> GuardSession

A view of this guard bound to one session, so tool calls are judged by what came before.

task puts the session under a task profile (see guardlayer.tasks), with task_args from the trusted request.

protect

protect(
    func: F | None = None,
    *,
    system_prompt: str | None = None,
    on_block: str = "raise",
    blocked_message: str = "Sorry, I can't help with that request.",
) -> Any

Decorate fn(prompt, *args, **kwargs) -> str (sync or async) with input and output filtering.

The wrapped function receives the (possibly redacted) prompt, and the caller gets the (possibly redacted) response. On a BLOCK it raises GuardBlocked, or returns blocked_message when on_block="message". Non-string responses pass through unscanned.

add_canary

add_canary(prompt: str, *, echo: bool = False) -> Canary

Embed a canary token in a (system) prompt; pass the returned Canary to scan_output.

add_hook

add_hook(hook: Callable[[ScanResult], None]) -> None

from_preset classmethod

from_preset(
    name: str, overrides: Mapping[str, Any] | None = None
) -> GuardLayer

Build a guard from a named preset (see guardlayer.presets), with optional config overrides.

from_config classmethod

from_config(
    source: str | Path | Mapping[str, Any],
) -> GuardLayer

Build a guard from a TOML/JSON file path or a config dict (see guardlayer.config).

guardlayer.GuardBlocked

GuardBlocked(result: ScanResult)

Bases: Exception

Raised by protect-wrapped calls when a prompt or response is blocked (or held for review).

Results

guardlayer.ScanResult dataclass

ScanResult(
    verdict: Verdict,
    score: float,
    direction: Direction,
    detections: list[Detection] = list(),
    text: str = "",
    modified: bool = False,
    id: str = (lambda: uuid.uuid4().hex)(),
    timestamp: float = time.time(),
    latency_ms: float = 0.0,
    timings_ms: dict[str, float] = dict(),
    errors: list[str] = list(),
    metadata: dict[str, Any] = dict(),
    shadow_verdict: Verdict | None = None,
    observed_rules: list[str] = list(),
)

The aggregated outcome of running the pipeline over one text.

needs_review property

needs_review: bool

True when a human must approve before proceeding.

allowed property

allowed: bool

True when it is safe to proceed without a human: the verdict is ALLOW or FLAG.

effective_verdict property

effective_verdict: Verdict

The stricter of the enforced verdict and the observe-mode shadow verdict.

guardlayer.Detection dataclass

Detection(
    scanner: str,
    rule: str,
    category: str,
    severity: float,
    message: str,
    span: tuple[int, int] | None = None,
    metadata: dict[str, Any] = dict(),
    action: str | None = None,
)

A single signal raised by one scanner.

guardlayer.Verdict

Bases: str, Enum

Overall decision for a scanned piece of text.

guardlayer.Action

Bases: str, Enum

What the pipeline does with a detection of a given category.

guardlayer.Category

Bases: str, Enum

Threat categories. Detections carry the string value, so custom scanners may add their own.

Policies

guardlayer.Policy dataclass

Policy(
    flag_threshold: float = 0.4,
    block_threshold: float = 0.8,
    actions: dict[str, Action] = (
        lambda: dict(DEFAULT_ACTIONS)
    )(),
    fail_closed: bool = False,
    redaction_format: str = "[REDACTED:{rule}]",
    mode: Literal["enforce", "observe"] = "enforce",
    observe: list[str] = list(),
    enforce: list[str] = list(),
)

How detections become a verdict.

actions maps a category — or a "direction:category" pair, which takes precedence — to an Action. Unlisted categories are scored. A detection that carries its own action (tool-policy rules do) uses that instead.

mode = "observe" records everything but enforces nothing (shadow mode). observe lists detections to only observe even in enforce mode; enforce lists detections to keep enforcing in observe mode. Entries are globs matched against the rule name, scanner:rule and the category, e.g. "heuristics:*", "egress_raw_ip", "pii".

is_observed

is_observed(detection: Detection) -> bool

True when this detection is recorded but not enforced.

guardlayer.ToolPolicy

ToolPolicy(
    *,
    allowlist: Iterable[str] | None = None,
    denylist: Iterable[str] = (),
    capabilities: Mapping[str, Iterable[str]] | None = None,
    capability_actions: Mapping[str, Action | str]
    | None = None,
    rules: Iterable[ToolRule | Mapping[str, Any]] = (),
    include_default_rules: bool = True,
    disabled_rules: Iterable[str] = (),
    rule_actions: Mapping[str, Action | str] | None = None,
    egress_allowlist: Iterable[str] | None = None,
    block_exfil_services: bool = True,
    flag_raw_ips: bool = True,
    infer: bool = True,
    remote_tools: Iterable[str] = (),
    include_default_remote_tools: bool = True,
    arguments: Iterable[
        ArgumentRule | Mapping[str, Any]
    ] = (),
)

Evaluate a proposed tool call against allow/deny lists, capability actions, argument rules and egress rules.

is_remote

is_remote(tool: str) -> bool

True when the tool talks to something outside this machine: its results are untrusted content and its arguments leave the machine. Network/exec-capable and untagged tools are remote; so are inferred or untagged tools matching remote_tools (e.g. mcp__*, *search*), even when their names sound read-only. Explicit capabilities win over the default patterns, not over tools you list in remote_tools yourself.

is_declared

is_declared(tool: str) -> bool

Whether the tool's capabilities were declared (config, or an integration's own tools), not guessed.

reads_nothing

reads_nothing(tool: str) -> bool

Declared with no way to read anything (no read, network or exec capability), e.g. a "think" or "finish" tool: its output can only echo the agent, so nobody outside can have written it. Inferred capabilities don't count: an unknown tool may read anything.

vouches

vouches(tool: str) -> bool

Whether an integration vouches for this tool as its own (see vouched).

resolve

resolve(tool: str) -> tuple[frozenset[str], bool]

(capabilities, tagged). An explicit empty list tags a tool as harmless; untagged tools match every rule.

can_act

can_act(tool: str) -> bool

True if the tool may have side effects or reach the network (anything beyond reading).

guardlayer.ToolRule dataclass

ToolRule(
    name: str,
    action: Action | str = Action.BLOCK,
    pattern: str | None = None,
    tools: tuple[str, ...] | None = None,
    capabilities: frozenset[str] | None = None,
    category: str = Category.TOOL_MISUSE.value,
    severity: float = 0.9,
    message: str = "",
    flags: int = re.IGNORECASE,
)

A rule over tool calls.

Fires when the tool matches tools (glob patterns; None = any tool), the tool has one of capabilities (None = any; untagged tools always match), and pattern (a regex over the call's argument text; None = always) is found.

guardlayer.SessionPolicy dataclass

SessionPolicy(
    enabled: bool = True,
    actions: dict[str, Action] = (
        lambda: dict(DEFAULT_SESSION_ACTIONS)
    )(),
    untrusted_tools: list[str] = list(),
    trusted_tools: list[str] = list(),
    hostile_min_verdict: Verdict = Verdict.FLAG,
    allow_egress: dict[str, list[str]] = dict(),
    sources: dict[str, dict[str, Any]] = dict(),
    sinks: dict[str, dict[str, Any]] = dict(),
    default_integrity: str = "declared",
    destinations: list[dict[str, Any]] = list(),
    tasks: dict[str, Any] = dict(),
    trifecta_on_pii: bool = False,
    after_injection_scope: str = "consequence",
    trifecta_scope: str = "destination",
    consequences: dict[str, str] = dict(),
    destination_args: dict[str, list[str]] = dict(),
    untrusted_destination: str = "outbound",
)

What counts as untrusted, and what to do when a tainted session tries to act.

  • untrusted_tools: tool-name globs whose results count as untrusted. By default a tool result is untrusted when the tool can reach the network or is untagged; scan_context input is always untrusted. trusted_tools excludes tools from both untrusted and hostile.
  • actions: action per session rule (sensitive_data_egress, trifecta, after_injection); set one to "log" to switch it off.
  • allow_egress: tool-name glob -> data types (detection rule names such as iban or email) that tool may send out. Those types don't trigger sensitive_data_egress, nor trifecta when they are the only sensitive data in the session. It also means an undetected injection could direct that tool to send that data type; keep it narrow.
  • sources: tool-name glob -> the label of what that tool returns, e.g. {"get_customer": {"confidentiality": "private"}, "read_issue": {"integrity": "untrusted"}}. Declarations only raise the session label; content detections can raise it further.
  • sinks: tool-name glob -> what that tool accepts: accepts_untrusted = false (untrusted content must not drive it) and/or max_confidentiality (the most sensitive data it may receive).
  • default_integrity: "declared" (default: a tool's result is trusted only if its capabilities were declared, by you or by an integration for its own tools, or you named it in trusted_tools / sources, and it is local; a name can't establish trust), "trusted" (tools inferred local from their names are trusted too; the behaviour before 0.9), or "untrusted" (every tool result is untrusted unless listed in trusted_tools).
  • trifecta_on_pii: personal data found in tool output counts as sensitive for trifecta (as secrets always do). Off by default: it makes the session private instead; exact copies leaving are still caught.

declared_consequence

declared_consequence(tool: str | None) -> str | None

The consequence declared for tool (the most severe matching glob), or None.

declared_trusted

declared_trusted(tool: str | None) -> bool

Listed in trusted_tools, or declared integrity = "trusted" in sources (and nowhere untrusted).

allowed_kinds

allowed_kinds(tool: str | None) -> frozenset[str]

Data types tool may send out (union over every matching allow_egress pattern).

source_label

source_label(tool: str | None) -> Label | None

The declared label of what tool returns (most restrictive over matching patterns), or None.

sink

sink(
    tool: str | None,
) -> tuple[bool, Confidentiality | None]

(accepts_untrusted, max_confidentiality) for tool; the strictest over matching patterns.

call_cap

call_cap(
    tool: str | None,
    arguments: Mapping[str, Any] | str | None,
) -> Confidentiality | None

The most sensitive data this particular call may carry.

Destinations can allow more for matching values (internal recipients may receive private data); every other value gets the tool's sink cap. The call's cap is the lowest across its values.

declares

declares(tool: str | None) -> bool

Whether the user stated who writes this tool's results: trusted_tools, untrusted_tools, or a source's integrity. Its capabilities or how confidential its data is don't say that.

guardlayer.GuardSession

GuardSession(
    guard: GuardLayer, session_id: str | None = None
)

A guard bound to one session: every scan reads and updates the session's taint.

set_task

set_task(
    name: str,
    task_args: Mapping[str, Any] | None = None,
    *,
    approved: bool = False,
) -> None

Put the session under the task profile name. Call it from trusted code with the user's request.

Setting a first task, or a narrower one, needs nothing. Switching to a task that allows a tool the current one doesn't (widening) raises PermissionError unless approved=True (a human agreed), and the approval is logged.

narrow

narrow(tools: list[str]) -> None

Keep only these of the current task's tools (never adds any). No approval needed.

clear_task

clear_task(*, approved: bool = False) -> None

Remove the task restriction. That widens what the agent may do, so it needs approved=True.

Audit and evidence

guardlayer.AuditLogger

AuditLogger(
    path: str | Path | None = None,
    *,
    stream: IO[str] | None = None,
    min_verdict: Verdict = Verdict.ALLOW,
    include_text: bool = False,
    use_logging: bool = False,
    chain: bool = True,
    signer: AuditSigner | str | Path | None = None,
)

guardlayer.AuditSigner

AuditSigner(private_key: Any)

Signs audit entry hashes with an Ed25519 private key.

guardlayer.verify_audit_log

verify_audit_log(
    path: str | Path,
    *,
    public_key: Any | str | Path | bytes | None = None,
    expected_head: str | None = None,
) -> AuditVerification

Check a chained audit log: sequence, hash links, entry hashes and (with public_key) signatures.

expected_head, when given, must be the hash of the last entry — this detects truncation.

guardlayer.compliance.build_evidence

build_evidence(
    path: str | Path,
    *,
    public_key: Any | str | Path | bytes | None = None,
    expected_head: str | None = None,
    frameworks: Iterable[str] | None = None,
) -> EvidencePack

Verify an audit log and map each entry to framework controls.

Always returns a pack; check pack.verification.ok before relying on it (the CLI refuses unverified logs unless told otherwise). frameworks limits the controls to those frameworks.

guardlayer.compliance.EvidencePack dataclass

EvidencePack(
    source: str,
    source_sha256: str,
    verification: AuditVerification,
    frameworks: list[str],
    records: list[dict[str, Any]] = list(),
    generated_at: str = (
        lambda: (
            datetime.now(timezone.utc)
            .isoformat()
            .replace("+00:00", "Z")
        )
    )(),
)

control_summary

control_summary() -> list[dict[str, Any]]

Per control: how many entries evidence it, by verdict, and the time range covered.

Integrations

guardlayer.integrations.tools.guard_tool

guard_tool(
    guard: GuardLayer,
    fn: F | None = None,
    *,
    name: str | None = None,
    session: Any = None,
    approve: Callable[[ScanResult], bool] | None = None,
    on_block: str = "message",
    withhold_at: Verdict | str = Verdict.BLOCK,
    on_injection: str = "withhold",
) -> Any

Wrap fn (sync or async). Usable as guard_tool(guard, fn) or @guard_tool(guard, ...).

session is a session id, a GuardSession, or a zero-argument callable that returns one (for per-request sessions). name defaults to the function name.

guardlayer.integrations.langgraph.guard_tools

guard_tools(
    guard: GuardLayer,
    tools: Iterable[Any],
    *,
    session: Any = None,
    on_review: str = "interrupt",
    withhold_at: Verdict | str = Verdict.BLOCK,
    on_injection: str = "withhold",
) -> list[Any]

Return guarded copies of LangChain tools (anything with .name, .invoke and .ainvoke).

session defaults to the graph's thread_id. Pass a string, a GuardSession or a callable to override it.

guardlayer.integrations.openai_agents.guardrails

guardrails(
    guard: GuardLayer,
    *,
    session: str
    | Callable[[Any], str | None]
    | None = None,
    withhold_at: Verdict | str = Verdict.BLOCK,
    on_injection: str = "withhold",
) -> GuardrailSet

Build the four guardrails. Needs pip install openai-agents.

Extending

guardlayer.BaseScanner

BaseScanner(directions: Iterable[str] | None = None)

Convenience base class: holds name/directions and builds Detections.

guardlayer.Rule dataclass

Rule(
    name: str,
    pattern: str,
    category: str,
    severity: float,
    message: str,
    directions: frozenset[str] = _IN,
    ignore_case: bool = True,
    multiline: bool = False,
)

guardlayer.LLMJudgeScanner

LLMJudgeScanner(
    judge: JudgeFn,
    *,
    threshold: float = 0.5,
    category: str = Category.PROMPT_INJECTION.value,
    directions: Iterable[str] | None = None,
)

Bases: BaseScanner

Delegate to an LLM (or any scoring function) you supply.

judge(text, context) returns a probability, or (probability, reason). A typical implementation calls your model with build_judge_prompt(text) and parses the reply with parse_judge_score. It is the slowest layer, so enable it selectively.

guardlayer.default_scanners

default_scanners(
    canaries: CanaryManager | None = None,
) -> list[Scanner]

The zero-dependency default ensemble.