Skip to content

Configuration

You may not need a file at all: a preset alone is a working setup. When you do, one TOML or JSON file configures everything. Load it with GuardLayer.from_config("guardlayer.toml"), --config on the CLI, or GUARDLAYER_CONFIG for the REST API. A preset sets the baseline; everything in the file overrides it.

Describe your tools (the part that matters most)

GuardLayer's strongest protection comes from knowing your tools: which ones bring in outside content, which ones can act or send data out. Say it once per tool:

preset = "balanced"

[tool.read_email]
capabilities = ["read"]            # read | write | network | exec
output = "untrusted"               # others can write what it returns
output_data = "private"            # public | private | restricted

[tool.read_docs]
output = "trusted"                 # your own content: never untrusted

[tool.send_email]
capabilities = ["network"]
accepts_untrusted = false          # untrusted content must not drive it
max_data = "private"               # the most sensitive data it may receive
may_send = ["email"]               # data it is meant to send (exempt from egress rules)
arguments = [{ argument = "to", allow = ["*@mycompany.com"], action = "review" }]

[tool."mcp__github__*"]            # names can be globs
output = "untrusted"
Key Meaning
capabilities what the tool can do: read, write, network, exec (inferred from the name when omitted)
output "trusted" or "untrusted": can someone else write what it returns?
output_data how sensitive its output is: public, private (business data), restricted (secrets)
accepts_untrusted false: refuse (review) when the session has read untrusted content
max_data the most sensitive data it may receive
may_send data types it is meant to send out, e.g. ["iban"] for a payment tool
remote true: its arguments leave the machine though it only reads (a search API)
arguments rules for argument values: allow / deny globs, action
destinations allow more data for some values: { argument, match, max_data }

Don't know where to start? Record a few days in observe mode with an audit log (min_verdict = "allow"), then draft the declarations from what the agent actually used, and correct the draft:

guardlayer --config guardlayer.toml policy draft audit.jsonl --claude-code -o tools.toml

Check what you declared with guardlayer --config guardlayer.toml policy check: it lists every tool and what GuardLayer assumes about it.

Everything else

The sections below are the full reference. [tool.NAME] is a shorter way to write the per-tool parts of [tools], [session] and [labels]; both forms work together.

preset = "balanced"            # observe | balanced | strict | airgap

[guard]
flag_threshold = 0.4
block_threshold = 0.8
fail_closed = false            # block when a scanner errors (strict and airgap set this)
auto_learn = true              # add blocked prompts to the similarity corpus
mode = "enforce"               # or "observe": record what would happen, enforce nothing
observe = ["egress_raw_ip"]    # observe only these rules or categories (globs), even in enforce mode
enforce = []                   # keep enforcing these in observe mode

[actions]                      # category, or "direction:category" -> score | block | flag | review | redact | log
secret = "redact"
"output:pii" = "redact"
policy = "block"

[tools]                        # agent tool-call policy
allowlist = ["search", "bash", "mcp__github__*"]
denylist = ["delete_repo"]
egress_allowlist = ["api.github.com"]   # a domain covers its subdomains
remote_tools = ["kb_*"]                 # added to the defaults (mcp__*, *search*, *web*, ...)
capabilities = { run_sql = ["write"], lookup = ["read"], TodoWrite = [] }
capability_actions = { exec = "review" }
rule_actions = { egress_raw_ip = "block" }
disabled_rules = []
rules = [{ name = "no_prod", pattern = "prod-db", action = "block" }]

[session]                      # taint tracking
actions = { trifecta = "review", after_injection = "review", sensitive_data_egress = "block" }
trusted_tools = ["read_docs"]           # results never count as untrusted or hostile
untrusted_tools = ["read_email"]        # results always count as untrusted ("*" for every tool)
allow_egress = { send_money = ["iban"] } # data types a tool may send out (exempt from egress and trifecta)
store = "memory"                        # or "file", with dir = "...", for checks in separate processes
ttl_seconds = 86400

[labels]                                # information flow (see Concepts: Labels)
default_integrity = "declared"          # results trusted only from local tools you or an integration vouched for
                                        # (output = "trusted"); "trusted": also tools whose names sound local (before
                                        # 0.9); "untrusted": none unless in trusted_tools
sources = { get_customer = { confidentiality = "private" }, read_issue = { integrity = "untrusted" } }
sinks = { post_comment = { max_confidentiality = "public" }, write_file = { accepts_untrusted = false } }
destinations = [{ tool = "send_email", argument = "to", match = "*@mycompany.com", max_confidentiality = "private" }]

[[tools.arguments]]                     # allowed values for one argument (globs; lists split)
tool = "send_email"
argument = "to"
allow = ["*@mycompany.com"]
action = "review"

[audit]                        # tamper-evident audit log
path = "guardlayer-audit.jsonl"         # "{hostname}" and "{pid}" are filled in
min_verdict = "flag"
signing_key = "audit.key"               # optional Ed25519 key (`signing` extra)
include_text = false                    # store scanned text, not only its hash (avoid)

[scanners.heuristics]
disabled_rules = ["fake_role_header"]
rules_file = "my_rules.toml"

[scanners.similarity]
threshold = 0.55
corpus_file = "attacks.txt"
max_windows = 256                       # windows compared per text; lower is faster on very long texts

[scanners.pii]
entities = ["email", "credit_card", "aadhaar"]

[scanners.links]
allowed_domains = ["example.com"]

[scanners.denylist]            # opt-in: on when the section is present
terms = ["project nightingale"]

[scanners.classifier]          # opt-in, `ml` extra
threshold = 0.7
# device = 0                   # GPU index
# model = "org/model"          # a different Hugging Face classifier
# revision = "<commit sha>"    # pin it (the default model is pinned for you)
# runtime = "onnx"             # run a local ONNX model dir instead (`multilingual` extra, no PyTorch)
# model_file = "onnx/model_quantized.onnx"

Every scanner section accepts enabled and directions. The default scanners are on unless disabled; the opt-in ones (denylist, classifier, relevance) switch on when their section is present.

Environment variables

Variable Overrides
GUARDLAYER_PRESET preset
GUARDLAYER_MODE [guard] mode
GUARDLAYER_FLAG_THRESHOLD, GUARDLAYER_BLOCK_THRESHOLD the thresholds
GUARDLAYER_FAIL_CLOSED [guard] fail_closed
GUARDLAYER_AUTO_LEARN [guard] auto_learn
GUARDLAYER_CONFIG config path for the REST API
GUARDLAYER_API_KEY required X-API-Key for the REST API
GUARDLAYER_STATE_DIR session directory for the Claude Code hook

In code

Everything in the file has a Python equivalent:

from guardlayer import GuardLayer, Policy, SessionPolicy, ToolPolicy

guard = GuardLayer(
    policy=Policy(block_threshold=0.8, fail_closed=True),
    tool_policy=ToolPolicy(egress_allowlist=["api.github.com"], capability_actions={"exec": "review"}),
    session_policy=SessionPolicy(untrusted_tools=["read_email"]),
)

Or build from a dict: GuardLayer.from_config({"preset": "strict", "tools": {"egress_allowlist": ["api.github.com"]}}).