Skip to content

Built-in rules

Generated from the code (guardlayer rules prints the same list). Severities feed the score; tool rules carry their own action. Switch rules off with disabled_rules, change a tool rule's action with rule_actions, and add your own with a rule pack.

Content rules

Rule Category Severity Directions What it catches
ignore_previous_instructions prompt_injection 0.90 context, input Attempt to override prior instructions.
forget_everything prompt_injection 0.85 context, input Tells the model to discard everything that came before.
change_instructions prompt_injection 0.80 context, input Attempts to rewrite the model's instructions.
new_instructions prompt_injection 0.60 context, input Injects a replacement instruction block.
new_instructions_follow prompt_injection 0.60 context, input Announces a replacement set of instructions.
from_now_on prompt_injection 0.45 context, input Attempts to redefine the assistant's future behaviour.
instruction_priority_claim prompt_injection 0.70 context, input Claims precedence over existing instructions.
fake_special_tokens prompt_injection 0.80 context, input, output Chat-template control tokens embedded in text.
fake_role_header prompt_injection 0.50 context, input Forged system/developer role header.
ai_directed_instruction prompt_injection 0.70 context Text addresses an AI reader directly (indirect injection).
ai_must_instruction prompt_injection 0.65 context Embedded directive aimed at an AI (indirect injection).
addressed_to_ai prompt_injection 0.60 context Data speaks to the AI reading it (indirect injection).
task_precondition_instruction prompt_injection 0.75 context Data inserts a task the AI must do first (indirect injection).
fake_system_marker prompt_injection 0.70 context Forged system-message marker inside content.
forged_chat_turn prompt_injection 0.70 context Forged conversation turn or role tag inside content (indirect injection).
forged_safety_verdict prompt_injection 0.80 context Content forges a safety-check or judge verdict to get past defences.
summary_anchored_action prompt_injection 0.60 context Instruction chains an extra action onto the agent's current task.
split_letter_obfuscation prompt_injection 0.50 context Letters separated to evade keyword filters.
agent_goal_statement prompt_injection 0.60 context Content states a goal for the AI agent reading it.
hidden_html_instruction prompt_injection 0.60 context Instruction hidden in an HTML comment.
hidden_css_text prompt_injection 0.30 context Visually hidden text in markup.
encoded_instruction prompt_injection 0.70 context, input Asks the model to decode and then act on a payload.
payload_splitting prompt_injection 0.50 context, input Payload-splitting: reassemble fragments then act on them.
tool_abuse_instruction prompt_injection 0.60 context, input Directs an agent to misuse a tool.
ignore_instructions_multilingual prompt_injection 0.90 context, input Instruction override in a non-English language.
ignore_instructions_indic prompt_injection 0.90 context, input Instruction override in an Indian language (Hindi, Marathi, Bengali, Gujarati, Punjabi, Tamil, Telugu, Kannada, Malayalam, romanised Hindi).
new_instructions_multilingual prompt_injection 0.60 context, input Announces replacement instructions (non-English).
reveal_prompt_multilingual system_prompt_leak 0.75 context, input Asks the model for its own instructions (non-English).
unethical_ai_persona jailbreak 0.75 context, input Defines an AI persona without ethics or restrictions.
has_no_rules jailbreak 0.50 context, input Describes the model as having no rules or limits.
jailbreak_marker jailbreak 0.80 context, input Known jailbreak control marker.
jailbreak_persona jailbreak 0.90 context, input Persona/role override (jailbreak persona).
do_anything_now jailbreak 0.85 context, input Known jailbreak persona reference.
dan_reference jailbreak 0.60 context, input References the DAN jailbreak family.
privileged_mode jailbreak 0.70 context, input Request to enter a privileged/unrestricted mode.
disable_safety jailbreak 0.80 context, input Attempt to disable safety controls.
hypothetical_framing jailbreak 0.70 context, input Hypothetical framing used to elicit restricted output.
refusal_suppression jailbreak 0.55 context, input Refusal suppression.
no_rules_claim jailbreak 0.60 context, input Claims the model is no longer bound by its rules.
opposite_mode jailbreak 0.45 context, input Alternate-persona jailbreak pattern.
reveal_system_prompt system_prompt_leak 0.85 context, input Attempt to extract the system prompt.
ask_for_own_instructions system_prompt_leak 0.60 context, input Asks the model for its own instructions.
repeat_text_above system_prompt_leak 0.60 context, input Asks the model to regurgitate its context.
prompt_beginning system_prompt_leak 0.70 context, input Asks what came at the start of the prompt.
system_prompt_disclosure system_prompt_leak 0.60 output Response appears to disclose the system prompt.
exfiltration_to_destination data_exfiltration 0.90 context, input, output Instruction to exfiltrate sensitive data to a destination.
exfiltration_request data_exfiltration 0.70 context, input Request to leak sensitive data.
jailbreak_success_marker jailbreak 0.75 output Response signals a jailbroken persona.
destructive_delete unsafe_command 0.85 context, output Recursive force-delete of a root/home path.
remote_script_pipe unsafe_command 0.70 context, output Pipes a remote script straight into an interpreter.
reverse_shell unsafe_command 0.90 context, output Reverse-shell pattern.
encoded_powershell unsafe_command 0.85 context, output Obfuscated/remote PowerShell execution.
fork_bomb unsafe_command 0.90 context, output Fork bomb.
disk_wipe unsafe_command 0.85 context, output Disk format/wipe command.
destructive_sql unsafe_command 0.35 context, output Destructive SQL statement.
world_writable_root unsafe_command 0.60 context, output Makes the filesystem root world-writable.

Tool-call rules

Rule Category Default action Applies to What it catches
destructive_command tool_misuse block exec Destructive command (recursive delete of a root/home directory, disk wipe or fork bomb).
risky_command tool_misuse review exec Risky command (history rewrite, data deletion, privilege escalation, publishing or shutdown).
persistence tool_misuse review exec, write Persistence mechanism (startup files, cron, services or scheduled tasks).
credential_file tool_misuse block any tool Access to a credential store (SSH keys, cloud credentials, password stores or system secrets).
remote_code_execution tool_misuse review exec Downloads code and runs it straight away.
credential_access tool_misuse review exec Retrieves a stored credential, which puts the secret in the agent's hands.
dotenv_file tool_misuse review any tool Access to a .env file, which usually holds secrets.

Session rules

Rule Default action
sensitive_data_egress block
trifecta review
after_injection review
untrusted_destination review
untrusted_to_protected_sink review
confidentiality_exceeds_sink review
untrusted_file_executed review
out_of_task review
task_argument_not_allowed review
injection_driven_action review
intent_check_failed log

See Sessions and taint for when they fire.