Skip to content

Other languages

What's in the core (no extra install)

The default rules include "ignore previous instructions" and its variants in several European languages, and, since 0.8, in Hindi, Marathi, Bengali, Gujarati, Punjabi, Tamil, Telugu, Kannada, Malayalam and romanised Hindi, in either word order. Checked on 2,056 normal Wikipedia introductions in those languages (256 chosen because they use the rule's own words): the rule flagged none of them.

The containment layers don't depend on language at all: labels, task profiles, the tool policy, and the behavioural check work whatever language an attack is written in.

Optional: a multilingual classifier on ONNX

For wider coverage you can run a multilingual prompt-injection classifier through the classifier scanner, on ONNX Runtime instead of PyTorch (under 100 MB installed, instead of PyTorch's gigabytes):

pip install "guardlayer[multilingual]"

GuardLayer never downloads a model for you. Fetch the files at a fixed revision and check them, so what runs is what you reviewed. For Horizon Labs' prompt-injection-guard-small (Apache-2.0, 30 languages, 268 MB quantized):

REV=b27472d551844d8c22ed5ad7acd72fe558831bc6
BASE=https://huggingface.co/Horizon-Labs/prompt-injection-guard-small/resolve/$REV
mkdir -p models/pig-small/onnx && cd models/pig-small
curl -fLO $BASE/config.json && curl -fLO $BASE/tokenizer.json
curl -fL -o onnx/model_quantized.onnx $BASE/onnx/model_quantized.onnx
sha256sum tokenizer.json onnx/model_quantized.onnx
# 47834a7dbbb0c4324fe9569fe47955241594a9ee539c9e478755d6fbc51e5f46  tokenizer.json
# 541471ab9de1af3e0db63b8208ec7fdc3ef74202b76c5470115e153c052dae02  onnx/model_quantized.onnx
[scanners.classifier]
model = "models/pig-small"   # the local directory
runtime = "onnx"
threshold = 0.9

Should you turn it on? Our measurement says: not for blocking

We run every candidate model through the same test before recommending it: more attacks caught on data it wasn't trained on, and no new false alarms. This one catches far more attacks, and fails the false-alarm test. Flagged texts at threshold 0.7 (0.9 in brackets), default rules alone vs. default + classifier:

Set Texts Default + classifier
deepset test (attacks) 60 14 32 (24)
jailbreak-classification test (attacks) 139 101 120 (110)
Gandalf test (attacks) 112 64 108 (106)
Public benign prompts 4,509 0 3 (1)
AgentDojo tool outputs (benign) 404 0 11 (8)
LLMail-Inject normal emails 238 0 92 (50)
Indian-language Wikipedia (benign) 2,056 7 8 (7)

Latency on a laptop CPU: 78 ms median, 1.3 s at the 95th percentile (long documents are scanned in several windows).

The model's training data includes LLMail-Inject, so its perfect score on LLMail attacks isn't a fair test and isn't shown. Even so it flags 39% of LLMail's normal emails at 0.7: for an agent that reads email, that is too many.

Our advice: if you want the extra recall, run it in observe mode or as a signal for review, not as a blocker, and measure it on your own traffic first. For Indian languages the core rule already covers the common override phrasing with no false alarms in our tests. The benchmark is in benchmarks/classifier_eval.py; rerun it for any other model you consider.