Are you an LLM? You can read better optimized documentation at /docs/agent-guard/policies/writing-rules.md for this page in Markdown format
Writing rules
Agent Guard's policy engine fails closed. An assertion that throws denies the request, and it denies with an engine error rather than your message. That single property shapes how rules are written: every expression has to be safe against data that is not there.
This page covers the patterns that keep rules from misfiring, and the ones that keep them from firing on the wrong thing.
Guard everything you touch
Arguments
Tools do not share an argument schema. Bash sends command, Read sends file_path, Grep sends pattern. Reading an argument the tool did not send is an evaluation error.
An unguarded rule written for Bash does not quietly skip a Read call — it blocks it, and the agent is told an expression failed rather than why the operation was refused. Open every rule with a membership check:
yaml
assert: |
!("command" in tool.arguments) ||
!tool.arguments.command.contains("rm -rf")The || short-circuits: when the argument is absent the rule passes without evaluating the rest. To cover several arguments, guard each one:
yaml
assert: |
(!("command" in tool.arguments) || !tool.arguments.command.contains("rm -rf")) &&
(!("file_path" in tool.arguments) || !tool.arguments.file_path.contains("/etc/"))Optional context fields
Four context fields can be absent entirely: artifact.kitfile, artifact.layers, tool.session, and request.user. Guard them with has():
yaml
assert: |
!has(artifact.kitfile) ||
artifact.kitfile.package.license in ["Apache-2.0", "MIT", "BSD-3-Clause"]The decision is the same either way — a missing Kitfile fails the policy — but only one version tells the operator which rule refused and why. The other says no such key: kitfile.
Scanner scores
This is the most damaging unguarded access available, because it does not fail one request, it fails all of them.
yaml
# Wrong. With no scanner attached, the key is absent, CEL raises,
# and fail-closed evaluation blocks every model request in the deployment.
assert: |
guardrail.scores["toxicity"] < 0.7Written with a membership guard, the rule lies dormant until a scanner supplies the score, then takes effect on its own:
yaml
assert: |
!("toxicity" in guardrail.scores) ||
guardrail.scores["toxicity"] < 0.7This lets you ship threshold rules before wiring a scanner, which is usually the right order: get the policy reviewed and deployed, then turn on the signal.
Two ways to write a rule
The convention is that assert describes the safe case and the action fires when it is false. For a condition on the request itself, that reads naturally:
yaml
spec:
match:
tool:
names: ["Bash"]
action: Enforce
rules:
- name: no-force-push
assert: |
!("command" in tool.arguments) ||
!tool.arguments.command.contains("push --force")
message: "Force push is prohibited"For a blocklist, inverting every entry gets unreadable fast. Put the condition in match instead, and let the rule assert the literal "false" so everything the match selects fires the action:
yaml
spec:
match:
expression: |
!(artifact.registry in ["jozu.ml", "registry.internal.example.com"])
action: Enforce
rules:
- name: registry-not-on-allowlist
assert: "false"
message: "Registry is not on the organization's allowlist. Mirror the artifact into an approved registry."Note the quotes. assert: false is a YAML boolean and will not parse as CEL; assert: "false" is a CEL expression that always fails. The same trick gives you an audit-everything policy: match broadly, assert "false", action Audit.
Use the match-based form when the question is "which requests does this apply to," and the rule-based form when the question is "what about this request is unsafe."
Regular expressions
Patterns are RE2. There are no backreferences and no lookahead, so some expressions cannot be written as one pattern and have to be split across rules or restructured.
Prefer character classes to backslash escapes. [.]env and \.env mean the same thing to RE2, but [.] passes through YAML quoting unchanged while \. depends on how the string is quoted.
.contains() is substring matching, not regex: .contains("curl") matches concurly. When you need a word boundary, use .matches():
yaml
assert: |
!("command" in tool.arguments) ||
!tool.arguments.command.matches("(^|[^\\w-])curl([\\s;|&]|$)")A literal regex that does not compile fails the policy at load time rather than at evaluation time, so gerty validate catches it before it can block anything.
Avoiding false positives
The usual failure mode for a content rule is not a missed match. It is a rule that fires on v0.0.0-20260519132957-10bb9b174f44 and gets switched off a week later. A policy nobody trusts is worse than no policy.
Two pattern shapes avoid most of it.
Structure-anchored, where the format is distinctive on its own. A Social Security number's 3-2-4 grouping, an ITIN's 9xx-7x-xxxx shape, or a Medicare beneficiary identifier, whose alphabet omits S, L, O, I, B and Z precisely so it cannot be confused with other identifiers.
Context-anchored, where the value alone is ambiguous and a nearby keyword is required. Ten bare digits are far more often a timestamp than a government ID, so the rule looks for a label near the number rather than the number alone.
Some things are better left unmatched. Bare email addresses appear constantly in commit metadata and package manifests, so matching them produces noise rather than protection. A mailing-address rule needs a postal code within a short distance of the match, or "allocate 512 MB Ram Drive for cache" reads as a street address.
Test both directions. A rule needs a case that must trip it and a neighbouring case that must not — see Test and publish a policy. A rule set that only proves it blocks things is half-tested, and the untested half is what determines whether anyone keeps it switched on.
Tightening safely
When you add or narrow a rule, ship it as Audit first and review a week of events before promoting it to Enforce. That tells you what the rule would have blocked before it starts blocking anything, on your traffic rather than on the examples you thought of while writing it.
The same applies to thresholds. Attach the scanner, watch the scores it produces on real requests, then set the number.
