Are you an LLM? You can read better optimized documentation at /docs/agent-guard/policies/test-and-publish.md for this page in Markdown format
Test and publish a policy
A policy that does not load is worse than no policy at all. Evaluation fails closed, so a source that fails to parse blocks every agent that pulls it — including agents belonging to people who were not expecting a policy change that morning.
This page covers the loop between "the YAML looks right" and "the team is running it": validating that a policy loads, checking that it decides the way you think, and packaging it for distribution.
Get the CLI
Validation and evaluation run through the policy engine's command-line tool, which is separate from the Agent Guard install. Contact Jozu to obtain it — see Support for the channels.
Agent Guard itself does not need it. Everything on this page is for the person authoring policies, not for the developers running agents against them.
Does it load?
bash
gerty --policy ./policies validateThis parses every YAML file in the directory and compiles the CEL in every rule. It catches the failures that would otherwise surface as a blocked agent: an unknown field, a malformed expression, a regex that does not compile, an assert: false that YAML turned into a boolean before CEL saw it.
Run it before every push. It is the one step that separates a policy bug from an outage.
Does it decide correctly?
Loading is not the same as being right. A rule can compile perfectly and still match nothing, or match everything.
eval runs one request against the loaded policies and prints the decision, which rule fired, and the message returned. It takes a JSON context file as a positional argument:
bash
gerty -p ./policies eval ./case.jsonThe context file is the context the rule will read, plus a kind discriminator telling the engine which policies to run it against:
json
{
"kind": "tool",
"name": "Bash",
"arguments": { "command": "rm -rf /workspace/build" }
}kind is artifact, tool or guardrail. The rest of the object is the matching context — the fields in the CEL context reference, without the leading tool. or artifact. prefix. An artifact case is as small as:
json
{ "kind": "artifact", "registry": "docker.io", "repository": "acme/tool", "tag": "v1" }The decision comes back as JSON, naming the policy and rule that fired:
json
{
"allowed": false,
"action": "Elicit",
"violations": [
{
"policy": "z1-destructive-data-operations",
"rule": "confirm-recursive-delete",
"message": "Recursive delete (rm -r) requires operator confirmation"
}
],
"durationMs": 0,
"elicitationMessage": "This command destroys data irreversibly. Review the exact command before approving -- it will run as written."
}violations is the list, not just the first entry, so a request that trips several rules reports all of them. That is what you check against when a rule is supposed to be the one that fired.
Add --request (-r) with a second JSON file when a rule also reads the request context, for identity or client IP.
Test both directions
Every rule needs two cases: one that must trip it, and a neighbouring one that must not.
Testing only the blocking cases leaves the false positives undetected, and false positives are what cause a policy set to be switched off. A rule that catches rm -rf /workspace is worth nothing if it also catches confirm-rm-behavior.test.ts, because within a week someone will remove it.
Keep the cases in a table next to the policies and run them as a suite, so the whole set is re-checked whenever any rule changes. The reference policy bundle ships with a suite built this way if you want a shape to copy.
Test locally before you distribute
Policies in ~/.agentguard/policies/global/ are picked up on every run, with no publishing step. Edit the YAML, save, run the agent, watch the behavior change. That is the right loop while you are still figuring out an assertion.
bash
agentguard run claude-code --workspace ~/projects/myappPackage and publish
A policy bundle is an OCI artifact, pushed the same way any other artifact is. Validate first — a broken push blocks every agent that pulls it.
bash
gerty --policy . validate
kit pack . -t jozu.ml/myorg/policies:v1
kit push jozu.ml/myorg/policies:v1Always include the registry hostname. A ref without one resolves against the local store, which is what you want while testing and not what you want once published.
Roll it out
Two ways to consume a published bundle.
For a single run, without changing anyone's configuration:
bash
agentguard run claude-code --policy-ref jozu.ml/myorg/policies:v1As a standing source, re-synced before every subsequent run:
bash
agentguard policy add jozu.ml/myorg/policies:v1Use --policy-ref to try a bundle, and policy add once you have decided. See Manage policies for scoping to a workspace, signature verification, and sync behavior.
Audit first
When rolling out a new rule to other people, ship it with action: Audit and review a week of events before promoting it to Enforce. You find out what the rule would have blocked before it blocks anything, measured on real traffic rather than on the cases you imagined.
agentguard logs shows decisions as they happen. See Logs and audit for the audit record.
Versioning
Tags move. A bundle published as :v1 can be republished, and every agent picks up the change on its next run — which is what makes urgent policy changes fast, and what makes accidental ones fast too.
Pin by digest where a run has to see exactly one version:
bash
agentguard policy add jozu.ml/myorg/policies@sha256:abc123...This matters most in CI and shared deployments, where "it worked yesterday" needs to mean something specific.
Rolling back
Publish the previous version under the tag developers are already using, and the next agentguard run syncs it. There is nothing to uninstall on each machine.
If the bad bundle went out under a fixed tag rather than a moving one, developers need agentguard policy remove on the bad ref and policy add on the good one, so prefer moving tags for anything you may need to withdraw quickly.
