Installation
The SDK ships inside the same crate as the CLI binary. Add it to your project:Passthrough mode
If your agent runs withoutnanny run, every macro is a no-op. The function executes normally with no enforcement overhead:
#[tool]: declare a governed tool
Mark a function as a tool that Nanny should govern:
fetch_page:
- Nanny checks: is
fetch_pagein the[tools] allowedlist? - Nanny checks: has
fetch_pageexceeded[tools.fetch_page] max_calls? - Nanny records the call, with the labels the operator declared for it.
- If any check fails, execution stops immediately, the function body never runs.
Matching the tool allowlist
The tool name used for allowlist checks is the function name as declared in Rust:nanny::http_get: built-in HTTP tool
nanny::http_get is a built-in governed HTTP GET function. It requires no #[tool] annotation. Nanny applies allowlist, call-count, and token enforcement automatically.
- Checked against the
[tools] allowedlist (tool name:"http_get") - Subject to
[tools.http_get] max_callslimit
nanny run), nanny::http_get makes the request directly with no enforcement overhead.
nanny::report_usage: report LLM token usage
To record the tokens an LLM used, report them after the call with nanny::report_usage. Hand Nanny the counts already present on the response.
Measurement only. Nothing here stops a run: tokens are recorded so you can answer what a run cost, and Nanny never makes a decision from them.
input and output are required. You can optionally attach model and provider labels, identifiers only, never prompt or response content, and, if your provider’s response reports it, cache_read/cache_write (a finer split of input, never additional tokens beyond it):
LlmUsageRecorded event. report_usage is fire-and-forget: it never blocks your agent and never panics. In passthrough mode (no nanny run) it is a no-op.
Rust reports usage explicitly, one call per LLM response. In Python,
nanny_sdk.instrument(client) does the same thing automatically by wrapping the client, Rust cannot patch a client at runtime, so the reporting is an explicit call instead.#[rule]: declare an enforcement rule
A rule is a function that returns a verdict on whether execution should continue.
Return true to allow, false to deny:
false, Nanny stops execution with:
PolicyContext fields
Thectx parameter gives you a snapshot of the current execution state:
Rules are evaluated before the tool runs, so
requested_tool is the call
being checked, not one already made.
tool_labels covers every allowed tool because a rule usually needs to ask what
an already-called tool was. “Did anything that reads untrusted content run
before this?” cannot be answered from the pending call alone.
now_ms is supplied rather than read from the clock, so a rule about when an
action is permitted stays a pure function of its inputs and can be tested.
Two helpers
tool_has returns false for an unknown tool and for an unknown label, which
is the safe answer in both directions: a rule asking about a tool the operator
never declared should not fire, and a misspelled label must not match
everything.
#[agent]: name a phase of the run
Mark a function so every verdict produced inside it is attributed to that phase:
AgentScopeEntered and AgentScopeExited bracket the events in between, so the
audit log can answer which phase a denial came from. Nesting is supported.
nanny::run_scope: an independent run inside one process
A run is Nanny’s unit of governance: one rule set, one stop state, final
once stopped. A long-lived host that handles many independent requests wants
each one to be its own run, so a stop in one never affects another.
nanny run one process is always exactly one run, so this is a no-op there.
Mirrors nanny_sdk.run_scope() on the Python side.
nanny::set_app: declare which app this process is
AppIdentified, so runs from several apps sharing one governance server
are attributed separately rather than collapsing into one stream. A process
that declares nothing inherits the governor’s identity.
nanny run already does this for the process it launches, reading the
committed .nanny/app.json, under both --serve and --join. Call it
yourself only for a process nanny run did not start, such as one that joins a
governor from its own entrypoint. Deduped bridge-side, so calling it on every
request is safe. Mirrors nanny_sdk.set_app() on the Python side.
nanny::set_harness: declare what is running this agent
version is optional:
unknown, which is worth setting even
for a first-party engine: an application whose agent loop is its own code is
not an unknown harness, it simply is not a framework.
Deduped bridge-side. Mirrors nanny_sdk.set_harness() on the Python side.
Complete example
What happens on stop
When Nanny stops execution inside an instrumented function, the macro propagates the stop signal by panicking with a structured message. The panic is caught by the Nanny runtime, your agent process exits cleanly with a non-zero code and anExecutionStopped event in the log.
You do not need to handle stop reasons in your agent code. Nanny handles the exit path.