@tool, @rule, and @agent decorators that govern each call before it executes.
Passthrough mode
When running outsidenanny run, every decorator is a no-op. The function executes normally with no enforcement overhead:
@tool: declare a governed tool
Mark a function as a tool that Nanny should govern:
fetch_page:
- Nanny checks: is
fetch_pagein the[tools] allowedlist? - Nanny checks: has
fetch_pageexceeded[tools.fetch_page] max_calls? - Nanny records the call, with the labels the operator declared for it.
- If any check fails, a
NannyStopexception is raised, the function body never runs.
Matching the tool allowlist
The tool name used for allowlist checks is the function name as declared in Python:instrument: automatic LLM token tracking
Call nanny_sdk.instrument(client) once at startup to report LLM token usage automatically. Every completion response is intercepted and its counts recorded.
Measurement only. Nothing here stops a run: tokens are recorded so you can answer what a run cost, and Nanny never decides from them.
- OpenAI, Groq, Together AI, Azure OpenAI, LiteLLM: any client with a
chat.completions.createmethod - Anthropic:
client.messages.create - Mistral:
client.chat.complete - Google Gemini (
google-genaiSDK),client.models.generate_content - Cohere v2:
cohere.ClientV2
instrument returns the same client unchanged, it patches the method in-place. In passthrough mode (outside nanny run), it is a no-op: no wrapping, no overhead.
For providers that report prompt-caching usage (OpenAI, Anthropic, DeepSeek, Gemini), instrument also captures it and reports it as cache_read/cache_write, a finer split of input, never additional tokens beyond it, and never used for enforcement (Nanny still debits input + output exactly as before). Every provider names and shapes this data differently, so instrument normalizes each one’s own vocabulary into these two generic fields; absent, not zero, for a provider/response that doesn’t report cache usage at all. Exists purely so a downstream cost calculator can price cache-hit tokens at their real, much cheaper rate, see LlmUsageRecorded for the full field reference.
@rule: declare an enforcement rule
A rule is a function that returns a verdict on whether execution should continue.
Return True to allow, False to deny:
False, Nanny raises:
PolicyContext fields
Thectx parameter gives you a snapshot of the current execution state:
Rules are evaluated before the tool runs,
requested_tool is set to the tool name being checked. Use last_tool_args for content-based enforcement:
tool_labels covers every allowed tool, not just the pending one, because a
rule usually needs to ask what an already-called tool was:
ctx.tool_has(tool, label) returns False for an unknown tool and for an
unknown label, which is the safe answer in both directions.
ctx.tools_with(label) returns every allowed tool carrying a label, sorted.
now_ms is supplied rather than read from the clock, so a rule about when an
action is permitted stays a pure function of its inputs and can be tested.
@agent: name a phase of the run
In a multi-agent system each phase does different work, and a denial in one is worth attributing to that phase rather than to the process as a whole. @agent names the phase, and every event between entering and exiting belongs to it.
AgentScopeEntered and AgentScopeExited bracket the events in between, so the
audit log can answer which phase a denial came from. Nesting is supported.
Works identically for async functions. A scope does not change what the agent may do; it labels which phase each verdict belongs to. The scope exits whether the function returns normally or raises.
run_scope: an independent run inside one process
A run is Nanny’s unit of governance: one rule set, one stop state, final
once stopped. A long-lived server wants each request to be its own run, so a
stop in one never affects another.
nanny run --serve
or --join), which keys state per run. Under local nanny run one process is
always exactly one run, so this is a no-op and code that might run under either
mode does not need to branch.
get_run_events: read a run’s events from your own code
Returns the events a governance server has buffered for one run:
run_id explicitly rather than it being read from the current
scope, because the caller is usually a background worker polling several runs
and has no run of its own.
Every call returns the run’s full buffered list, so track how many you have
already consumed, by index or by seq. Returns [] for a run with no events
yet and for an unreachable server: this is a side channel, and it should never
crash a session over a blip that governed calls already handle by failing
closed.
set_app: declare which app this process is
AppIdentified, so runs from several apps sharing one governance server
are attributed separately. A process that declares nothing inherits the
governor’s identity.
nanny run already does this for the process it launches, under both
--serve and --join. Call it yourself only for a process nanny run did not
start, such as a worker that joins a governor from its own entrypoint.
set_harness: declare what is running this agent
instrument() detects the well-known frameworks from the call stack and the
imported modules, and reports what it found alongside each LLM call. Call this
when detection cannot help, which is not a rare case: an application whose
agent loop is its own code matches no framework, so detection correctly reports
nothing and every run arrives unattributed. A first-party agent is not an
unknown harness, it simply is not a framework.
An explicit declaration wins over detection for the life of the process, since
a framework being importable does not mean it drove the call. Deduped
bridge-side, so declaring on every request is safe, and a no-op in passthrough
mode like the decorators.
What happens on stop
When Nanny stops execution, it raises aNannyStop exception. All stop reasons are distinct subclasses:
BridgeUnavailable is raised when Nanny’s enforcement layer is active but unreachable during rule evaluation or a tool call. Nanny fails closed, the agent does not continue ungoverned. Like all NannyStop subclasses it extends BaseException, so it propagates through broad except Exception handlers in agent frameworks without being swallowed.
A process that has never reached the bridge retries for 30 seconds first,
with backoff, so a container started before the governor it joins waits rather
than failing the work it was handed. That retry is armed once and disarms
permanently on the first success: after a governor has answered, a later
failure raises immediately, because retrying then would let the agent keep
calling tools while nothing was authorising them.
Under plain nanny run (single process, no --serve), you do not need to handle stop reasons in most agent code. They propagate up the call stack and terminate the process, since nanny run is watching that process and exits it for you. Catching them there is useful mainly in test code and at the CLI entry point.
This is not true once you’re running under nanny run --serve / --join. When your process joins a governance server, the wrapping nanny run --join command does not watch your process for stops at all: no polling loop, no kill on a stop, nothing. That’s deliberate. The stop guarantee is that a stop ends one run, not your whole process, so a long-lived server (a web app, a Discord bot, anything handling more than one request per process) has to keep running after another request’s run gets stopped. Nothing does that for you. If a request handler lets NannyStop propagate uncaught, your framework’s own crash behavior decides what happens next, which for many servers means the whole process going down over one governed request. Catch it at the boundary of whatever “one governed unit of work” means for your app, typically each request handler, and turn it into a failure response for that one request, not an unhandled exception:
--serve/--join.
Complete example
Multi-agent pattern
A pipeline where each phase has a role, its own tools, and rules that apply across all of them.- Least privilege: a tool outside
[tools] allowedraisesToolDeniedbefore any rule runs - Loop detection: the
@rulefires before Nanny’s enforcement layer is contacted, so the denied tool never runs - Attribution: every verdict is bracketed by the
@agentscope that produced it - Full audit trail: every tool call and every stop reason logged to NDJSON
Framework integration
LangChain
Stack@lc_tool (outer) and @nanny_tool (inner). LangChain registers the function for dispatch; Nanny intercepts every call regardless of which model or API style invoked it:
tool.run(args) → LangChain validates args → Nanny wrapper intercepts → enforcement check → if allowed, file is read.
CrewAI
Same stacking pattern. CrewAI’s@tool decorator and Nanny’s @tool decorator both wrap the function. Nanny’s wrapper fires on every tool.run() call inside the crew: