Skip to main content
The Python SDK brings the same enforcement model as the Rust SDK to Python, @tool, @rule, and @agent decorators that govern each call before it executes.

Passthrough mode

When running outside nanny run, every decorator is a no-op. The function executes normally with no enforcement overhead:

@tool: declare a governed tool

Mark a function as a tool that Nanny should govern:
When the agent calls fetch_page:
  1. Nanny checks: is fetch_page in the [tools] allowed list?
  2. Nanny checks: has fetch_page exceeded [tools.fetch_page] max_calls?
  3. Nanny records the call, with the labels the operator declared for it.
  4. If any check fails, a NannyStop exception is raised, the function body never runs.
Works identically for async functions:

Matching the tool allowlist

The tool name used for allowlist checks is the function name as declared in Python:

instrument: automatic LLM token tracking

Call nanny_sdk.instrument(client) once at startup to report LLM token usage automatically. Every completion response is intercepted and its counts recorded. Measurement only. Nothing here stops a run: tokens are recorded so you can answer what a run cost, and Nanny never decides from them.
Supported clients (detected by duck-typing, no provider package is imported):
  • OpenAI, Groq, Together AI, Azure OpenAI, LiteLLM: any client with a chat.completions.create method
  • Anthropic: client.messages.create
  • Mistral: client.chat.complete
  • Google Gemini (google-genai SDK), client.models.generate_content
  • Cohere v2: cohere.ClientV2
instrument returns the same client unchanged, it patches the method in-place. In passthrough mode (outside nanny run), it is a no-op: no wrapping, no overhead. For providers that report prompt-caching usage (OpenAI, Anthropic, DeepSeek, Gemini), instrument also captures it and reports it as cache_read/cache_write, a finer split of input, never additional tokens beyond it, and never used for enforcement (Nanny still debits input + output exactly as before). Every provider names and shapes this data differently, so instrument normalizes each one’s own vocabulary into these two generic fields; absent, not zero, for a provider/response that doesn’t report cache usage at all. Exists purely so a downstream cost calculator can price cache-hit tokens at their real, much cheaper rate, see LlmUsageRecorded for the full field reference.

@rule: declare an enforcement rule

A rule is a function that returns a verdict on whether execution should continue. Return True to allow, False to deny:
Rules are evaluated client-side on every tool call, before Nanny’s enforcement layer is contacted. When a rule returns False, Nanny raises:
The denied tool never runs.

PolicyContext fields

The ctx parameter gives you a snapshot of the current execution state: Rules are evaluated before the tool runs, requested_tool is set to the tool name being checked. Use last_tool_args for content-based enforcement:
All fields are live: the SDK fetches current state before every rule evaluation. tool_labels covers every allowed tool, not just the pending one, because a rule usually needs to ask what an already-called tool was:
ctx.tool_has(tool, label) returns False for an unknown tool and for an unknown label, which is the safe answer in both directions. ctx.tools_with(label) returns every allowed tool carrying a label, sorted. now_ms is supplied rather than read from the clock, so a rule about when an action is permitted stays a pure function of its inputs and can be tested.

@agent: name a phase of the run

In a multi-agent system each phase does different work, and a denial in one is worth attributing to that phase rather than to the process as a whole. @agent names the phase, and every event between entering and exiting belongs to it.
AgentScopeEntered and AgentScopeExited bracket the events in between, so the audit log can answer which phase a denial came from. Nesting is supported. Works identically for async functions. A scope does not change what the agent may do; it labels which phase each verdict belongs to. The scope exits whether the function returns normally or raises.

run_scope: an independent run inside one process

A run is Nanny’s unit of governance: one rule set, one stop state, final once stopped. A long-lived server wants each request to be its own run, so a stop in one never affects another.
Safe under concurrency: the run id lives in a context variable, so threads and async tasks each get their own rather than sharing one process-global value. The scope self-cleans on exit. Only meaningful when governed through a governance server (nanny run --serve or --join), which keys state per run. Under local nanny run one process is always exactly one run, so this is a no-op and code that might run under either mode does not need to branch.

get_run_events: read a run’s events from your own code

Returns the events a governance server has buffered for one run:
You pass the run_id explicitly rather than it being read from the current scope, because the caller is usually a background worker polling several runs and has no run of its own. Every call returns the run’s full buffered list, so track how many you have already consumed, by index or by seq. Returns [] for a run with no events yet and for an unreachable server: this is a side channel, and it should never crash a session over a blip that governed calls already handle by failing closed.

set_app: declare which app this process is

Emits AppIdentified, so runs from several apps sharing one governance server are attributed separately. A process that declares nothing inherits the governor’s identity. nanny run already does this for the process it launches, under both --serve and --join. Call it yourself only for a process nanny run did not start, such as a worker that joins a governor from its own entrypoint.

set_harness: declare what is running this agent

instrument() detects the well-known frameworks from the call stack and the imported modules, and reports what it found alongside each LLM call. Call this when detection cannot help, which is not a rare case: an application whose agent loop is its own code matches no framework, so detection correctly reports nothing and every run arrives unattributed. A first-party agent is not an unknown harness, it simply is not a framework. An explicit declaration wins over detection for the life of the process, since a framework being importable does not mean it drove the call. Deduped bridge-side, so declaring on every request is safe, and a no-op in passthrough mode like the decorators.

What happens on stop

When Nanny stops execution, it raises a NannyStop exception. All stop reasons are distinct subclasses:
Catch them by category or individually:
BridgeUnavailable is raised when Nanny’s enforcement layer is active but unreachable during rule evaluation or a tool call. Nanny fails closed, the agent does not continue ungoverned. Like all NannyStop subclasses it extends BaseException, so it propagates through broad except Exception handlers in agent frameworks without being swallowed. A process that has never reached the bridge retries for 30 seconds first, with backoff, so a container started before the governor it joins waits rather than failing the work it was handed. That retry is armed once and disarms permanently on the first success: after a governor has answered, a later failure raises immediately, because retrying then would let the agent keep calling tools while nothing was authorising them. Under plain nanny run (single process, no --serve), you do not need to handle stop reasons in most agent code. They propagate up the call stack and terminate the process, since nanny run is watching that process and exits it for you. Catching them there is useful mainly in test code and at the CLI entry point. This is not true once you’re running under nanny run --serve / --join. When your process joins a governance server, the wrapping nanny run --join command does not watch your process for stops at all: no polling loop, no kill on a stop, nothing. That’s deliberate. The stop guarantee is that a stop ends one run, not your whole process, so a long-lived server (a web app, a Discord bot, anything handling more than one request per process) has to keep running after another request’s run gets stopped. Nothing does that for you. If a request handler lets NannyStop propagate uncaught, your framework’s own crash behavior decides what happens next, which for many servers means the whole process going down over one governed request. Catch it at the boundary of whatever “one governed unit of work” means for your app, typically each request handler, and turn it into a failure response for that one request, not an unhandled exception:
See Long-lived processes and NannyStop for the full explanation of why this only applies under --serve/--join.

Complete example

Run it under Nanny:
Run it without Nanny (decorators silent, agent runs normally):

Multi-agent pattern

A pipeline where each phase has a role, its own tools, and rules that apply across all of them.
The key properties this gives you:
  • Least privilege: a tool outside [tools] allowed raises ToolDenied before any rule runs
  • Loop detection: the @rule fires before Nanny’s enforcement layer is contacted, so the denied tool never runs
  • Attribution: every verdict is bracketed by the @agent scope that produced it
  • Full audit trail: every tool call and every stop reason logged to NDJSON
For cross-process and cross-machine enforcement, use the governance server.

Framework integration

LangChain

Stack @lc_tool (outer) and @nanny_tool (inner). LangChain registers the function for dispatch; Nanny intercepts every call regardless of which model or API style invoked it:
Execution order: your code calls tool.run(args) → LangChain validates args → Nanny wrapper intercepts → enforcement check → if allowed, file is read.

CrewAI

Same stacking pattern. CrewAI’s @tool decorator and Nanny’s @tool decorator both wrap the function. Nanny’s wrapper fires on every tool.run() call inside the crew:
Rule packs ship the same shape as the rules above. See Rule packs.