reza

·5 min read

Conductor: making multi-agent Claude Code runs inspectable

A small Claude Code plugin for multi-agent runs: who did what with which tools, live progress, claims checked before they reach me, and one place where I am asked to decide.

Most of my longer work with Claude Code splits one task across many subagents. Past a handful of agents I lose track of the basics, and results like "the tests pass" come back up the chain with no provenance, or worse, without anyone having run the tests. Conductor is a small plugin that keeps track for me:

  • Who did what. Every agent, the task it was given, who dispatched it, and what it reported.
  • With which tools. Every file read, command run, page fetched and MCP call, per agent.
  • Where things stand. A live view of the plan, with active and finished agents.
  • What to trust. Every claim is checked by a separate agent, against sources the agents actually touched, before it reaches me.
  • When I am needed. Agents that need a decision stop and ask, through the one session I talk to, instead of guessing.

Conductor

The code is on GitHub.

What it does#

  • One orchestrator. The session I talk to plans the work, dispatches agents and asks me when a decision is mine. It can read anything but cannot edit the project, so every change goes through an agent that gets checked.
  • Agents that can nest. Workers make changes, analysts only read, and both can delegate to sub-agents.
  • A check before anything travels up. A worker states its claims with evidence. A separate auditor agent, starting from a fresh context, checks each claim against the source and marks it verified, refuted or unverified. Only then may the worker report.
  • Ways to keep the cost down. Low-stakes tasks, like listing files, can skip the review; their claims are then shown as asserted, never as verified. The auditor can also run on a cheaper model than the workers.
  • A record of everything. Every dispatch, file read, command, claim and verdict goes into one append-only log. A live dashboard and a provenance query are built from that log.

Conductor architecture: work flows down, only verified claims flow up

How it works#

The rules are not instructions in a prompt. They are enforced by Claude Code hooks: small scripts that Claude Code runs at fixed moments, such as before a tool is used or when a subagent tries to finish. A hook can refuse an action or send the agent back with a reason.

If an agent tries tothe hook
dispatch work without a task id and acceptance criteriarefuses, and shows the format
finish while its own sub-agents are still runningsends it back to wait
report "done" without an independent review of its claimssends it back to get one
cite a file it never read, a command it never ran, or an MCP call it never madesends it back: "C1 cites cmd:… but you never ran it"
keep a claim the auditor refuted, or reword one after reviewsends it back until it is fixed or withdrawn

The fourth row is the one I care about most. The hooks log every file an agent reads, every command it runs and every MCP call it makes (a database query, for example), so a citation can be checked against what actually happened. After three failed attempts the agent is let through, but its result is labelled unverified and shows up as an item for me to review.

Inside, each Claude Code event starts one short Python process. It answers with a decision, appends a line to the log, and triggers a rebuild of the dashboard. Nothing else writes the log.

Conductor internals: only the hook writes the log; every view replays it

Last words of caution#

  • The auditor can be wrong. The gate guarantees that a check happened, against sources the agents actually touched. It does not guarantee that the check was correct; the auditor is also a model.
  • The evidence check is mechanical. It confirms a command was run, not that the command shows what the claim says.
  • It may break with Claude Code updates. It relies on details of Claude Code's hook payloads that could change between releases.