dottie-os — local multi-agent floor

The hive

One directory of plain files runs a whole team of agents: a roster, a shared blackboard, a task ledger, mailboxes, and append-only logs. No database, no watch-train theatrics — files you can read with cat. The System One sidecar decides; the hive keeps the floor on disk.

01Why this exists

Pair-programming still needs a roster, a blackboard, and a ledger. Those stay on your machine as plain files — next to the local sidecar, not as a public train dashboard or a chat control plane.

Every job still lands in one place. Choice / Score / Noul decide the next move. The hive records who did what. That is the floor. It is not a public train dashboard.

02What the hive is

A directory on your machine. The design was audited from the Munder Difflin project and rebuilt for Dottie's stdlib-local world — the reusable core turned out to be the plain-file coordination, not the Electron app around it.

hive/
  registry.json    # roster: every agent, role, capabilities, status
  board.md         # shared blackboard — one designated scribe writes it
  tasks.json       # task ledger: id, assignee, spec, status, result
  log.jsonl        # append-only event feed (drives the activity stream)
  costs.jsonl      # append-only cost ledger
  agents/<id>/
    identity.md    # who I am, my role, my capabilities
    memory.md      # long-term memory — read at start, appended as I learn
    inbox/         # messages delivered to me, one JSON file each
    outbox/        # messages I want to send — the router drains these
    cursor.json    # where I am in my inbox, so nothing is read twice

Three rules keep it from falling apart. One process commits to git — agents never touch it, so concurrent writers can't corrupt the repo. Each agent writes only inside its own directory. And messages move by a router that picks files out of one agent's outbox/ and drops them into another agent's inbox/, written atomically — no file is ever written by two processes at once.

03How messages move — and what happens when they can't

An agent that needs something writes a message file to its outbox. The router — polling about every 1.5 seconds — delivers it to the recipient's inbox and appends the event to the log. When an agent finishes a turn, it drains its inbox before stopping, so work keeps flowing without anyone watching.

Delivery is fail-closed. Nothing is silently dropped:

  • A message that can't be delivered bounces to the orchestrator, which decides what to do with it.
  • A malformed message is quarantined and logged — never delivered, never deleted quietly.
  • Conversations carry a hop count. Past the cap (12), the orchestrator steps in instead of letting two agents loop forever.
  • Only requests, queries, and proposals obligate a reply. Plain notifications are terminal — they can't start a ping-pong.

04Four autonomy gates

Agents run on their own, but four layers decide how far that goes:

  1. Orchestrator triage

    A standing orchestrator reads cross-agent traffic and resolves the routine stuff itself — clarifications, data requests, plan tweaks. Only the critical items escalate. The escalation policy lives in its prompt, so you tune the prompt, not the code.

  2. Tool permissions are the human gate

    Agents ask permission to use tools the way a single assistant would. Those prompts are the human-in-the-loop checkpoint — and they work from a phone, not just a desk.

  3. Pause / gate / steer / halt at the boundary

    A control registry enforces four states on any agent at the boundary, regardless of what the agent wants to do. Policy is enforced where actions cross into the world, not inside the agent's head.

  4. Destructive ops need explicit confirm words

    Deleting, spending, or changing scope requires a deliberate confirmation — a bare “yeah” or “ok” never authorizes anything, mass targets are forbidden, and a confirmation expires after two minutes.

05Budgets and the breaker

Every agent runs on a work-token budget — counted on real work, not cache re-reads. When an agent misbehaves, a circuit breaker walks a ladder: steer (nudge it back on track), then constrain (narrow what it may do), then stop. One level per tick, and it de-escalates on recovery. It never kills an agent unless a human opts in.

Spawning new agents is off by default, and any spawn request runs isolated unless you say otherwise.

06What's in, what's out

In

  • File-based hive layout and message contract
  • Fail-closed routing and append-only ledgers
  • The breaker ladder and work-token budgets
  • Confirm tiers for destructive operations
  • Boundary-enforced pause / gate / steer / halt
  • Isolate-by-default spawn queue

Out — deliberately

  • Electron app and the 2D office theatrics
  • Voice control
  • Provider-shim sprawl
  • Auto-approve mode
  • Heavyweight vector memory at this scale — markdown first

07One open question

Where should human approvals surface?

The gates above need a human at the other end. That surface is configurable — the dashboard, a chat app, the terminal — and it's your call. Nothing here hard-codes one channel.

08The shape of it

Lean stdlib core. Local-first — the hive lives on your machine, in files you own. The website, the local runtime, and the orchestrator harness are one project, built in that order. This page is step one.

Design audited from the Munder Difflin project (read-only) and rebuilt for Dottie's stdlib-local world. Spec numbers above — the 1.5s router poll, the 12-hop cap, the 2-minute confirm expiry — come from that audit.