Skip to content

Runbooks that finish what they start.

Turn incident response, patching and routine checks into agent jobs that run on the machines involved. A job survives restarts and lost nodes, waits for people when it must, and shows every step it took.

Live outputCompleted

$ oz0 run incident-responder '{"service": "billing-api", "alert": "p95 latency"}'

incident-responder ⇢ host-inspector (job inc-2291/host-inspector-db2)

host-inspector → system.disk_usage {"path":"/var/lib/postgresql"}

host-inspector ← system.disk_usage 97% used, 4.1 GB free

host-inspector → postgres.replication_slots {}

host-inspector ← postgres.replication_slots analytics: inactive since 03:12

host-inspector: WAL files pile up behind the inactive slot "analytics".

incident-responder: Root cause on db-2: an inactive replication slot keeps WAL files. Drop it once the analytics job is confirmed gone.

JobCost
incident-responder$0.05
host-inspector$0.02
Example output. The tools come from plugins: your own MCP servers or existing packages.

Why it belongs on your own machines.

Work that takes all night
A patch window or an incident can run for hours. Every step is saved, so a crash or a deploy picks up after the last finished step instead of starting over.
On the machine with the problem
The agent reads logs and runs checks on the server itself, not through a jump host and a pile of shared credentials.
Any machine, one fleet
Linux servers, cloud VMs, GPU boxes and the Raspberry Pi in the closet join the same fleet with one command each.

Agents teams build.

Each one is a plugin: an agent definition, the tools it may use and the limits it runs under. Install it from Git and every node that should run it gets it.

  • incident-responderGathers logs, metrics and recent changes from every affected machine and proposes a fix.Runs on Hands work to each hostKept in line by Budget per job
  • patch-runnerPatches machines in waves, checks health after each wave and stops at the first one that fails.Runs on role=serverKept in line by Approval per wave
  • disk-janitorFinds what is filling a disk and cleans up what the policy allows.Runs on The full hostKept in line by Allowed paths only
  • morning-checkReports on backups, certificates and free space across the fleet every morning.Runs on Every nodeKept in line by Runs on a schedule

A night of patching

One job patches the whole fleet in waves, and a server restart in the middle of it costs nothing.

  1. Plan the waves

    The patch-runner groups machines by label and starts with a few canaries.

    Any node

  2. Patch and check

    Each node patches itself, reboots and reports its health. The job waits for it to come back.

    role=server nodes

  3. Survive the night

    If the server running the job restarts, the job continues after the last finished step. No machine is patched twice.

    Temporal

  4. Morning report

    When the last wave is done, the team gets a summary, the machines that need attention and what it cost.

    The job’s history

What makes it work.

  • Durable steps

    Each model and tool call is its own step, so a run continues after the last one that finished.

  • Work handed across machines

    An agent starts child jobs on other nodes and gets their answers back, under one job.

  • Schedules that keep firing

    Recurring checks run on time, even while management is down.

  • People in the loop

    A wave, a restart or a cleanup can wait for an approval, for as long as it takes.

Where it is going

Sites that keep running offline.

Today jobs run on the machines involved and survive restarts. Next, a store, a factory or a ship gets a small edge of its own, so runbooks keep running through a WAN outage and the site catches up with the rest of the fleet when the line returns.

Store 12Store 31Factory 2
Servers
Local edge
Till
Office PC
Local edge
Till
Kiosk
Local edge
Gateway
Line PC
Store 31 has lost its line. Its local edge keeps the runbooks going and catches up when the line returns.
Offline is not out
The local edge keeps the jobs, plugins and certificates its nodes need. Checks, patches and fixes go on while the line is down, and every step syncs back afterwards.
Every machine its own identity
A till, a PLC gateway or a kiosk joins with its own certificate, so you always know which machine ran which step, and can cut one off without touching the rest.
Built for many small sites
Two hundred stores with five machines each is one fleet with two hundred labels, rolled out in waves and watched from one place. The edge is stateless, so it scales out with the sites.

Built for

  • IT and infrastructure teams
  • Site reliability engineers
  • Platform teams
  • Managed service providers