Runbooks that finish what they start.
Turn incident response, patching and routine checks into agent jobs that run on the machines involved. A job survives restarts and lost nodes, waits for people when it must, and shows every step it took.
$ oz0 run incident-responder '{"service": "billing-api", "alert": "p95 latency"}'
incident-responder ⇢ host-inspector (job inc-2291/host-inspector-db2)
host-inspector → system.disk_usage {"path":"/var/lib/postgresql"}
host-inspector ← system.disk_usage 97% used, 4.1 GB free
host-inspector → postgres.replication_slots {}
host-inspector ← postgres.replication_slots analytics: inactive since 03:12
host-inspector: WAL files pile up behind the inactive slot "analytics".
incident-responder: Root cause on db-2: an inactive replication slot keeps WAL files. Drop it once the analytics job is confirmed gone.
| Job | Cost |
|---|---|
| incident-responder | $0.05 |
| host-inspector | $0.02 |
Why it belongs on your own machines.
- Work that takes all night
- A patch window or an incident can run for hours. Every step is saved, so a crash or a deploy picks up after the last finished step instead of starting over.
- On the machine with the problem
- The agent reads logs and runs checks on the server itself, not through a jump host and a pile of shared credentials.
- Any machine, one fleet
- Linux servers, cloud VMs, GPU boxes and the Raspberry Pi in the closet join the same fleet with one command each.
Agents teams build.
Each one is a plugin: an agent definition, the tools it may use and the limits it runs under. Install it from Git and every node that should run it gets it.
- incident-responderGathers logs, metrics and recent changes from every affected machine and proposes a fix.Runs on Hands work to each hostKept in line by Budget per job
- patch-runnerPatches machines in waves, checks health after each wave and stops at the first one that fails.Runs on role=serverKept in line by Approval per wave
- disk-janitorFinds what is filling a disk and cleans up what the policy allows.Runs on The full hostKept in line by Allowed paths only
- morning-checkReports on backups, certificates and free space across the fleet every morning.Runs on Every nodeKept in line by Runs on a schedule
A night of patching
One job patches the whole fleet in waves, and a server restart in the middle of it costs nothing.
Plan the waves
The patch-runner groups machines by label and starts with a few canaries.
Any node
Patch and check
Each node patches itself, reboots and reports its health. The job waits for it to come back.
role=server nodes
Survive the night
If the server running the job restarts, the job continues after the last finished step. No machine is patched twice.
Temporal
Morning report
When the last wave is done, the team gets a summary, the machines that need attention and what it cost.
The job’s history
What makes it work.
Durable steps
Each model and tool call is its own step, so a run continues after the last one that finished.
Work handed across machines
An agent starts child jobs on other nodes and gets their answers back, under one job.
Schedules that keep firing
Recurring checks run on time, even while management is down.
People in the loop
A wave, a restart or a cleanup can wait for an approval, for as long as it takes.
Where it is going
Sites that keep running offline.
Today jobs run on the machines involved and survive restarts. Next, a store, a factory or a ship gets a small edge of its own, so runbooks keep running through a WAN outage and the site catches up with the rest of the fleet when the line returns.
- Offline is not out
- The local edge keeps the jobs, plugins and certificates its nodes need. Checks, patches and fixes go on while the line is down, and every step syncs back afterwards.
- Every machine its own identity
- A till, a PLC gateway or a kiosk joins with its own certificate, so you always know which machine ran which step, and can cut one off without touching the rest.
- Built for many small sites
- Two hundred stores with five machines each is one fleet with two hundred labels, rolled out in waves and watched from one place. The edge is stateless, so it scales out with the sites.
Built for
- IT and infrastructure teams
- Site reliability engineers
- Platform teams
- Managed service providers