Agents on every network you run.
Put a node in each site and its agents reach the devices behind it, with nothing opened inwards. Audit configurations, chase outages and roll out changes, with every step on record.
$ oz0 run outage-triage '{"site": "got-2", "link": "core-1:xe-0/0/3"}'
outage-triage → netbox.device {"name":"core-1"}
outage-triage ← netbox.device mx204 in site got-2, rack B4
outage-triage → junos.show {"device":"core-1","cmd":"interfaces xe-0/0/3 extensive"}
outage-triage ← junos.show input errors 18 204, CRC 18 197, rx power -17.9 dBm
outage-triage → tickets.open {"queue":"NET","severity":"P2"}
outage-triage ← tickets.open NET-4821
outage-triage: CRC errors and low receive power on xe-0/0/3 point to a dirty or failing optic. Details in NET-4821.
| Job | Cost |
|---|---|
| outage-triage | $0.04 |
Why it belongs on your own machines.
- The devices are not on the internet
- Management interfaces live on private networks. A node inside each site can reach them, and it only dials out to the edge, so there is no VPN to build and no inbound rule to open.
- Changes need a window and a yes
- An agent prepares a change on its own. Pushing it waits for the maintenance window and for an approval, and the job holds no machine while it waits.
- Many sites, one fleet
- Every site is a label. Send a job to one site, to every site in a region, or to all of them at once, from the same command.
Agents teams build.
Each one is a plugin: an agent definition, the tools it may use and the limits it runs under. Install it from Git and every node that should run it gets it.
- config-auditorCompares running configurations with your source of truth and opens a ticket for every drift.Runs on site=*Kept in line by Read-only tools
- outage-triagePulls counters, logs and neighbour tables from the devices around a failing link and writes up what it found.Runs on The site with the alarmKept in line by Budget per job
- change-runnerApplies a reviewed change site by site and checks reachability after each one, stopping at the first failure.Runs on site=*Kept in line by Approval before every push
- capacity-reportWrites a weekly report of link use and growth per site.Runs on Any nodeKept in line by Runs on a schedule
One change, forty sites, one night
A change request becomes a single job that plans, waits, applies and reports, and that you can stop at any point.
Plan
The change-runner reads the request and works out the commands for every site.
Any node
Wait for the window
The job sleeps until 02:00 and for an approval in the web UI. No machine is busy while it waits.
Temporal timer
Apply, site by site
Each site’s own node pushes the change to its devices and checks that everything still answers.
site=* nodes
Stop or close
A failed check stops the rollout and names the site. Otherwise the job closes the change with a summary.
The job’s history
What makes it work.
Nodes that only dial out
A node behind NAT or a strict firewall needs ports 7233 and 7443 out to the edge, and nothing in.
A label per site
Labels such as site=got-2 or region=eu-north decide which node takes a job.
Approvals and windows
Push tools wait for a person to say yes. Timers wait for the window without holding a machine.
Your tools, as plugins
Wrap your device APIs, NetBox or ticketing as MCP tools, or install an existing package.
Where it is going
A node on every router and switch.
Today a node sits in each site and reaches the devices behind it. Next, the node moves onto the devices themselves, one per router, switch and firewall, and they talk to each other directly. A site that loses its uplink keeps finding its own faults, and the full story reaches the edge when the line is back.
- Working when the network is not
- The switch next to a fault tells the router, the router tries the backup line, and the agents on both agree on what broke, without a data centre in the loop.
- Every device its own identity
- Each node already carries its own certificate from your cluster, renewed every 24 hours. Between devices, both sides prove who they are, and a blocked device is shut out everywhere at once.
- Built for scale
- A node is one more worker on a task queue and the edge is stateless, so adding devices adds capacity, not complexity. Thousands of switches are routed by label and seen in one place.
Built for
- Enterprise network teams
- Internet service providers
- Managed service providers
- Campus and data centre operators