Skip to content

Agents on every network you run.

Put a node in each site and its agents reach the devices behind it, with nothing opened inwards. Audit configurations, chase outages and roll out changes, with every step on record.

Live outputCompleted

$ oz0 run outage-triage '{"site": "got-2", "link": "core-1:xe-0/0/3"}'

outage-triage → netbox.device {"name":"core-1"}

outage-triage ← netbox.device mx204 in site got-2, rack B4

outage-triage → junos.show {"device":"core-1","cmd":"interfaces xe-0/0/3 extensive"}

outage-triage ← junos.show input errors 18 204, CRC 18 197, rx power -17.9 dBm

outage-triage → tickets.open {"queue":"NET","severity":"P2"}

outage-triage ← tickets.open NET-4821

outage-triage: CRC errors and low receive power on xe-0/0/3 point to a dirty or failing optic. Details in NET-4821.

JobCost
outage-triage$0.04
Example output. The tools come from plugins: your own MCP servers or existing packages.

Why it belongs on your own machines.

The devices are not on the internet
Management interfaces live on private networks. A node inside each site can reach them, and it only dials out to the edge, so there is no VPN to build and no inbound rule to open.
Changes need a window and a yes
An agent prepares a change on its own. Pushing it waits for the maintenance window and for an approval, and the job holds no machine while it waits.
Many sites, one fleet
Every site is a label. Send a job to one site, to every site in a region, or to all of them at once, from the same command.

Agents teams build.

Each one is a plugin: an agent definition, the tools it may use and the limits it runs under. Install it from Git and every node that should run it gets it.

  • config-auditorCompares running configurations with your source of truth and opens a ticket for every drift.Runs on site=*Kept in line by Read-only tools
  • outage-triagePulls counters, logs and neighbour tables from the devices around a failing link and writes up what it found.Runs on The site with the alarmKept in line by Budget per job
  • change-runnerApplies a reviewed change site by site and checks reachability after each one, stopping at the first failure.Runs on site=*Kept in line by Approval before every push
  • capacity-reportWrites a weekly report of link use and growth per site.Runs on Any nodeKept in line by Runs on a schedule

One change, forty sites, one night

A change request becomes a single job that plans, waits, applies and reports, and that you can stop at any point.

  1. Plan

    The change-runner reads the request and works out the commands for every site.

    Any node

  2. Wait for the window

    The job sleeps until 02:00 and for an approval in the web UI. No machine is busy while it waits.

    Temporal timer

  3. Apply, site by site

    Each site’s own node pushes the change to its devices and checks that everything still answers.

    site=* nodes

  4. Stop or close

    A failed check stops the rollout and names the site. Otherwise the job closes the change with a summary.

    The job’s history

What makes it work.

  • Nodes that only dial out

    A node behind NAT or a strict firewall needs ports 7233 and 7443 out to the edge, and nothing in.

  • A label per site

    Labels such as site=got-2 or region=eu-north decide which node takes a job.

  • Approvals and windows

    Push tools wait for a person to say yes. Timers wait for the window without holding a machine.

  • Your tools, as plugins

    Wrap your device APIs, NetBox or ticketing as MCP tools, or install an existing package.

Where it is going

A node on every router and switch.

Today a node sits in each site and reaches the devices behind it. Next, the node moves onto the devices themselves, one per router, switch and firewall, and they talk to each other directly. A site that loses its uplink keeps finding its own faults, and the full story reaches the edge when the line is back.

Site got-2
Edge
Firewall
Router
Core 1
Core 2
Access 1
Access 2
Wi-Fi
The uplink is down. The devices still talk to each other, find the fault and keep the record until the edge is back.
Working when the network is not
The switch next to a fault tells the router, the router tries the backup line, and the agents on both agree on what broke, without a data centre in the loop.
Every device its own identity
Each node already carries its own certificate from your cluster, renewed every 24 hours. Between devices, both sides prove who they are, and a blocked device is shut out everywhere at once.
Built for scale
A node is one more worker on a task queue and the edge is stateless, so adding devices adds capacity, not complexity. Thousands of switches are routed by label and seen in one place.

Built for

  • Enterprise network teams
  • Internet service providers
  • Managed service providers
  • Campus and data centre operators