On this page
Stado — command gap analysis
Based on the 2026-07-24 control-host incident (full disk → wedged launchd → dead weles-api, no reboot path, no service adoption path).
Host lifecycle
- host reboot TARGET — graceful reboot through the approved channel.
(Shipped as
stado host reboot. The Rust portdeploy/host_reboot.rswas complete but UNREACHABLE:deploy/mod.rsnever declared the module and no CLI variant dispatched to it, so the command recorded here as implemented did not exist. Both halves are wired now.) - host uptime TARGET — uptime, load averages, logged-in users.
(Shipped as
stado host uptime [--json],deploy/host_uptime.rs. Reads the load averages from the kernel rather than scraping theuptimeline, whose shape differs between macOS and Linux.) - host ping TARGET — reachability: ssh check + beacon age in one verdict.
(Shipped as
stado host ping [--json],deploy/host_ping.rs. Reports both signals and takes the WORSE as the verdict and the exit status, so the box in this incident — answering ssh with a five-day-old beacon — fails it. Staleness isconfig::HEARTBEAT_STALE_MINUTES, the crate's existing tolerance for a one-minute writer, and the age is rendered by the samecli::registry::human_agethatregistry beacon-ageuses.) - host disk TARGET — current disk usage plus the registry cleanup policy
state (last pass, freed bytes, next scheduled pass).
(Shipped as
stado host disk [--json],deploy/host_disk/mod.rs.df -Pk /for usage, the registry's ownDiskCleanupPolicyfor the declaration, and the janitor's own state file — located viaproviders::local::disk_cleanup::state_relative_path— for last pass, freed bytes and next scheduled pass. No second schema.) - host cleanup TARGET --dry-run — preview what the registry cleanup would
delete, without deleting.
(Shipped as
stado host cleanup TARGET --dry-run [--json],deploy/host_cleanup.rs, which contains NO cleanup policy. It runs the host's own stado — found throughhost_recovery::WC_CANDIDATES— asdisk-cleanup --once --dry-run, i.e.providers::local::disk_cleanup::preview_cleanup_once: the janitor's own planning phase with anenforcepolicy pinned down to its ownreportmode and no state written.--dry-runis mandatory; the enforcing pass stays with the janitor's interval andhost recover.) - host exec TARGET -- CMD — run a fixed read-only command on a host
through the approved channel (with an allowlist, not free shell).
(Shipped as
stado host exec TARGET [--json] -- CMD…,deploy/host_exec/allowlist/mod.rs. Three barriers: shell-metacharacter rejection, an exact match againstAPPROVED_COMMANDS, and execution of the matched entry's own fixed absolute argv — the operator's words select an entry, they never become part of the command line. Every entry carries awhyjustifying it as read-only and argument-free.)
All six ride one channel, deploy/host_channel.rs, which derives its ssh
option set from host_reboot::ssh_reboot_argv (it calls it and drops the
trailing program) rather than copying it, so the commands cannot drift apart.
Two more joined them after a 2026-08 diagnosis that the six could not finish. The first reads what a host actually has; the second is the only write in the group, and carries out what the registry declares:
6a. host inventory TARGET — the stado-managed binaries under
$HOME/.stado/bin, the $HOME/.stado/forwards/*.url markers, the
listening loopback ports, and whether the installed stado knows a fixed
list of subcommands.
(Shipped as stado host inventory TARGET [--json],
deploy/host_inventory/report/collect.rs. None of those facts were reachable: they all
need $HOME, and item six's contract is a fixed argv of absolute paths
with NO operator-supplied path in it, so extending that allowlist was the
wrong fix — reading them instead meant a raw
ssh user@ip '<inline script>' with a hardcoded address, plus pgrep -fl
and printenv, the two things the allowlist deliberately withholds
because argv and environments carry secrets. The command takes a registry
target name and nothing else — no path, file name, port or pattern — and
its remote program is one compile-time script with no interpolation in
it. Its real output is the reconciliation: for each forward marker,
whether anything is listening on the port the marker names. It was
written because on control-host the stado-weles-api.url marker
said http://127.0.0.1:8766 while the admission API was on 8794, and
nothing in the fleet read the markers. Listener ownership comes from
netstat -anv -p tcp — already justified in item six's table — as bare
pids; no lsof, no pgrep -f, no argv, no environment.)
6b. host release TARGET --binary NAME --version X.Y.Z — put one
registry-declared managed binary on a host.
(Shipped as
stado host release TARGET --binary NAME --version X.Y.Z [--platform P] [--dry-run] [--json], deploy/host_release.rs. This is the gap
ARCHITECTURE.md names outright: "no system in the pack currently owns
'get this build onto that host'. That is a gap". Item 6a reads what a
host HAS and targets[].managed_versions declares what it SHOULD have;
nothing carried out the difference, so closing it meant copying a binary
by hand. The order is Weles's shipped auto-deploy order
(weles/scripts/worker/deploy/README.md) applied to one binary: fetch
the exact coordinate through /api/release/object, verify the
OPERATOR-configured SHA-256 — not one computed from the download, and not
a manifest fetched from the same endpoint as the artifact, which is where
this is deliberately stricter than bootstrap — check the layout, stage
under $HOME/.stado/releases/<binary>/<version>/<platform>/, and only
then hard-link and rename(2) the artifact over
$HOME/.stado/bin/<binary> and restart the registry-declared unit
through the same program service restart uses. The three remote phases
are three separate compile-time programs on the shared channel rather
than one script, so "nothing activates before it verified" is a property
of which programs were sent and is observable at the Runner seam; a
failed fetch or a mismatched digest leaves the running version untouched.
--binary and --platform are closed compile-time tables and --version
is an exact semantic version, so the operator's words select entries and
never become path or URI segments — item six's rule, kept. It refuses
when the registry declares no version for that host and binary, or
declares a different one: delivery carries out a declaration, it does not
stand in for one. Idempotent by reporting rather than by repeating —
an already-current host answers already_active and receives no further
program.)
Service management (the "full service management" layer)
- service list — every registry-managed service across all hosts with state (active/inactive/failed/missing) from the latest beacons.
- service status NAME — one service's state everywhere it is managed.
- service restart NAME [--host TARGET] — restart one managed service without the full host-recovery pass.
- service adopt UNIT --host TARGET — adopt an existing LaunchAgent into the managed set (the weles-api gap: it exists on the host but stado does not manage it).
- service retire UNIT --host TARGET — remove a service from management (bootout + forget, files kept).
- service deploy NAME --host TARGET --from PATH — install a new service unit under management (plist + bootstrap + registry note).
- service logs NAME [--host TARGET] [--lines N] — tail a unit's log without ssh-ing by hand.
- service env NAME — show the effective environment a managed service runs with (from its plist), secrets redacted.
All eight are shipped: stado service list|status|restart|adopt|retire| deploy|logs|env, engine in deploy/service/mod.rs, surface in
cli/service/mod.rs. list and status answer from the health beacons alone,
so the fleet-wide question costs no ssh. The managed set has two sources,
shown in a SOURCE column: the per-target services array in the registry
(what adopt/retire/deploy edit) and the fixed MANAGED_AGENTS list
every host recover pass reloads — genuinely managed, so listed, but owned
by that program, which is why retire refuses them.
missing (the beacon exists and does not carry the unit) is kept distinct
from unknown (no beacon at all): conflating a silent host with a vanished
unit is precisely the failure this group exists to stop. Mutations go
through cli::registry::push_document, which validates before it writes, so
an edit that would produce an invalid registry uploads nothing. env
redacts secret-shaped keys before the report is built, so no rendering path
can print a value.
Registry truth
- registry doctor — diff registry declarations against live host state: unmanaged agents, missing plists, stale beacons, hosts with no heartbeat.
- registry host add HOST --ssh DEST --release-platform PLATFORM — onboard
a new machine into the canonical registry (validated). Both flags are
required;
stado fleet enrollis the path that probes the machine for them. - registry beacon-age — one table: every host and its last beacon timestamp (the "hasn't reported in 5 days" detector).
All three shipped. registry doctor reports no-heartbeat, stale-beacon,
missing-plist, unit-not-active, unmanaged-host and unmanaged-agent, sourced
from the beacon prefix and the capacity broadcasts — never ssh — and exits
non-zero when declaration and reality disagree. registry host add reuses
push_document's validation rather than a second implementation.
registry beacon-age gives every registry target a row, including targets
with no beacon at all, worst first.
Related, and the reason these were reachable at all: registry push/pull
were pinned to a hardcoded GCS bucket, so on an Azure-only deployment the
registry the coordinator's survival check depends on could not be repaired.
All registry readers and writers now go through targets::RegistryStore,
which keeps the GCS path byte-identical and routes every other backend
through the configured store.
Jobs / queue
- job rerun ID — resubmit a finished/failed job with identical spec.
(Shipped as
stado job reruninstado-rs/src/cli/job/rerun/mod.rs. Replays the spec throughqueue::submit::submit_batchrather than hand-writing a job document, so routing, the run manifest and the listing metadata are stamped by the same code a freshstado submituses.) - job watch ID — stream a running job's log (machine logs exists, but
is cursor-based; watch wraps it into a tail).
(Shipped as
stado job watch [--follow]in the same module. Carries the byte cursor forward across polls, tails to a terminal prefix and exits with the job's outcome.) - queue pause / queue resume — maintenance mode: stop/start dispatching
without cancelling queued jobs.
(Shipped as
stado queue pause|resume|status|drain, state inqueue/control.rs. Pausing gates three paths, not one: scheduler dispatch, the local agent's claim loop, and box admission — the third was found while implementing, and without itdrain --waitcould watchrunning/grow while waiting. Already-running jobs finish untouched; the agent's cooperative-yield eviction is also gated, since evicting a running job to free room for a claim that can never happen would destroy work.drain --waitblocks untilrunning/empties and exits non-zero on timeout. This is the supported pre-migration drain thatdeploy/MIGRATE_TO_STADO.mdpreviously enforced with an honour-system environment variable.)
Blockers found while implementing
The earlier port's deploy/host_recovery.rs file was removed when recovery
was split into modules. The current module entry is
stado-rs/src/deploy/host_recovery/mod.rs. Use the declared
repair capability and operations
for supported recovery commands rather than reconstructing incident scripts.
Second gap set — the billing-outage incident
The GCP billing account was closed and every GCS call began returning
accountDisabled. Six independent defects turned that into a total outage,
and each of them surfaced as a crash loop or a silently empty UI rather than
as a check. The commands below close what that revealed. All are shipped.
- doctor —
stado doctor [--json] [--fix-hints], probes indoctor.rs. An ordered preflight over config, storage auth plus a real write/read/delete round trip, provider auth, live quota, the release channel, agent-template rendering, Azure VM identity, registry reachability, queue pause state and alert channels. Every probe is fault-isolated and deadline-bounded, so one black-holed endpoint is one FAIL row rather than a hung command. The template check renders through the dispatcher's own code path, which is the only way it can prove anything about what a real dispatch would produce. - storage ls | stat | cat | verify — the outage question was "is the queue
empty, or is the store unreachable?", and nothing could answer it, because
BlobBackend::existsmaps every error to false.stattherefore probes through the path that surfaces the error and reportsunreachabledistinctly fromabsent.verifyis the object-for-object comparison the removeddeploy/MIGRATE_TO_STADO.mddemanded and never provided. - storage copy — there was no way to move queue state between backends at all. It carries blob metadata, not just bodies: the scheduler prefilters on those stamps, so a body-only copy leaves jobs visible while degrading every tick into downloading the whole queue.
- instances list | reap — orphaned cloud VMs bill forever and were
invisible. Implementing it exposed a live bug: the Azure provider's
enumeration existed but was never wired to the
Providertrait, so the base default applied and Azure agent VMs were invisible to the dead-agent reaper as well as to any CLI. - cancel --terminate — cancelling a job left its VM running and billing.
Plain
cancelis unchanged; the instance reference is resolved from the job document first and the provider lease second. - secrets put | get | ls | rm — the Azure billing service principal lived
in GCP Secret Manager, so GCP dying also blinded us to the Azure credit
balance. Values now come from the separate Skarbiec service;
putreads STDIN only because argv is visible in process listings and shell history. - billing watch — the alert that should have warned us could not fire:
the depletion signal is computed only inside the
"status": "ok"branch, so a closed account or dead credential produced silence rather than an alarm. There is now an account-health signal independent of the balance threshold, alerting on transition and de-duplicated through the billing blob, plus a foreground watchdog that is deliberately runnable OUTSIDE the cloud it watches — a collector that dies with its provider cannot warn you about that provider.
Known dead code, needs a decision
queue/secrets.rs and monitor::billing::{fetch_azure_sp_with, no_credentials_section} are now #[cfg(test)]-only: production reads the
billing service principal from Skarbiec and nowhere else, deliberately, so no
fallback can quietly recreate the cross-cloud coupling that caused the outage.
What remains is a whole module plus a status-message builder kept
alive by a single test asserting text production can no longer emit. Removing
them means deleting that test, which needs the owner's approval — hence this
note instead of a commit.
Third gap set — the Skarbiec key-loss incident
add-user --role owner registered a freshly generated key as owner without
re-encrypting anything and without changing the vault's owner field, and the
key every stored item was actually encrypted to left the keyring the same
night. Every surface reported success: the broker's /health answered ok
without touching key material, reads of an undecryptable item dropped the TCP
connection instead of returning a status, and recovery-status printed a
fingerprint whether or not the offline material still existed. Diagnosis then
took hours of one-off shell pipelines — list the keyring, map fingerprints to
keygrips, guess which recipient the ciphertext names, hunt for the file that
would open it. None of that was a command, so none of it survived the session
that produced it, which is the defect this set closes.
- secrets doctor —
stado secrets doctor [--json], surface incli/secrets/diagnostics/doctor.rs, engine in Skarbiec's ownkey-doctor. It runs the installed binary rather than reimplementing the check: the vault and the keyring belong to Skarbiec, and during an outage a second opinion that disagrees with the program actually performing the decryption is worse than no opinion. One table gives every recipient, its role, whether the vault document really names it owner, whether its secret half is in this keyring, and the exactprivate-keys-v1.d/<KEYGRIP>.keypath a restore has to produce — the encryption subkey, because that is the file decryption needs and the primary will not do. A readable vault exits zero; anything else exits non-zero carrying the remedy. It is answered before any Skarbiec client is constructed, since a grant, a token or a live service is precisely what may be broken, andSKARBIEC_BINoverrides discovery — the only way to diagnose a build before it is installed, which is the case whenever the installed binary is itself what is stale.
Shipped in Skarbiec, where the knowledge belongs: key-doctor (the engine
above, reading the vault document and keyring directly and never the HTTP API,
proving the verdict by opening a deterministic canary item and discarding the
plaintext), rotate-owner (which re-encrypts every current and historical
ciphertext onto a new owner and preserves the recovery recipient — the
operation add-user was mistaken for, and which did not exist in the shipped
source), an honest recovery-status, a /health that opens key material
instead of reporting liveness, and a refusal on add-user --role owner.
Two findings came out of running the new command rather than reasoning about
it, which is the argument for shipping commands over notes: the vault document
still named the previous owner because add-user writes the recipients map and
never the owner field, and two worker recipients hold items whose secret halves
may live on other machines — recovery avenues nobody had listed.
secrets harvest —
stado secrets harvest [--json] [--all] [--restore NAME], surface incli/secrets/diagnostics/harvest.rs, engine intranscripts/queries/inventory.rs. Agent runtimes persist every tool call and result, and those results include process listings, environment dumps and file reads, so credentials the fleet never meant to write down sit in plain text on disk, dated, in files nobody prunes. During this incident that was the only surviving copy of several live values. The scan has two uses at once: the recovery inventory for a vault whose key material is gone, and the exposure inventory to shrink afterwards.It reports names, counts, distinct-value counts, dates and file counts, and NEVER a value — the defect being measured is values reaching places that only needed names, and a tool that printed them would add a terminal, a shell history and its own transcript to that list. A value moves only through
--restore NAME, which streams it intoskarbiec setover stdin.The transcript lake under
~/.transcript-lakeis deliberately not a source: its ingest masks high-entropy fields, so armored key material arrives there already destroyed. The raw per-session stores keep it intact, which is the finding worth acting on.--restorerefuses whenkey-doctorsays the vault cannot be opened. Encrypting needs only public halves, so writing into a dead vault SUCCEEDS and produces one more unreadable item — the recovered value would be buried in the hole it was pulled out of.The engine reads the session schemas rather than scanning the files as text, and that distinction is the whole difference between an inventory and a pile of guesses. Both stores record which tool produced each payload: omp names the tool on the result event, Claude names the call and carries the tool name in the earlier assistant
tool_useblock, so the file is read in order and the id-to-name map moves forward with it. A payload from a shell or an evaluator observed the live machine — environments, process tables, command output. A payload from a read or a search merely quoted a file, so itsKEY-ish names are identifiers in the repository, not credentials that were in use. Only the first kind is scanned by default and only the first kind can be restored;--allwidens to the second. A flat text scan cannot tell them apart, which is exactly how the first version of this reported hundreds of variable names as recoverable secrets.
Source: this website