vps-health-api/agent
le king fu e181a9691c fix(review): close the three gaps found in stack review
Follow-up to the /pr-review pass on PRs #18-#22. Three findings, none of
which changed behaviour, all of which weakened a guarantee the stack was
supposed to provide.

1. auth.test.js carried "any new route must be added here" but /hosts was
   never added when #13 introduced it, so no test asserted GET /hosts -> 401
   on a missing header. The drift class the file exists to catch slipped on
   its first outing. ROUTES now covers /hosts, and a dedicated block pins
   both directions of the token separation: a read token cannot write, an
   ingest token cannot read.

2. ingestOversized() folded any transport error into 413, so the four
   oversized-body tests would have stayed green if the server had stopped
   writing the status and merely killed the socket - on the one path where
   "a status, not a dead socket" is the whole client contract. Transport
   errors are now surfaced instead of swallowed.

3. HOST_AGENT_TIMEOUT_MS was read by push-metrics.js but never exported by
   run-push.sh, so setting it in the documented env file did nothing. Now
   exported and documented.

Also drops a claim from agent/README.md that the review proved false: the
ingest/read token split buys no containment on the ThinkPad, which already
stores HEALTH_TOKEN in cleartext for defenseur-auto. Losing that laptop
compromises both, so they rotate together.
2026-08-16 14:24:43 -04:00
..
push-metrics.js feat(agent): push workstation metrics to /hosts/<id> from cron 2026-08-16 12:21:59 -04:00
README.md fix(review): close the three gaps found in stack review 2026-08-16 14:24:43 -04:00
run-push.sh fix(review): close the three gaps found in stack review 2026-08-16 14:24:43 -04:00

Workstation agent

Collects one snapshot of the machine it runs on and pushes it to POST /hosts/<id> on the health API. Cron runs it every 5 minutes; the dashboard reads the result through GET /hosts.

Zero dependencies — node:http / node:https and the metrics.js already shared with the server. Node 22+.

File Role
push-metrics.js Collect, build the payload, POST once, exit
run-push.sh Cron wrapper: load the env file, check the install, exec node

This directory is not part of the server image: the Dockerfile copies package.json index.js metrics.js one by one, and a test pins that it stays that way. The agent ships by copying files onto a workstation, never by deploy.

What it deliberately does not do

No queue, no replay. One run is one attempt. A push that fails is dropped, logged, and the process exits non-zero — the next beat is five minutes away and a heartbeat from five minutes ago describes a machine that no longer exists. Nothing is spooled to disk, so nothing can ever be replayed to make the dashboard show a past that is no longer true.

No secret in the logs. run-push.sh is meant to be piped into logger, so everything the agent prints ends up in syslog and journald permanently. The failure path therefore reports only err.code and the HTTP status — never an error object, never the request options, never a response body, and never a header. That is also why run-push.sh must never gain a set -x: the shell would echo the token as it sources the env file.

The cost is real and accepted: a 400 tells you the server rejected the payload, not why. The reason is in the server's logs, and --dry-run (below) shows you the exact payload that was sent.

Install on a new workstation

1. Freeze a copy

The cron must point at a copy, never at the git working tree. The server is deployed by a manual trigger, so the checked-out repo drifts freely: a branch checkout in ~/claude-code/vps-health-api would silently change — or break — the heartbeat of the workstation, and nothing would say so.

push-metrics.js loads metrics.js from one directory up, so the copy keeps the repo layout:

mkdir -p ~/.local/share/maximus-host-agent/agent
cp metrics.js ~/.local/share/maximus-host-agent/
cp agent/push-metrics.js agent/run-push.sh ~/.local/share/maximus-host-agent/agent/
chmod +x ~/.local/share/maximus-host-agent/agent/run-push.sh

Re-run those three cp after any change to the agent — that copy step is the deployment.

2. Write the env file

~/.config/maximus-host-agent.env, mode 600 (the agent warns on any other mode). Use printf, not a heredoc: a trailing newline or space inside the token value is the most common install failure.

HOSTS_API_URL=https://health.lacompagniemaximus.com
HOSTS_INGEST_TOKEN=<the ingest token>
HOST_ID=thinkpad
# NODE_BIN=/usr/bin/node          # only if cron cannot find node (see below)
# HOST_AGENT_TIMEOUT_MS=10000     # push timeout, wall-clock; default 10000
chmod 600 ~/.config/maximus-host-agent.env

HOSTS_INGEST_TOKEN is the write-only token: it can push snapshots for an allowlisted id and nothing else. It is deliberately not HEALTH_TOKEN, so that the two can be rotated independently and so a workstation that only pushes never needs the read token.

Do not read more into that separation than it gives you. On the ThinkPad it buys no containment at all: that machine already stores HEALTH_TOKEN in cleartext for the defenseur-auto cron, so losing the laptop compromises both tokens. Rotate them together — see secret-rotation-ops.md.

HOST_ID must match ^[a-z0-9][a-z0-9-]{0,31}$ and be listed in the server's HOSTS_ALLOWED_IDS, otherwise the push comes back 403.

3. Dry run

Collect and print the payload without touching the network or needing any config:

node ~/.local/share/maximus-host-agent/agent/push-metrics.js --dry-run

Then one real push, by hand, before installing the cron:

~/.local/share/maximus-host-agent/agent/run-push.sh
# host-agent: ok id=thinkpad status=204 hostname=thinkpad-x1

Confirm the server side (from a machine that holds HEALTH_TOKEN):

curl -H "Authorization: Bearer $HEALTH_TOKEN" \
  https://health.lacompagniemaximus.com/hosts

The entry for your id should show online: true and a small ageSeconds.

4. Install the cron

crontab -e
*/5 * * * * $HOME/.local/share/maximus-host-agent/agent/run-push.sh 2>&1 | logger -t host-agent

Read what it did:

journalctl -t host-agent --since -1h

If the line works by hand but produces nothing under cron, it is almost always node: cron's PATH is /usr/bin:/bin, and a node installed by nvm or under /usr/local is not on it. Set NODE_BIN to an absolute path in the env file — the wrapper says so explicitly rather than failing silently.

One id per workstation

The server keeps one file per host id and the last write wins. Two workstations configured with the same HOST_ID therefore overwrite each other every five minutes, and nothing anywhere reports an error: the dashboard shows one host that looks perfectly healthy while its numbers come from whichever machine pushed last.

The signal to watch for is the hostname field in GET /hosts changing between polls while the id stays the same — plus CPU and memory that jump around without pattern. If you see that, two machines are sharing an id.

Give every workstation its own id and add it to HOSTS_ALLOWED_IDS on the server before installing the agent on it.

Exit codes

Code Meaning
0 Snapshot accepted (server answered 2xx, normally 204)
1 Install problem: env file missing, a variable unset, malformed HOST_ID, node not found, metrics.js not found next to the agent
2 The push did not land: network error (code=…) or a non-2xx answer (HTTP …)

A 1 needs a human on the workstation. A 2 usually fixes itself on the next beat; a 2 that repeats for an hour means the API or the link is down.

Payload

Exactly the GET /health body minus logto and timestamp:

{ hostname, uptime,
  cpu:    { model, cores, loadAvg, usagePercent },
  memory: { totalGB, usedGB, freeGB, usagePercent },
  disk:   { totalGB, usedGB, freeGB, usagePercent } }

Sharing metrics.js with the server keeps the two in step, but it is not the guarantee — the frozen copy can lag behind a deployed server. The real contract is the server's own validation, and __tests__/agent.test.js checks the agent's payload against it (including a real snapshot through the real handler).