The Daily Ops Digest: Twelve Collectors, One Rolling Issue
The ops digest is a daily health check for the entire homelab. It runs without an LLM in the loop, uses twelve independent collectors, and posts results to one rolling issue that only changes when something real changes. This post covers where each collector gets its data, which rules turn data into a finding, why the newer features run in shadow mode, and how the digest fits into the agent pipeline described in the agents overview. For context on the system itself, see the Hermes Agent post.
Purpose
A homelab with a Proxmox host, Kubernetes, Docker hosts, and external services generates enough signals that checking them by hand every day is tedious. The ops digest automates the first pass: it gathers state from twelve sources, applies thresholds to decide what is noteworthy, produces a structured report, and posts only when findings change. If nothing changed since yesterday, the issue stays as it is, so there is nothing new to scroll past.
Architecture
The pipeline moves data in one direction: sources feed collectors, collectors feed rules, and the output goes through a gate tool that validates everything before it is posted. The ops-digest-run script is the runner. It only reads from the homelab; the one thing it changes is the rolling issue in a private ops repository, and it does that through the report-issue gate tool. Each update edits the issue body and adds one comment with the counts of new and resolved rows.
The twelve collectors break down into three phases:
Phase 1: Direct reads from Kubernetes via the hermes-ro read-only service account (argocd app health and sync status, pod crash loops and restarts, warning events, Longhorn volume states, certificate expiry) plus Gatus uptime checks and Docker container counts per host through a read-only Portainer wrapper.
Phase 2: Facts pushed by the Proxmox host, which the agent has no login on. Every morning at 05:10 the host pushes one JSON file over SSH, using a key that is pinned to a single forced command and to the host’s own address. The file covers guests and backup jobs (is every running guest in a job, and did its last backup succeed), OOM kills per guest, ZFS pool health, capacity and scrubs, and the nightly backup-server cycle.
Phase 3: Graylog health (journal freshness, notification state, indexer failures), agent self-checks (whether every scheduled job ran and no queue is stuck), the latest Kubescape triage summary, and, in shadow mode, log error signatures and a summary of public edge traffic. Shadow mode means these sections show their data but raise no findings until they have enough history to tell normal noise from a real problem.
Each collector produces either a JSON file on success or an .error file with a single line of the failure message on failure. The ops-digest.py script turns the output into one Markdown report; it is a pure function with no external dependencies. If a collector fails entirely, that failure becomes a finding instead of being dropped.
flowchart LR
K3S["k3s read only SA"] -->|"argocd, pods<br/>events, longhorn,certs"| C1(Phase 1)
GA[Gatus API] --> C1
PT[Portainer wrapper] -->|"Docker counts"| C1
PXE[Proxmox host] -->|"pushed facts"| PF["ops-facts-receive"]
PF --> C2(Phase 2)
GR[Graylog API] --> C3(Phase 3)
HERM[Hermes cron/kanban] --> C3
KU[Kubescape summary] --> C3
C1 --> R["Rules thresholds"]
C2 --> R
C3 --> R
R --> BR[Report body]
BR --> GT["Gate tool<br/>report-issue"]
GT --> GI["Rolling issue<br/>(private repo)"]
GI --> HU[Human review]
Permissions and roles
Every step runs under its own identity with only the permissions it needs:
- Kubernetes: The hermes-ro service account can list deployments, pods, events, and CRDs. It can’t read secrets or exec into pods. If a collector fails (kubectl error, authentication failure), that is reported as a finding rather than dropped silently.
- Graylog: A read-only service account token stored in a file owned by root:hermes (0640) gives read access to the Graylog API for health checks and query aggregation but cannot create streams or write logs.
- Facts receiver: The
ops-facts-receivescript runs as a dedicated Unix user (opsfacts) with a forced-command SSH key. It accepts only from authorized hosts, validates JSON structure and enforces a size cap (under 2 MB), and writes the file atomically into a setgid directory, so hermes can read it but not write it. If the host name is not allowed or the payload fails validation, the push is rejected. - Publish gate: The
report-issuetool validates file ownership (with its own sudo rule). It checks that the body passes gitleaks and (for public repos) the sanitize check before upserting the issue.
How it runs
The ops digest runs daily at 05:30, twenty minutes after the Proxmox host pushes its facts.
The runner follows these steps (from its header comment):
- Clone or update the private ops repository using
GIT_CONFIG_*environment variables so the token never appears in command-line arguments. - Run each collector, writing
<name>.jsonon success or<name>.errorwith a single line of the error message on failure. - Pass all output files to
ops-digest.py, which applies the thresholds fromops-digest.jsonand writesreport.md(a pure function that uses only the Python standard library). - Only if the report fingerprint differs from the last posted one, call
report-issue(as a separate user via its sudo rule) to upsert the rolling issue body plus one comment listing new/resolved row counts. - If a critical finding is new, attempt to mail an alert through the notify tool; if that fails, the keys are forgotten so the next run retries.
The fingerprint covers each finding’s severity, key and text, not its details. A counter that keeps growing, such as a restart count, therefore doesn’t rewrite the issue every morning.
How it relates to the other agents
The ops digest shares the same three-layer pattern as every other agent in the crew: deterministic collectors gather facts, rules decide what is a finding, and a gate tool publishes. No LLM is involved in collecting, judging or publishing. The language model only writes prose: the docs briefs and drafts.
Specifically with the other scheduled jobs:
- Config compliance runs at 05:20 and keeps its own rolling issues (timezones, logging coverage, workload provenance, repo inventory). The ops digest only checks that the job ran on time.
- Kubescape scans at 06:00. Its triage step compares the controls with a 7-day baseline (newly failing, fixed, score change), and the next morning’s digest reads that summary through its “kubescape” collector.
- Docs drift detector runs on Sundays at 03:00 and keeps its own rolling issue. The ops digest only checks that the job ran.
Data flows in one direction. The ops digest watches the other agents’ health but never changes their state, and every output ends with a person reading an issue or reviewing a PR.
Notes
The first runs found a container that was in no backup job, and the OOM history of another container that nobody had noticed.
Two newer features run in shadow mode: log triage groups error lines from Graylog into normalized signatures (variable parts masked with regex), and compares them against a 7-day history file. The edge watch summarizes the traffic that reaches the public Traefik routers (volume per router, bursts of failed logins, sensitive paths that answered, routers that became public) from outside clients only. Neither raises a finding until it has enough history to tell noise from a real pattern.
The trade-off is speed for safety. A finding discovered at 05:30 does not get fixed automatically. It sits in the rolling issue until someone reviews it, decides on a fix, or acknowledges why it is expected with an expiration date (expired acknowledgments reappear as findings). In exchange, nothing runs with more permissions than it needs to complete its job, no secrets can escape the wrappers, and every write is a visible artifact that survives beyond a single session.