Post

Docs Drift Detector: Keeping Posts Honest About Versions

Docs Drift Detector: Keeping Posts Honest About Versions

Purpose

Posts go stale the moment Renovate bumps an image tag. A weekly job checks which versions each post documents against what is actually running, generates a rolling issue with the drift report, and any fixes follow the same docs pipeline as a pull request that references the issue rather than closing it. See also the base system post and the agents overview.

Architecture

The drift detector is tools/docs-drift.py, a pure function that takes a repo clone and inventory files as input and produces a report. It has no network access, no credentials, and makes no remote calls — the Hermes runner gathers the inventory separately and feeds it in.

It compares image references found in every post (backticked repo:tag strings, YAML snippets, and table Image rows) against a live inventory built from two sources:

  • k3s workloads via a read-only Kubernetes service account (kubectl get deploy,sts,ds -A -o json)
  • Docker hosts via a read-only Portainer MCP wrapper that normalizes container images

Comparison happens by repository path — the registry host is ignored so docker.io/library/postgres:16 matches a post saying just postgres:16. Four drift classes are tracked: version_drift (post names a tag no workload uses), latest_documented (post says latest but Renovate pinned it), not_running (the image runs nowhere), and undocumented (something runs that no post mentions — informational only). A config file, tools/drift-config.json, controls ignored namespaces, intentionally ignored images, and maps host names to public-safe roles.

Permissions and roles

The pattern used across all Hermes agents mirrors the base system: a deterministic collector gathers facts without an LLM, rules decide what is a finding, and a root-owned gate tool validates before publishing. The drift detector runs under its own Unix user with only the rights it needs — the read-only k3s service account cannot execute into pods or read Secrets, the Portainer wrapper has no write operations, and report-issue validates file ownership, blocks symlinks, and checks output against both gitleaks and the site’s sanitize check before touching the repo.

How it runs

The job fires every Sunday at 03:00 in Hermes “no-agent” cron mode — the script runs on schedule and its stdout is delivered directly, with no LLM turn involved. The flow:

  1. The drift detector runs against freshly gathered inventory facts
  2. It produces a Markdown report listing drifting posts
  3. report-issue --report docs-drift upserts a single rolling issue in the Doc-Site-Build repo; if the body is identical to the last run, nothing changes
  4. When drift rows appear, the docs queue picks them up — one card per post per run
  5. A doc writer fixes the stale references and opens a PR that says Refs #N so the rolling issue stays open until the next weekly run drops the fixed row

The first fix PR updated a monitoring post’s image version after the initial rollout.

flowchart LR
    A["Post repository<br/>(clone)"] --> C["docs-drift.py<br/>(compare)"]
    B["Running workloads<br/>(k3s SA + Portainer MCP)"] --> C
    C --> D["Drift rows<br/>(version, latest, not running)"]
    D --> E["report-issue<br/>(gate: sanitize + gitleaks)"]
    E --> F["Rolling issue<br/>(Doc-Site-Build)"]
    F --> G["Docs queue card<br/>(one post per run)"]
    G --> H["Doc writer<br/>(fix references)"]
    H --> I["PR Refs #N<br/>(merge, no close)"]
    I --> J["Next Sunday run<br/>(drops fixed rows)"]

How it relates to the other agents

The drift detector was the first automated report to use the report-issue gate tool. Compliance checks (timezone, deploy provenance, logging coverage, repo inventory) and ops digest reports reuse the same pattern: deterministic collection → rule evaluation → validated upsert into a rolling issue. Unlike those, the drift report goes to this public repository so it passes the sanitize check along with gitleaks; the ops reports target a private repo and skip sanitization since they carry real hostnames and IPs.

Notes

A post tagged retired is skipped during extraction. Single-name images like postgres:16 are only flagged if that image actually runs somewhere in the cluster — otherwise the regex matches too many library references that aren’t deployment facts. The config ignore list can also carve out intentional differences, such as a development build deliberately running an older tag.

This post is licensed under CC BY 4.0 by the author.