Post

Hermes Agent: A Local AI Ops Agent with Guardrails

Hermes Agent: A Local AI Ops Agent with Guardrails

Hermes Agent

An AI agent with real access to a homelab, limited by the identities and permissions it holds instead of whatever prompt says. You can hand an LLM meaningful operational reach without trusting it to stay in its lane because every boundary is enforced by RBAC, file ownership, fork isolation, or a wrapper the model cannot modify. Prompt instructions are the last line of defense.

Purpose

Hermes Agent (open-source agent framework from Nous Research) runs as an always-on assistant that reads cluster state, triggers security scans, searches documentation repositories, and drafts posts for the public site – without ever being able to escalate, exfiltrate secrets, or run arbitrary code on infrastructure hosts. The goal is to show what safe, local AI ops access looks like in practice.

Architecture

The agent runs in an unprivileged Debian 13 LXC container on the Proxmox host, version v0.21.5. The language model is fully local: Qwen 3.6 27B served by Ollama on the GPU Docker host via its OpenAI-compatible endpoint.

The model requires a context window of at least 64K tokens. Ollama’s endpoint ignores the num_ctx parameter in requests, so the fix was a derived model with PARAMETER num_ctx 65536 baked into the Modelfile. It consumes about 20 GB of the GPU’s 24 GB VRAM – when a voice assistant workload needs the same GPU, Ollama swaps models. That contention is an accepted trade-off for running everything locally.

Web search goes through a self-hosted SearXNG instance. No third-party AI or search APIs are involved.

The toolchain on the box includes linters, security scanners (Trivy, Semgrep, Checkov, Gitleaks, Grype), IaC CLIs (kubectl, Helm, Terraform, Ansible), language servers, and documentation tooling (Pandoc, mermaid-cli, draw.io). The enabled toolsets include web search/extract, terminal, file operations, code execution, vision, skills, delegation, cron jobs, and four MCP servers (postgres, gitea, portainer, gitea_docs). Skills provide reusable instruction sets for the homelab map, a docs writer workflow, a humanizer pass, and structured planning and debugging procedures.

Guardrails

The agent can see things but only ever do what its identities allow. Each integration has an out-of-model enforcement mechanism:

IntegrationScopeEnforced by
Kubernetesget/list/watch on workloads, logs, and CRDs (ArgoCD, Traefik, Longhorn, cert-manager, MetalLB). No Secrets, no exec/attach/port-forward. HelmChart CRDs excluded (values may carry credentials)Dedicated ServiceAccount RBAC. Verified with kubectl auth can-i; get secrets -A returns no, create pods/exec returns no
Security scanDaily Kubescape scan runs under a separate system user. Agent triggers only a root-owned wrapper (no arguments). Scanner RBAC lives in a different repo from results, so the scanner cannot widen its own permissions through GitOpsSudo rule + file ownership + repo separation
Git (read)Restricted bot account with read-only token scoped to repos, issues, PRs, and CI runsGitea token scopes. The account sees only what it is explicitly given
Git (docs writing)Agent works in its own fork and opens PRs. Fork PRs do not receive repo CI secrets; Actions are disabled on the fork (self-hosted runners have Docker access, so a workflow in any branch equals code execution on that runner). The write token is held by another system user behind a wrapper with a tool allowlist (no CI config, no deletes, no repo creation). CI runs on the fork PR and waits for owner approval. main requires passing checks, only the owner mergesFork model + wrapper allowlist + branch protection + manual workflow approval
Container platform (Portainer)Wrapper exposes Docker inventory only – Kubernetes access is exclusively through the dedicated ServiceAccount (Kubernetes row), so the no-Secrets rule cannot be bypassed via Portainer. Raw Docker/Kubernetes proxy tools are disabled (GET-only was not enough: GET through a raw proxy can still read Kubernetes Secrets or download container files). Community Edition has no read-only role; environment variables redacted; API key readable only by the wrapper’s own userWrapper + file permissions
DatabaseRead-only MCP server using PostgreSQL role with pg_read_all_data and default_transaction_read_only, behind ingress basic authDatabase role
flowchart LR
    HA[Hermes Agent<br/>unprivileged LXC]

    subgraph DirectReadOnly["Direct read-only identities"]
        K8S["Kubernetes API<br/>(ServiceAccount)"]
        GITR["Gitea (read token)"]
        DB["PostgreSQL MCP<br/>(read-only role)"]
    end

    subgraph WrapperHeld["Wrapper-held credentials<br/>(agent cannot read secret)"]
        PORT["Portainer MCP<br/>(GET only wrapper)"]
        GITW["Docs Git MCP<br/>(fork-only wrapper)"]
        SCAN["Kubescape scan<br/>(root-owned wrapper)"]
    end

    subgraph DocsPath["Documentation write path"]
        FORK[Fork branch] --> PR{Pull Request}
        PR -->|Owner approves| CI[CI checks<br/>build, lint, sanitize, gitleaks]
        CI -->|Passes| MERGE["Owner merges →<br/>live public site"]
    end

    HA -- "get/list/watch<br/>no secrets" --> K8S
    HA -- "repos, issues, PRs" --> GITR
    HA -- "pg_read_all_data<br/>read_only=on" --> DB
    HA -- "tool allowlist" --> PORT
    HA -- "fork + PR only" --> GITW
    HA -- "no-arg wrapper" --> SCAN
    GITW --> FORK

    style HA fill:#1a1a2e,stroke:#e94560,stroke-width:2px
    style DirectReadOnly fill:#f0f0f0,stroke:#888
    style WrapperHeld fill:#fff3cd,stroke:#ffc107
    style DocsPath fill:#d4edda,stroke:#28a745

Skills

Skills are the reusable context that keeps each session from starting at zero. The important ones:

  • Homelab map: a compact reference of hosts, IPs, where things run, and which read-only access path to use. The agent skips the discovery phase and goes to the right tool.
  • Doc-writer: conventions, sanitization rules, template sections, check commands, and the fork-to-PR workflow for the public documentation site. This post was drafted through it.
  • Humanizer: a prose pass that rewrites AI-sounding text into something a human operator would actually publish.
  • Planning and debugging workflows: structured approaches for multi-step implementation plans and systematic root-cause analysis.

Lessons learned

GET-only is not the same as safe. Making the Portainer account an admin was needed to see containers, and the MCP was GET-only, but its raw proxy tools could still read Kubernetes Secrets and container files. The fix was to remove capabilities at the wrapper (Docker-only, no proxy) rather than rely on “read-only”.

“Read-only” is not always a product feature. Portainer CE has no read-only role, so the solution was a wrapper that only exposes GET/HEAD methods and redacts environment variables. When the tool does not have fine-grained roles, build the boundary outside it.

Self-hosted CI runners execute whatever a branch’s workflow says. That fact drove the fork-only design for documentation writes. The runner has Docker access, so a malicious or prompted workflow in any fork branch would be code execution on that host. Fork PRs intentionally do not receive repo secrets, Actions are disabled on the fork itself, and separately CI for a fork PR on the main repo waits for the owner’s approval.

Auto-merge bots merged before checks ran. Branch protection rules initially allowed merges without requiring status checks, so making checks required was the fix. No one can merge to main without green CI now.

Local LLM plumbing has its own issues. The Ollama context-size problem (endpoint ignoring request parameters) required baking the config into a derived model. Having one 20 GB inference load next to other AI workloads on the same GPU means models swap in and out, adding latency when the workload changes.

Notes

The Hermes Agent framework gives you the scaffolding; the guardrails are what you build around it. Prompt instructions alone do not make a system safe – an LLM ignores them under pressure, jailbreak attempts, or simply bad reasoning. What matters is that every credential is scoped to the minimum identity needed, and every boundary lives in a config file the model does not control.

The toolchain includes CI checks on documentation PRs: build validation, HTML proofing, markdown linting, secret scanning with Gitleaks, sanitization rules, and diagram rendering verification. The pipeline catches problems before merge, and the merge gate requires human approval.

Setup version: Hermes Agent v0.21.5 with Qwen 3.6 27B on Ollama. Fully local – no cloud AI APIs, no third-party search providers, no telemetry leaving the homelab.

This post is licensed under CC BY 4.0 by the author.