Post

Immich: Self-Hosted Photos with GPU Machine Learning

Immich: Self-Hosted Photos with GPU Machine Learning

Purpose

Immich is a self-hosted photo and video management system built as a Google Photos replacement. It runs on the GPU Docker host in a four-container compose stack: uploads, smart search, face recognition, and background processing all go through Traefik and ship logs to Graylog using the same pipeline as the rest of the lab.

Deployment

Immich runs on the dedicated GPU Docker host (NVIDIA), not in Kubernetes. Portainer manages the four containers as a compose stack:

ContainerImage RoleVersion
immich_serverAPI, web UI, background jobsghcr.io/immich-app/immich-server:v3.1.0
immich_machine_learningFace detection/recognition, CLIP smart searchghcr.io/immich-app/immich-machine-learning:v3.1.0-cuda
immich_postgresDatabase with vector extensions for embeddingsghcr.io/immich-app/postgres:14-vectorchord0.3.0-pgvectors0.2.0
immich_redisTask queue and caching (Redis-compatible)valkey/valkey:8-bookworm

The ML container uses the CUDA variant, which mounts the host’s NVIDIA GPU directly for hardware-accelerated inference. Immich ships its own Postgres image with vectorchord and pgvectors preloaded — no manual setup needed, unlike the SQL Server Postgres instance which runs in a separate VM with custom tuning.

GPU acceleration

The ML service handles two GPU workloads:

  • Face detection and recognition using deepface models on CUDA
  • CLIP smart search, which does embedding-based semantic queries (“photos at the beach” without relying on tags)

Video transcoding runs on CPU in this deployment — no hardware acceleration profile is enabled for the server container.

Architecture

flowchart LR
  subgraph Clients
    W["Web UI"]
    M["Mobile App"]
  end

  subgraph Immich["Immich Containers (GPU Docker Host)"]
    S[Server<br>v3.1.0]
    ML["ML (CUDA)<br>v3.1.0-cuda"]
    PG["Postgres + vectors<br>14-vectorchord/pgvectors"]
    RK["Valkey 8<br>(Redis-compatible)"]
  end

  subgraph Storage["Photo Library (ZFS)"]
    L[Library]
  end

  S -->|queries| PG
  S -->|tasks| RK
  S -->|infer| ML
  S -->|store/retrieve| L

  ML -- GPU device --> G[(NVIDIA GPU)]

  W & M -->|HTTPS via Traefik| S

  subgraph Logging
    GD["gelf driver<br>(all 4 containers)"]
    GU["GELF UDP input"]
    GP["Immich pipeline<br>(normalize levels + ANSIs)"]
    GR[Graylog]
  end

  Immich --> GD --> GU --> GP --> GR

Configuration

Each container uses environment variables from the stack’s .env file. Notable settings:

  • Server logs in JSON format, so every line is a structured record with level, context, pid, and stack fields — ready for Graylog to parse without extra configuration.
  • ML has NO_COLOR set and an explicitly wide COLUMNS value to suppress color codes from Python’s logging library. The framework still emits some residual escape sequences, so the log pipeline has to clean them up anyway.

The Postgres container uses the same internal Docker network as the others, with no external port exposure. Clients talk only to the server.

Storage

The photo library is on the host’s ZFS pool. Backups fall under the broader Ceph cluster strategy. The database lives alongside other service data in regular backups.

Unlike Kubernetes workloads that use PVCs with local-path or Longhorn, this Docker compose stack mounts volumes directly from the host — no storage class to configure.

Centralized Logging

All four containers send logs to Graylog via Docker’s native gelf log driver, reusing the same UDP input. Each container carries a node_name tag so Graylog can filter by host.

The Immich processing pipeline

A standalone Graylog pipeline processes Immich logs:

  1. Drop stray extractor fields — pre-existing extractors without conditions were catching Postgres timestamp lines and creating invalid date fields, which caused silent indexing failures.
  2. Parse per container — the server’s JSON log lines get unpacked into discrete fields (level, context, pid, stack trace) instead of staying as one big message blob.
  3. Clean ML ANSI codes and levels — even with NO_COLOR, Python emits residual [2m sequences. Docker’s GELF driver strips the leading escape byte but leaves the bracket code behind. The pipeline removes those remnants and normalizes level names (log to info, warn to warning).

Gotcha: extractor without a condition

A Graylog extractor ran on every message instead of only matching ones it understood. Timestamp-prefixed Postgres lines hit the unconditional extractor, created invalid date fields, and got silently rejected at indexing — no error in the UI, no backpressure, just missing search results. The Immich pipeline’s first step drops this field before any other stage runs, preventing the same data loss for all four containers.

Access

The web UI routes through Traefik on the GPU Docker host, behind a wildcard TLS certificate shared with internal homelab services. The mobile app connects directly to the server instead of going through Traefik, avoiding an extra hop for large file uploads.

Authentication uses standard Immich user accounts — no SSO or external auth providers. The service lives at an internal subdomain of homelab.example, same pattern as everything else in this series.

Notes

  • The -cuda ML tag bundles the full CUDA toolkit with PyTorch and model weights, making it 4.8 GB versus about 2 GB for the server image.
  • Valkey is a Redis drop-in replacement. Immich treats it like any other Redis connection — same protocol, newer maintenance.
  • The custom Postgres image ships vectorchord and pgvectors preloaded, so container startup works without manual extension setup.
This post is licensed under CC BY 4.0 by the author.