Files
deploy_home/ansible/roles/podman
Bastian de BylandClaude Opus 5 91d9f6b419 fix: actually bound the Zomboid log archive, which logrotate could not
The logrotate config added with the Sophie overhaul claimed to bound PZ's own
logs. It did not. data/Logs had reached 18 GB across 112 entries and 146k
files, and rotation had never once fired there.

Two reasons, both wrong assumptions on my part. PZ rolls its logs into dated
directories -- logs_2025-12-14/ through logs_2026-08-26/, up to 279 MB each --
so the Logs/*.txt glob only ever matched a handful of loose files at the top.
And it rotates by size: not one file in that tree exceeds 100M, because the
growth is in the number of files, not the size of any of them.

Age-based pruning is the right tool, so this adds a daily zomboid-log-prune
timer keeping 14 days and deleting the emptied directories behind it. It runs
under podman unshare, since the files belong to the container's UID.

logrotate keeps server-console.txt, which is a single ever-growing file and
genuinely is what it is good at. Narrowed its scope to say so rather than
implying coverage it never had.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 02:34:30 -04:00
..

podman role

Container orchestration for the home server. Containers are defined under tasks/containers/{base,home,skudak,debyltech}/ and wired from tasks/main.yml, which is where every image tag is pinned.

Backups

Every Nextcloud-style instance (Nextcloud, Gitea, BookStack, partsy) shares one backup engine: tasks/containers/cloud-backup.yml plus templates/nextcloud/cloud-backup.sh.j2. Each instance includes it with its own vars, producing /usr/local/bin/<name>-backup.sh and a systemd timer.

Stages, in order — the ordering is deliberate, see the comments in the template:

  1. Database dump inside the container (mariadb / mysql / postgres branches), gzipped to /var/backups/nextcloud/<name>/db/. Promoted over yesterday's dump only after passing gzip -t and a completion-trailer grep.
  2. SQLite snapshots (.backup, then pragma integrity_check) where used.
  3. rsync of the data tree, config, and db dumps to TrueNAS.

Failures raise status=failed on the nextcloud-backup syslog tag. The mail itself is sent by the unit's own OnFailure= handler (templates/nextcloud/nextcloud-backup-alert.sh.j2) straight through sendmail, so alerting does not depend on Graylog and is unaffected by graylog_enabled being off. A Graylog rule matching the same tag is a secondary, dashboard-side copy of that signal.

Do not rename that tag: it is shared by every instance including Gitea, and renaming it here silently stops the Graylog-side alerting for all of them.

Restore

Untested backups are not a control. Rehearse this into a scratch location before you need it, and record the date you last did.

1. Database

Dumps live on the host at /var/backups/nextcloud/<name>/db/<name>-YYYYMMDD.sql.gz and on TrueNAS at <remote_path>/_backup/db/. TrueNAS in turn cloud-syncs /mnt/glacier/skudakcloud to Skudak's own iDrive e2 bucket, so a third copy exists there — but restoring from it means going through the TrueNAS console, not this host.

Verify the dump before trusting it:

gzip -t <name>-YYYYMMDD.sql.gz
gunzip -c <name>-YYYYMMDD.sql.gz | tail -c 512   # expect the completion trailer

Replay into the running database container. MariaDB/MySQL:

sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \
  | podman exec -i <db_container> sh -c \
    "exec env MYSQL_PWD=\$MYSQL_ROOT_PASSWORD mariadb -u root \$MYSQL_DATABASE"'

Postgres dumps are taken with --clean --if-exists --no-owner, so they replay into an existing database:

sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \
  | podman exec -i <db_container> sh -c \
    "exec env PGPASSWORD=\$POSTGRES_PASSWORD psql -U \$POSTGRES_USER \$POSTGRES_DB"'

2. Data tree

# from TrueNAS
rsync -az -e "ssh -i /etc/ssh/backup_keys/<name>" \
  <ssh_user>@truenas.localdomain:<remote_path>/ <volumes>/<instance>/data/

Then fix ownership — the containers run as uid 33 inside a rootless userns:

sudo -H -u podman bash -c 'cd; podman unshare chown -R 33:33 <volumes>/<instance>/data'

3. Reconcile

sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ maintenance:mode --on'
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ files:scan --all'
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ maintenance:mode --off'

A DB snapshot slightly older than the files degrades to "files the app has not indexed yet" and is repaired by files:scan. A DB snapshot newer than the files references blobs that were never backed up, which surfaces as broken shares and dead file entries — this is why the dump runs first.

4. LibreSign-specific

The signing CA lives in the data tree at data/appdata_*/libresign/pki/<instance>_<n>_openssl/, so a data-tree restore brings it back with everything else. After restoring, confirm it:

sudo -H -u podman bash -c 'cd; podman exec -u www-data skudak-cloud php occ libresign:configure:check'

Every check must report success. If openssl-configure reports an error, the certificate_engine / config_path app config is pointing somewhere without a CA — see the guarded generate task in tasks/containers/skudak/cloud.yml. Do not simply re-run libresign:configure:openssl on a restored instance without understanding why: it mints a new root CA and invalidates the trust chain on every document already signed under the old one.

LibreSign

Deployed on skudak-cloud only. LibreSign 14.1.0 requires Nextcloud server >=34.0.0,<35.0.0, which the pinned nextcloud:34.0.2-apache satisfies. If the Nextcloud tag is bumped to 35, LibreSign must be held or upgraded in step — the two instances are pinned independently in tasks/main.yml, so skudak-cloud can lag cloud if needed.

Dependency split, which drives what survives a container recreate:

Component Location Survives recreate?
Java (JRE 21), PDFtk, jSignPdf data/appdata_*/libresign/ Yes — persisted volume
Root CA / PKI data/appdata_*/libresign/pki/ Yes — persisted volume
poppler-utils, ghostscript /usr in the image No — reinstalled by Ansible each run

Certificate engine is OpenSSL, not CFSSL. CFSSL is the more common source of LibreSign setup failures and buys nothing at this scale.

Gotchas

  • Do not pass --ou to libresign:configure:openssl. LibreSign appends its own libresign-ca-id:... entry to the OU field, and the combined value overruns the 64-character ASN.1 limit for organizationalUnitName, failing with string too long.
  • A disabled app has no occ commands. If occ list | grep libresign returns nothing, the app is disabled, not missing — occ app:list will still show it under Disabled:. This is what a Nextcloud major upgrade does to an app it thinks is incompatible.
  • PHP_MEMORY_LIMIT must be raised above the 512M image default; signing fails opaquely mid-operation otherwise.
  • LC_ALL / LANG must be set or the JVM comes up as ANSI_X3.4-1968 and LibreSign warns that accented characters in signer names will be mangled (LibreSign issue #4872).

Logging

Every container runs with log_driver=journald, so container stdout lands in the host journal, capped at 500M by roles/common/templates/journald-size.conf.j2 (about 25 days at the current rate). journalctl CONTAINER_NAME=<name> is the day-to-day way to read it.

On top of that sits an optional Graylog stack — graylog, graylog-opensearch, graylog-mongo, fed by a host fluent-bit service that tails the journal and ships GELF to 127.0.0.1:12202, enriched with the MaxMind GeoIP database, and configured over the REST API by the separate graylog-config role.

It is off. The switch is graylog_enabled in inventories/home/hosts.yml, and it gates all five of those pieces at once:

graylog_enabled effect
true stack deployed, fluent-bit shipping, GeoIP downloaded, graylog-config runs, logs.debyl.io proxies to the UI and /gelf
false containers removed and their systemd user units disabled and deleted, fluent-bit stopped and disabled, GeoIP and graylog-config skipped, logs.debyl.io answers 503

It was turned off because the cost/benefit is bad on a 4-core box: two JVMs plus MongoDB held ~1.6 GB resident and ~3% of the CPU continuously to store roughly 3k messages a day — about 28 MB of real log data across four live indices — all of which journald already keeps for longer.

Turning it back on is graylog_enabled: true plus make deploy TAGS=graylog. Nothing is destroyed by the off path: the volumes under {{ graylog_path }} keep the indices, the Mongo database holding streams/pipelines/dashboards, and the node-id file, so the stack comes back with its configuration intact.

Two things genuinely stop while it is off, both by design:

  • Search and dashboards. The logs still exist in journald; the query interface over them does not.
  • External GELF ingest. The AWS Lambda that POSTs to logs.debyl.io/gelf for the debyltech-api stream gets a 503. Those events are dropped, not queued — nothing else records them.

Caddy config reloads

The reload caddy handler deliberately reads /config/Caddyfile, not the /etc/caddy/Caddyfile the container starts from, even though both are the same host file. /etc/caddy/Caddyfile is a single-file bind mount, which podman binds by inode, and Ansible's template module writes a temp file and renames it into place — so every deploy gives the host file a new inode while the container keeps seeing the one it was created with. Reloading from that path silently re-applied the previous config; changes only landed when something recreated the container. {{ caddy_path }}/config is also bind-mounted as a directory at /config, and directory mounts resolve names at open() time, so /config/Caddyfile is always the file Ansible just wrote.