The logrotate config added with the Sophie overhaul claimed to bound PZ's own logs. It did not. data/Logs had reached 18 GB across 112 entries and 146k files, and rotation had never once fired there. Two reasons, both wrong assumptions on my part. PZ rolls its logs into dated directories -- logs_2025-12-14/ through logs_2026-08-26/, up to 279 MB each -- so the Logs/*.txt glob only ever matched a handful of loose files at the top. And it rotates by size: not one file in that tree exceeds 100M, because the growth is in the number of files, not the size of any of them. Age-based pruning is the right tool, so this adds a daily zomboid-log-prune timer keeping 14 days and deleting the emptied directories behind it. It runs under podman unshare, since the files belong to the container's UID. logrotate keeps server-console.txt, which is a single ever-growing file and genuinely is what it is good at. Narrowed its scope to say so rather than implying coverage it never had. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
podman role
Container orchestration for the home server. Containers are defined under
tasks/containers/{base,home,skudak,debyltech}/ and wired from
tasks/main.yml, which is where every image tag is pinned.
Backups
Every Nextcloud-style instance (Nextcloud, Gitea, BookStack, partsy) shares one
backup engine: tasks/containers/cloud-backup.yml plus
templates/nextcloud/cloud-backup.sh.j2. Each instance includes it with its own
vars, producing /usr/local/bin/<name>-backup.sh and a systemd timer.
Stages, in order — the ordering is deliberate, see the comments in the template:
- Database dump inside the container (mariadb / mysql / postgres branches),
gzipped to
/var/backups/nextcloud/<name>/db/. Promoted over yesterday's dump only after passinggzip -tand a completion-trailer grep. - SQLite snapshots (
.backup, thenpragma integrity_check) where used. - rsync of the data tree, config, and db dumps to TrueNAS.
Failures raise status=failed on the nextcloud-backup syslog tag. The mail
itself is sent by the unit's own OnFailure= handler
(templates/nextcloud/nextcloud-backup-alert.sh.j2) straight through
sendmail, so alerting does not depend on Graylog and is unaffected by
graylog_enabled being off. A Graylog rule matching the same tag is a
secondary, dashboard-side copy of that signal.
Do not rename that tag: it is shared by every instance including Gitea, and renaming it here silently stops the Graylog-side alerting for all of them.
Restore
Untested backups are not a control. Rehearse this into a scratch location before you need it, and record the date you last did.
1. Database
Dumps live on the host at /var/backups/nextcloud/<name>/db/<name>-YYYYMMDD.sql.gz
and on TrueNAS at <remote_path>/_backup/db/. TrueNAS in turn cloud-syncs
/mnt/glacier/skudakcloud to Skudak's own iDrive e2 bucket, so a third copy
exists there — but restoring from it means going through the TrueNAS console,
not this host.
Verify the dump before trusting it:
gzip -t <name>-YYYYMMDD.sql.gz
gunzip -c <name>-YYYYMMDD.sql.gz | tail -c 512 # expect the completion trailer
Replay into the running database container. MariaDB/MySQL:
sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \
| podman exec -i <db_container> sh -c \
"exec env MYSQL_PWD=\$MYSQL_ROOT_PASSWORD mariadb -u root \$MYSQL_DATABASE"'
Postgres dumps are taken with --clean --if-exists --no-owner, so they replay
into an existing database:
sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \
| podman exec -i <db_container> sh -c \
"exec env PGPASSWORD=\$POSTGRES_PASSWORD psql -U \$POSTGRES_USER \$POSTGRES_DB"'
2. Data tree
# from TrueNAS
rsync -az -e "ssh -i /etc/ssh/backup_keys/<name>" \
<ssh_user>@truenas.localdomain:<remote_path>/ <volumes>/<instance>/data/
Then fix ownership — the containers run as uid 33 inside a rootless userns:
sudo -H -u podman bash -c 'cd; podman unshare chown -R 33:33 <volumes>/<instance>/data'
3. Reconcile
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ maintenance:mode --on'
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ files:scan --all'
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ maintenance:mode --off'
A DB snapshot slightly older than the files degrades to "files the app has
not indexed yet" and is repaired by files:scan. A DB snapshot newer than
the files references blobs that were never backed up, which surfaces as broken
shares and dead file entries — this is why the dump runs first.
4. LibreSign-specific
The signing CA lives in the data tree at
data/appdata_*/libresign/pki/<instance>_<n>_openssl/, so a data-tree restore
brings it back with everything else. After restoring, confirm it:
sudo -H -u podman bash -c 'cd; podman exec -u www-data skudak-cloud php occ libresign:configure:check'
Every check must report success. If openssl-configure reports an error, the
certificate_engine / config_path app config is pointing somewhere without a
CA — see the guarded generate task in tasks/containers/skudak/cloud.yml.
Do not simply re-run libresign:configure:openssl on a restored instance
without understanding why: it mints a new root CA and invalidates the trust
chain on every document already signed under the old one.
LibreSign
Deployed on skudak-cloud only. LibreSign 14.1.0 requires Nextcloud server
>=34.0.0,<35.0.0, which the pinned nextcloud:34.0.2-apache satisfies. If the
Nextcloud tag is bumped to 35, LibreSign must be held or upgraded in step — the
two instances are pinned independently in tasks/main.yml, so skudak-cloud
can lag cloud if needed.
Dependency split, which drives what survives a container recreate:
| Component | Location | Survives recreate? |
|---|---|---|
| Java (JRE 21), PDFtk, jSignPdf | data/appdata_*/libresign/ |
Yes — persisted volume |
| Root CA / PKI | data/appdata_*/libresign/pki/ |
Yes — persisted volume |
| poppler-utils, ghostscript | /usr in the image |
No — reinstalled by Ansible each run |
Certificate engine is OpenSSL, not CFSSL. CFSSL is the more common source of LibreSign setup failures and buys nothing at this scale.
Gotchas
- Do not pass
--outolibresign:configure:openssl. LibreSign appends its ownlibresign-ca-id:...entry to the OU field, and the combined value overruns the 64-character ASN.1 limit fororganizationalUnitName, failing withstring too long. - A disabled app has no
occcommands. Ifocc list | grep libresignreturns nothing, the app is disabled, not missing —occ app:listwill still show it underDisabled:. This is what a Nextcloud major upgrade does to an app it thinks is incompatible. PHP_MEMORY_LIMITmust be raised above the 512M image default; signing fails opaquely mid-operation otherwise.LC_ALL/LANGmust be set or the JVM comes up asANSI_X3.4-1968and LibreSign warns that accented characters in signer names will be mangled (LibreSign issue #4872).
Logging
Every container runs with log_driver=journald, so container stdout lands in
the host journal, capped at 500M by roles/common/templates/journald-size.conf.j2
(about 25 days at the current rate). journalctl CONTAINER_NAME=<name> is the
day-to-day way to read it.
On top of that sits an optional Graylog stack — graylog, graylog-opensearch,
graylog-mongo, fed by a host fluent-bit service that tails the journal and
ships GELF to 127.0.0.1:12202, enriched with the MaxMind GeoIP database, and
configured over the REST API by the separate graylog-config role.
It is off. The switch is graylog_enabled in
inventories/home/hosts.yml, and it gates all five of those pieces at once:
graylog_enabled |
effect |
|---|---|
true |
stack deployed, fluent-bit shipping, GeoIP downloaded, graylog-config runs, logs.debyl.io proxies to the UI and /gelf |
false |
containers removed and their systemd user units disabled and deleted, fluent-bit stopped and disabled, GeoIP and graylog-config skipped, logs.debyl.io answers 503 |
It was turned off because the cost/benefit is bad on a 4-core box: two JVMs plus MongoDB held ~1.6 GB resident and ~3% of the CPU continuously to store roughly 3k messages a day — about 28 MB of real log data across four live indices — all of which journald already keeps for longer.
Turning it back on is graylog_enabled: true plus make deploy TAGS=graylog.
Nothing is destroyed by the off path: the volumes under {{ graylog_path }}
keep the indices, the Mongo database holding streams/pipelines/dashboards, and
the node-id file, so the stack comes back with its configuration intact.
Two things genuinely stop while it is off, both by design:
- Search and dashboards. The logs still exist in journald; the query interface over them does not.
- External GELF ingest. The AWS Lambda that POSTs to
logs.debyl.io/gelffor thedebyltech-apistream gets a 503. Those events are dropped, not queued — nothing else records them.
Caddy config reloads
The reload caddy handler deliberately reads /config/Caddyfile, not the
/etc/caddy/Caddyfile the container starts from, even though both are the same
host file. /etc/caddy/Caddyfile is a single-file bind mount, which podman
binds by inode, and Ansible's template module writes a temp file and renames
it into place — so every deploy gives the host file a new inode while the
container keeps seeing the one it was created with. Reloading from that path
silently re-applied the previous config; changes only landed when something
recreated the container. {{ caddy_path }}/config is also bind-mounted as a
directory at /config, and directory mounts resolve names at open() time,
so /config/Caddyfile is always the file Ansible just wrote.