Rolling a world back meant hand-work over ssh: stop the service, move the live save aside, unzip the right archive, chown into the container's subuid range, relabel, start. That is the wrong shape of task to do by hand, and it is always done under time pressure -- by construction, because the archive you want is being deleted while you work. PZ keeps BackupsCount=10 per set and writes one every BackupsPeriod=30 minutes, so a periodic backup is reachable for about five hours and then gone. On 2026-09-05 the snapshot the admins asked for (05:16, four minutes before the incident) had about 90 minutes of life left when the request came in. Same shape as the wipe: the Discord bot writes a trigger file into its own rw volume, zomboid-restore.path notices it, and zomboid-restore.service runs the script as the podman user. The bot gets no ssh, no systemd, and keeps only its existing read-only mount of the Zomboid volume. Two details carry most of the correctness. Resolution is by mtime, not by index. The rotation renames the files -- today's backup_7.zip is backup_8.zip half an hour from now, and a new backup_7.zip holds a different world -- so an index is valid only while the listing is fresh, which is not long enough to survive a human reading a confirmation prompt. The trigger names a set and an mtime; the script resolves the path itself, whitelists the filename, and refuses if nothing matches. It never accepts a path. Everything that can fail is checked before the server is touched. A rotated-out target, an archive with no debbzoid world in it, a bad action, a traversal attempt in the set name: each aborts with the server still running and writes a result file the bot reports back. The live world is moved aside rather than deleted, so a restore is undoable and the last three are kept. One thing PZ does not advertise: its backups do not cover the whole save directory. blam/, a mod's own state, is in none of them -- not the 05:16 archive and not the newest one. Restoring only what the archive holds therefore lands the world slightly *behind* the target rather than on it, so anything present in the displaced world and absent from the archive is carried across. The gregtime tag moves to 3.17.0 for the bot half of this -- `backups`, `restore <n>`, `restore confirm`, `restore undo`, gated to the same two admins as the wipe. That image is built and running on the host; its source is not committed yet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
podman role
Container orchestration for the home server. Containers are defined under
tasks/containers/{base,home,skudak,debyltech}/ and wired from
tasks/main.yml, which is where every image tag is pinned.
Backups
Every Nextcloud-style instance (Nextcloud, Gitea, BookStack, partsy) shares one
backup engine: tasks/containers/cloud-backup.yml plus
templates/nextcloud/cloud-backup.sh.j2. Each instance includes it with its own
vars, producing /usr/local/bin/<name>-backup.sh and a systemd timer.
Stages, in order — the ordering is deliberate, see the comments in the template:
- Database dump inside the container (mariadb / mysql / postgres branches),
gzipped to
/var/backups/nextcloud/<name>/db/. Promoted over yesterday's dump only after passinggzip -tand a completion-trailer grep. - SQLite snapshots (
.backup, thenpragma integrity_check) where used. - rsync of the data tree, config, and db dumps to TrueNAS.
Failures raise status=failed on the nextcloud-backup syslog tag. The mail
itself is sent by the unit's own OnFailure= handler
(templates/nextcloud/nextcloud-backup-alert.sh.j2) straight through
sendmail, so alerting does not depend on Graylog and is unaffected by
graylog_enabled being off. A Graylog rule matching the same tag is a
secondary, dashboard-side copy of that signal.
Do not rename that tag: it is shared by every instance including Gitea, and renaming it here silently stops the Graylog-side alerting for all of them.
Restore
Untested backups are not a control. Rehearse this into a scratch location before you need it, and record the date you last did.
1. Database
Dumps live on the host at /var/backups/nextcloud/<name>/db/<name>-YYYYMMDD.sql.gz
and on TrueNAS at <remote_path>/_backup/db/. TrueNAS in turn cloud-syncs
/mnt/glacier/skudakcloud to Skudak's own iDrive e2 bucket, so a third copy
exists there — but restoring from it means going through the TrueNAS console,
not this host.
Verify the dump before trusting it:
gzip -t <name>-YYYYMMDD.sql.gz
gunzip -c <name>-YYYYMMDD.sql.gz | tail -c 512 # expect the completion trailer
Replay into the running database container. MariaDB/MySQL:
sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \
| podman exec -i <db_container> sh -c \
"exec env MYSQL_PWD=\$MYSQL_ROOT_PASSWORD mariadb -u root \$MYSQL_DATABASE"'
Postgres dumps are taken with --clean --if-exists --no-owner, so they replay
into an existing database:
sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \
| podman exec -i <db_container> sh -c \
"exec env PGPASSWORD=\$POSTGRES_PASSWORD psql -U \$POSTGRES_USER \$POSTGRES_DB"'
2. Data tree
# from TrueNAS
rsync -az -e "ssh -i /etc/ssh/backup_keys/<name>" \
<ssh_user>@truenas.localdomain:<remote_path>/ <volumes>/<instance>/data/
Then fix ownership — the containers run as uid 33 inside a rootless userns:
sudo -H -u podman bash -c 'cd; podman unshare chown -R 33:33 <volumes>/<instance>/data'
3. Reconcile
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ maintenance:mode --on'
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ files:scan --all'
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ maintenance:mode --off'
A DB snapshot slightly older than the files degrades to "files the app has
not indexed yet" and is repaired by files:scan. A DB snapshot newer than
the files references blobs that were never backed up, which surfaces as broken
shares and dead file entries — this is why the dump runs first.
4. LibreSign-specific
The signing CA lives in the data tree at
data/appdata_*/libresign/pki/<instance>_<n>_openssl/, so a data-tree restore
brings it back with everything else. After restoring, confirm it:
sudo -H -u podman bash -c 'cd; podman exec -u www-data skudak-cloud php occ libresign:configure:check'
Every check must report success. If openssl-configure reports an error, the
certificate_engine / config_path app config is pointing somewhere without a
CA — see the guarded generate task in tasks/containers/skudak/cloud.yml.
Do not simply re-run libresign:configure:openssl on a restored instance
without understanding why: it mints a new root CA and invalidates the trust
chain on every document already signed under the old one.
LibreSign
Deployed on skudak-cloud only. LibreSign 14.1.0 requires Nextcloud server
>=34.0.0,<35.0.0, which the pinned nextcloud:34.0.2-apache satisfies. If the
Nextcloud tag is bumped to 35, LibreSign must be held or upgraded in step — the
two instances are pinned independently in tasks/main.yml, so skudak-cloud
can lag cloud if needed.
Dependency split, which drives what survives a container recreate:
| Component | Location | Survives recreate? |
|---|---|---|
| Java (JRE 21), PDFtk, jSignPdf | data/appdata_*/libresign/ |
Yes — persisted volume |
| Root CA / PKI | data/appdata_*/libresign/pki/ |
Yes — persisted volume |
| poppler-utils, ghostscript | /usr in the image |
No — reinstalled by Ansible each run |
Certificate engine is OpenSSL, not CFSSL. CFSSL is the more common source of LibreSign setup failures and buys nothing at this scale.
Gotchas
- Do not pass
--outolibresign:configure:openssl. LibreSign appends its ownlibresign-ca-id:...entry to the OU field, and the combined value overruns the 64-character ASN.1 limit fororganizationalUnitName, failing withstring too long. - A disabled app has no
occcommands. Ifocc list | grep libresignreturns nothing, the app is disabled, not missing —occ app:listwill still show it underDisabled:. This is what a Nextcloud major upgrade does to an app it thinks is incompatible. PHP_MEMORY_LIMITmust be raised above the 512M image default; signing fails opaquely mid-operation otherwise.LC_ALL/LANGmust be set or the JVM comes up asANSI_X3.4-1968and LibreSign warns that accented characters in signer names will be mangled (LibreSign issue #4872).
Logging
Every container runs with log_driver=journald, so container stdout lands in
the host journal, capped at 500M by roles/common/templates/journald-size.conf.j2
(about 25 days at the current rate). journalctl CONTAINER_NAME=<name> is the
day-to-day way to read it.
On top of that sits an optional Graylog stack — graylog, graylog-opensearch,
graylog-mongo, fed by a host fluent-bit service that tails the journal and
ships GELF to 127.0.0.1:12202, enriched with the MaxMind GeoIP database, and
configured over the REST API by the separate graylog-config role.
It is off. The switch is graylog_enabled in
inventories/home/hosts.yml, and it gates all five of those pieces at once:
graylog_enabled |
effect |
|---|---|
true |
stack deployed, fluent-bit shipping, GeoIP downloaded, graylog-config runs, logs.debyl.io proxies to the UI and /gelf |
false |
containers removed and their systemd user units disabled and deleted, fluent-bit stopped and disabled, GeoIP and graylog-config skipped, logs.debyl.io answers 503 |
It was turned off because the cost/benefit is bad on a 4-core box: two JVMs plus MongoDB held ~1.6 GB resident and ~3% of the CPU continuously to store roughly 3k messages a day — about 28 MB of real log data across four live indices — all of which journald already keeps for longer.
Turning it back on is graylog_enabled: true plus make deploy TAGS=graylog.
Nothing is destroyed by the off path: the volumes under {{ graylog_path }}
keep the indices, the Mongo database holding streams/pipelines/dashboards, and
the node-id file, so the stack comes back with its configuration intact.
Two things genuinely stop while it is off, both by design:
- Search and dashboards. The logs still exist in journald; the query interface over them does not.
- External GELF ingest. The AWS Lambda that POSTs to
logs.debyl.io/gelffor thedebyltech-apistream gets a 503. Those events are dropped, not queued — nothing else records them.
Caddy config reloads
The reload caddy handler deliberately reads /config/Caddyfile, not the
/etc/caddy/Caddyfile the container starts from, even though both are the same
host file. /etc/caddy/Caddyfile is a single-file bind mount, which podman
binds by inode, and Ansible's template module writes a temp file and renames
it into place — so every deploy gives the host file a new inode while the
container keeps seeing the one it was created with. Reloading from that path
silently re-applied the previous config; changes only landed when something
recreated the container. {{ caddy_path }}/config is also bind-mounted as a
directory at /config, and directory mounts resolve names at open() time,
so /config/Caddyfile is always the file Ansible just wrote.