# podman role Container orchestration for the home server. Containers are defined under `tasks/containers/{base,home,skudak,debyltech}/` and wired from `tasks/main.yml`, which is where every image tag is pinned. ## Backups Every Nextcloud-style instance (Nextcloud, Gitea, BookStack, partsy) shares one backup engine: `tasks/containers/cloud-backup.yml` plus `templates/nextcloud/cloud-backup.sh.j2`. Each instance includes it with its own vars, producing `/usr/local/bin/-backup.sh` and a systemd timer. Stages, in order — the ordering is deliberate, see the comments in the template: 1. **Database dump** inside the container (mariadb / mysql / postgres branches), gzipped to `/var/backups/nextcloud//db/`. Promoted over yesterday's dump only after passing `gzip -t` **and** a completion-trailer grep. 2. **SQLite snapshots** (`.backup`, then `pragma integrity_check`) where used. 3. **rsync** of the data tree, config, and db dumps to TrueNAS. Failures raise `status=failed` on the `nextcloud-backup` syslog tag. The mail itself is sent by the unit's own `OnFailure=` handler (`templates/nextcloud/nextcloud-backup-alert.sh.j2`) straight through `sendmail`, so alerting does **not** depend on Graylog and is unaffected by `graylog_enabled` being off. A Graylog rule matching the same tag is a secondary, dashboard-side copy of that signal. Do not rename that tag: it is shared by every instance including Gitea, and renaming it here silently stops the Graylog-side alerting for all of them. ## Restore **Untested backups are not a control.** Rehearse this into a scratch location before you need it, and record the date you last did. ### 1. Database Dumps live on the host at `/var/backups/nextcloud//db/-YYYYMMDD.sql.gz` and on TrueNAS at `/_backup/db/`. TrueNAS in turn cloud-syncs each dataset offsite, so a third copy exists there — but restoring from it means going through the TrueNAS console, not this host: | Dataset | Offsite | |---|---| | `skudakcloud`, `skudakapps`, `skudakgit` | Skudak's own iDrive e2 bucket (excluded from the personal task) | | `nextcloud`, `gitea`, `debyltechcloud` | Personal iDrive e2 bucket, via the "iDrive E2 Backup" task over `/mnt/glacier` | Verify the dump before trusting it: ```bash gzip -t -YYYYMMDD.sql.gz gunzip -c -YYYYMMDD.sql.gz | tail -c 512 # expect the completion trailer ``` Replay into the running database container. MariaDB/MySQL: ```bash sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \ | podman exec -i sh -c \ "exec env MYSQL_PWD=\$MYSQL_ROOT_PASSWORD mariadb -u root \$MYSQL_DATABASE"' ``` Postgres dumps are taken with `--clean --if-exists --no-owner`, so they replay into an existing database: ```bash sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \ | podman exec -i sh -c \ "exec env PGPASSWORD=\$POSTGRES_PASSWORD psql -U \$POSTGRES_USER \$POSTGRES_DB"' ``` ### 2. Data tree ```bash # from TrueNAS rsync -az -e "ssh -i /etc/ssh/backup_keys/" \ @truenas.localdomain:/ //data/ ``` Then fix ownership — the containers run as uid 33 inside a rootless userns: ```bash sudo -H -u podman bash -c 'cd; podman unshare chown -R 33:33 //data' ``` ### 3. Reconcile ```bash sudo -H -u podman bash -c 'cd; podman exec -u www-data php occ maintenance:mode --on' sudo -H -u podman bash -c 'cd; podman exec -u www-data php occ files:scan --all' sudo -H -u podman bash -c 'cd; podman exec -u www-data php occ maintenance:mode --off' ``` A DB snapshot slightly **older** than the files degrades to "files the app has not indexed yet" and is repaired by `files:scan`. A DB snapshot **newer** than the files references blobs that were never backed up, which surfaces as broken shares and dead file entries — this is why the dump runs first. ### 4. LibreSign-specific The signing CA lives in the data tree at `data/appdata_*/libresign/pki/__openssl/`, so a data-tree restore brings it back with everything else. After restoring, confirm it: ```bash sudo -H -u podman bash -c 'cd; podman exec -u www-data php occ libresign:configure:check' ``` Every check must report `success`. If `openssl-configure` reports an error, the `certificate_engine` / `config_path` app config is pointing somewhere without a CA — see the guarded generate task in `tasks/containers/{skudak,debyltech}/cloud.yml`. **Do not** simply re-run `libresign:configure:openssl` on a restored instance without understanding why: it mints a *new* root CA and invalidates the trust chain on every document already signed under the old one. ## LibreSign Deployed on `skudak-cloud` and `debyltech-cloud`. LibreSign 14.1.0 requires Nextcloud server `>=34.0.0,<35.0.0`, which the pinned `nextcloud:34.0.3-apache` satisfies. If the Nextcloud tag is bumped to 35, LibreSign must be held or upgraded in step — each instance is pinned independently in `tasks/main.yml`, so the LibreSign instances can lag `cloud` if needed. The branding apps (`files/skudakmail`, `files/debyltechmail`) pin `max-version="34"` too and must be bumped alongside. Dependency split, which drives what survives a container recreate: | Component | Location | Survives recreate? | |---|---|---| | Java (JRE 21), PDFtk, jSignPdf | `data/appdata_*/libresign/` | **Yes** — persisted volume | | Root CA / PKI | `data/appdata_*/libresign/pki/` | **Yes** — persisted volume | | poppler-utils, ghostscript | `/usr` in the image | **No** — reinstalled by Ansible each run | Certificate engine is **OpenSSL**, not CFSSL. CFSSL is the more common source of LibreSign setup failures and buys nothing at this scale. ### Gotchas - **Do not pass `--ou`** to `libresign:configure:openssl`. LibreSign appends its own `libresign-ca-id:...` entry to the OU field, and the combined value overruns the 64-character ASN.1 limit for `organizationalUnitName`, failing with `string too long`. - **A disabled app has no `occ` commands.** If `occ list | grep libresign` returns nothing, the app is disabled, not missing — `occ app:list` will still show it under `Disabled:`. This is what a Nextcloud major upgrade does to an app it thinks is incompatible. - `PHP_MEMORY_LIMIT` must be raised above the 512M image default; signing fails opaquely mid-operation otherwise. - `LC_ALL` / `LANG` must be set or the JVM comes up as `ANSI_X3.4-1968` and LibreSign warns that accented characters in signer names will be mangled (LibreSign issue #4872). ## cloud.debyltech.com accounts Registration is off, so every account is created by an admin in Settings → Accounts. The isolation policy in `tasks/containers/debyltech/cloud.yml` is enforced on every deploy and checked by the verify script. **Staff**: add to the `admin` group (the only group allowed to share) and to `cloud_debyltech_staff_users` in `defaults/main.yml`, so the next deploy exempts them from the 0 B default quota. Until then, set the account's quota to *Unlimited* by hand. **Customer**: 1. Create a group per customer, e.g. `Customer - Acme`. Use a consistent prefix: autocomplete is off, so you share by typing the **exact** group name. 2. Create the account with only their email filled in and no password, in that group. Nextcloud sends a branded welcome email with a set-password link. 3. Share a folder (for example `Customers/Acme`) with the group. It defaults to **View only**; choose *Allow editing* only if they should upload. What a customer gets: - read-only access to what's shared with them - no personal storage (0 B quota) - no sharing or public links - no view of other accounts or groups: the share search and contacts menu return nothing Signature requests to customers need no account: use LibreSign with their email address. ## Logging Every container runs with `log_driver=journald`, so container stdout lands in the host journal, capped at 500M by `roles/common/templates/journald-size.conf.j2` (about 25 days at the current rate). `journalctl CONTAINER_NAME=` is the day-to-day way to read it. On top of that sits an optional Graylog stack — `graylog`, `graylog-opensearch`, `graylog-mongo`, fed by a host `fluent-bit` service that tails the journal and ships GELF to `127.0.0.1:12202`, enriched with the MaxMind GeoIP database, and configured over the REST API by the separate `graylog-config` role. **It is off.** The switch is `graylog_enabled` in `inventories/home/hosts.yml`, and it gates all five of those pieces at once: | `graylog_enabled` | effect | | --- | --- | | `true` | stack deployed, fluent-bit shipping, GeoIP downloaded, `graylog-config` runs, `logs.debyl.io` proxies to the UI and `/gelf` | | `false` | containers removed and their systemd user units disabled and deleted, fluent-bit stopped and disabled, GeoIP and `graylog-config` skipped, `logs.debyl.io` answers 503 | It was turned off because the cost/benefit is bad on a 4-core box: two JVMs plus MongoDB held ~1.6 GB resident and ~3% of the CPU continuously to store roughly 3k messages a day — about 28 MB of real log data across four live indices — all of which journald already keeps for longer. Turning it back on is `graylog_enabled: true` plus `make deploy TAGS=graylog`. Nothing is destroyed by the off path: the volumes under `{{ graylog_path }}` keep the indices, the Mongo database holding streams/pipelines/dashboards, and the node-id file, so the stack comes back with its configuration intact. Two things genuinely stop while it is off, both by design: - **Search and dashboards.** The logs still exist in journald; the query interface over them does not. - **External GELF ingest.** The AWS Lambda that POSTs to `logs.debyl.io/gelf` for the `debyltech-api` stream gets a 503. Those events are dropped, not queued — nothing else records them. ### Caddy config reloads The `reload caddy` handler deliberately reads `/config/Caddyfile`, not the `/etc/caddy/Caddyfile` the container starts from, even though both are the same host file. `/etc/caddy/Caddyfile` is a **single-file** bind mount, which podman binds by inode, and Ansible's `template` module writes a temp file and renames it into place — so every deploy gives the host file a new inode while the container keeps seeing the one it was created with. Reloading from that path silently re-applied the previous config; changes only landed when something recreated the container. `{{ caddy_path }}/config` is also bind-mounted as a *directory* at `/config`, and directory mounts resolve names at `open()` time, so `/config/Caddyfile` is always the file Ansible just wrote.