Compare commits

...

13 Commits

Author SHA1 Message Date
Bastian de Byl 773a2bbc9c harden TrueNAS CIFS mounts so immich self-heals
TrueNAS was power-cycled, the CIFS mounts failed, and systemd never retried
-- mount units are not restarted on failure. SMB came back, nothing
remounted, and immich-server served an empty library for days while its
database still listed 14,181 assets pointing at /mnt/media/originals.

fstab carried no _netdev, no nofail and no automount, so there was no path
back without a human. Now:

  x-systemd.automount  any access re-attempts the mount; failure stops
                       being terminal
  _netdev / nofail     ordered after network-online, dead NAS cannot block boot
  soft                 I/O errors instead of blocking forever, so the
                       container can be restarted rather than wedging in
                       uninterruptible sleep
  idle-timeout         unmount when unused, clearing stale handles
  resilienthandles     SMB3 rides out brief blips

ansible.posix.mount mounts directly and never starts the generated
.automount unit, leaving the on-access trigger inactive -- enable it
explicitly, or the headline fix silently does nothing.

The containers are systemd USER units while the mounts are SYSTEM units, so
RequiresMountsFor= is unavailable. cifs-watchdog bridges the scopes: checks
health, recovers, and restarts ONLY immich-server (the sole consumer of both
paths; postgres/redis/ML use local volumes).

Two bugs the umount test caught, both worth knowing:
- `ls` cannot test mountedness. An unmounted mount point is an ordinary
  empty directory, so ls succeeds and recovery was skipped entirely.
- A drop repaired within a single run leaves prev=healthy, so keying the
  restart solely on the stored state skipped it while the container still
  held its stale view.

Also moves the SMB password out of /etc/fstab, which is 0644 and was
readable by every local user, into a 0600 credentials file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 14:49:22 -04:00
Bastian de Byl 0939204061 docs(podman): add role README with backup, restore and LibreSign runbook
No restore procedure existed anywhere in this repo. The backup pipeline
is well commented but nothing described how to get data back, and the
README the backup script already referenced was missing.

Covers the shared backup engine and its stage ordering, a restore
procedure for both MariaDB and Postgres with the rootless-podman command
form, data-tree restore including the uid 33 chown, and the
maintenance-mode/files:scan reconcile.

Two things worth stating plainly:
- Untested backups are not a control. No rehearsal has been recorded.
- Do not blindly re-run libresign:configure:openssl on a restored
  instance -- it mints a new root CA and invalidates the trust chain on
  every document already signed under the old one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:55:04 -04:00
Bastian de Byl fec7d62acb feat(skudak-cloud): repair LibreSign, brand its mail, add Redis
LibreSign had been silently broken since it was first deployed in
January. Every step of the old before-starting hook ended in `|| echo`,
so six months of failures logged nothing.

LibreSign repair
- Root cause was a stale config_path: a valid OpenSSL root CA existed at
  generation 1, a failed CFSSL attempt left an empty generation 2, and
  config_path was left pointing at the empty one. Regenerated as
  "Skudak LLP" (was the pre-rename "Skudak Rennsport LLP").
- Deleted the hook. Java/PDFtk/jSignPdf live under data/appdata_*, a
  persisted volume, so they only ever needed installing once. Install and
  verification are now explicit tasks that actually fail.
- PHP_MEMORY_LIMIT 1024M -- the 512M image default fails opaquely
  mid-signature. LC_ALL/LANG so the JVM is not ANSI_X3.4-1968.
- signature_render_mode=GRAPHIC_ONLY. Any other mode halves the stamp
  width and overlays a name/date block that collides with the drawn mark
  and duplicates what our documents already typeset. The value must be
  exactly GRAPHIC_ONLY; a bare "GRAPHIC" is accepted by occ, matches no
  radio in the UI, and silently reverts to default.
- write_qrcode_on_footer=false, written with --type=boolean because
  FooterHandler reads it via getValueBool and the typed appconfig API
  does not coerce a string "0". The validation URL text is kept.
- identification_documents=0 -- the default gates signing behind an ID
  upload plus admin approval, so signers saw no way to sign.
- shareapi_restrict_user_enumeration_full_match=no, so an email owned by
  an existing account can be added as a signer. Root cause is in core
  (MailPlugin.php:163), not LibreSign. Do NOT set full_match_email=no --
  that disables email signer search entirely.

Mail branding (skudakmail app)
- Two supported extension points, no core patch and no LibreSign fork:
  mail_template_class for layout, subjects, button labels and the footer
  LibreSign never adds; and a BeforeMessageSent listener to embed the
  wordmark as a cid: part so it survives remote-image blocking.
- A third listener adds scoped CSS fixing the signing page being clipped
  on iOS Safari (100vh -> 100dvh). Patched upstream too.
- skudakmail-verify.php.j2 asserts all of the above through the real
  useTemplate() path and fails the play on drift. Every assertion was
  proven to fail when deliberately regressed.

Redis
- memcache.locking was unset, so Nextcloud used DBLockingProvider and
  every file lock became a MariaDB write -- the contention behind the
  intermittent multi-second stalls. Verified after: db locks static,
  redis keys growing.
- requirepass lives in a mounted 0640 conf, not --requirepass, which
  would leak it into podman inspect, the systemd unit and ps. The file is
  chowned to uid 999 because redis-server does not run as root and the
  :ro mount stops the image fixing it itself.
- No maxmemory: cache is evictable, locks are NOT, and evicting a held
  lock permits concurrent writers to one file. No persistence either --
  a restored RDB could reinstate locks whose owner is long dead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:54:52 -04:00
Bastian de Byl c184099b2d refactor(backup): remove duplicate direct-to-S3 stage
A host-to-iDrive S3 stage was added here and is now removed. It would
have written the same data into the same `backup-all` bucket that the
TrueNAS cloud-sync task already fills -- duplicate storage, two writers
to one prefix, for no additional coverage.

Offsite to business-owned storage was already solved: the rsync feeds
/mnt/glacier/skudakcloud and TrueNAS cloud-syncs that to Skudak's own
iDrive e2 account. The earlier note in skudak/cloud.yml proposed adding
S3 *and then dropping the rsync* -- replacement, not addition -- and
building both was a misreading of it.

If offsite is ever moved onto this host it must REPLACE the rsync, not
run beside it. Settle first whether the TrueNAS -> iDrive leg is
independently verifiable; keeping this chain means trusting it.

Also removes the now-orphaned /etc/backup_s3 credential file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:54:27 -04:00
Bastian de Byl f11391b28f bound log growth and reclaim ~52 GB of container disk
Caddy was rotating on implicit defaults (100MiB/keep 10/90d) that were not
holding -- 20 rotated files per stream and a 190-day-old .gz, 1.3 GB across
16 log streams. Made explicit at 10MiB/keep 3/7d.

Note roll_size et al are subdirectives of `output file`, NOT of `log`.
Getting that wrong does not degrade gracefully: Caddy refuses to start on a
bad config, so every site went down until it was corrected. Worth a
`caddy validate` gate before reload.

journald had no SystemMaxUse and had reached 4 GB, drifting toward its
10%-of-filesystem default (~190 GB on this root). Capped at 500M.

Both are safe to keep short because fluent-bit ships the journal and every
Caddy access log into Graylog -- though note its GELF output has been
erroring for days, which weakens that premise and wants investigating.

The larger find was unrelated to logs: 896 images totalling 59.6 GB with
75% unused (94 tags of greg-time-bot, 73 of fulfillr -- one per deploy) and
5.4 GB of dangling volumes, mostly 804 MB Nextcloud /var/www/html trees
orphaned by container recreations. Pruned to 22 images / 15.4 GB, and added
a weekly timer keeping 30 days so a rollback still needs no rebuild.

Also dropped the decommissioned 6379/tcp redis rule (nothing listening;
Immich's redis is on the shared podman network) and the orphaned nosql, s3
and searxng volume dirs.

Backup log exclusions turned out to be unnecessary: Gitea logs to console
so its log dirs are empty, Nextcloud already excludes its own, BookStack
mounts only uploads, and Caddy is not backed up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:13:35 -04:00
Bastian de Byl bc110ce69e back up Gitea + Skudak app data; drop PartKeepr and Pi-hole
Extends the Nextcloud backup machinery rather than adding a second
mechanism. cloud-backup.sh.j2 gains three guarded options, all no-ops for
the existing callers:

  backup_podman_user  Gitea runs rootless under `git`, not `podman`
  backup_db_type      postgres (Gitea) and mysql (BookStack) alongside
                      mariadb; each engine's completion trailer differs,
                      and grepping for the wrong one fails every run
  backup_sqlite_dbs   `sqlite3 .backup` for live WAL-mode SQLite, gated on
                      `pragma integrity_check` before promotion -- rsync
                      is either stale (no -wal) or torn (with it)

New instances: gitea-debyl, skudak-gitea, bookstack, partsy-skudak. The
alert handler is rendered once and shared, so its wording is now generic
rather than per-product; TAG stays nextcloud-backup because an external
Graylog rule matches on it.

`apply:` on the includes is load-bearing -- tags on a dynamic
include_tasks do not reach the tasks inside it.

Business data (skudak-gitea, bookstack, partsy-skudak) goes to TrueNAS
and on to Skudak's own iDrive account; the personal bucket's
/skudak*/** excludes are permanent, not a stopgap.

Removals: PartKeepr is superseded by Partsy, and its teardown never
finished -- it targeted /etc/systemd/system/podman-partkeepr*.service,
wrong prefix and wrong scope, leaving enabled user units in failed state.
Pi-hole's role was already orphaned (absent from deploy_home.yml); its
port 53 rule went with it after confirming nothing listens there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 18:08:47 -04:00
bastian a63bf5edec Backup hardening, UPS monitoring, Nextcloud cron, image bumps 2026-07-30 13:05:20 -04:00
Bastian de Byl 5776dbe1bf UPS monitoring, Nextcloud cron, container image bumps
ups role: NUT server on home.debyl.io for the CyberPower PR1500RT2U that backs
both it and truenas.localdomain, with staged shutdown (TrueNAS sheds at t+2min,
host at 10% charge) and best-effort IPMI power-on when mains returns. The 10%
threshold leans on ignorelb + override.battery.charge.low rather than a custom
poller, because CyberPower asserts its own low-battery flag far too early.
Credentials come from vault vars; nothing sensitive is templated in the clear.

Nextcloud background jobs: both instances have backgroundjobs_mode "cron", which
expects an external caller every ~5 minutes, and nothing was calling. The
personal instance had not run a background job since 2026-05-14 and skudak since
2024-11-20. Consequently trash and file versions never expired, stale chunked
uploads accumulated, calendar reminders never fired, and nextcloud.log was never
rotated -- which quietly made the existing log_rotate_size cap inert. Added a
systemd timer per instance, skipping cleanly when the container is down or in
maintenance so deploy windows don't show up as failed units.

Trash retention on the personal instance: the default "auto" only expires when
disk space demands it, so 66 GB of >30-day deletions sat on a host with 1.3 TB
free -- effectively unbounded. "auto, 30" makes the 30-day expiry unconditional
while still purging early under pressure.

Image bumps:
  nextcloud       33.0.0  -> 34.0.2   (both cloud and skudak-cloud)
  greg-time-bot   3.9.25  -> 3.10.0
  fulfillr        20260723.2044 -> 20260728.2155  (prod and dev)

The fulfillr bump records what is already deployed: both containers were rolled
to that image on 2026-07-29 for SCRUM-156 (digital product releases + customer
update campaign). Committing it keeps the repo from claiming an older tag than
the host is actually running, which would otherwise roll fulfillr backwards on
the next clean-checkout deploy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 13:00:51 -04:00
Bastian de Byl 16145fb6bd only email genuine backup failures
A spurious invocation carries no action for a reader, so mailing it just
trains them to skip past the subject line -- which defeats the point of the
alert. Send mail only when the unit actually reports failure; spurious
triggers still leave their journald record, so they stay greppable and can
still feed a Graylog rule.

Verified both paths: a spurious trigger leaves /var/log/msmtp.log untouched
and logs alert_mail=skipped reason=spurious, while a genuinely failed unit
still composes mail with the FAILED subject.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:34:45 -04:00
Bastian de Byl 6e99794d0f restore full dump flags; stop spurious alerts claiming FAILED
mariadb-upgrade --force has now been run against both instances and repaired
the system tables, so --routines and --events no longer abort the dump.
Restore them for completeness. Both verified: rc=0 with a clean completion
trailer, and both Nextcloud instances report installed with unchanged table
counts afterwards.

The upgrade still exits non-zero on these containers because it cannot create
the `sys` schema: /var/lib/mysql is owned by daemon rather than mysql, so
mysqld may not create top-level databases. `sys` is diagnostic only and
unused by Nextcloud, but the same permission would block creating any new
database, so it is recorded in the template comment.

The alert handler claimed FAILED in its subject line regardless of what the
unit actually reported, so starting it by hand mailed out a failure notice
for a run that succeeded. Derive the subject and log line from the real
Result, and emit status=spurious rather than status=failed so such triggers
cannot match a Graylog alert rule keyed on genuine failures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:04:32 -04:00
Bastian de Byl 7c72aee4f9 add home_smtp vault key for system mail
Referenced by roles/common/templates/msmtp/msmtprc.j2 to authenticate
outbound alert mail against the OpenSRS relay. Without this committed a
fresh checkout cannot render /etc/msmtprc.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 17:22:30 -04:00
Bastian de Byl 0ab423ca55 harden nextcloud backups: db dumps, alerting, drift fix
The data-only rsync left no way to restore a working instance: mysql/ and
config/ were never backed up, so a recovery would have files but no shares,
users or metadata. Dump the database before syncing files (a DB older than
the files is repairable with occ files:scan; a newer one references blobs
that never made it into the backup) and ship config/ alongside it.

Capture the --chmod=Du=rwx,Dgo=rx flag that had been hand-added to the
deployed skudak-cloud script. It was outside git, so every deploy silently
reverted it. It now lives in backup_rsync_extra_args.

Add OnFailure= alerting. The units failed silently before, which is how an
iDrive sync failure sat unnoticed since May. msmtp rather than the esmtp
already installed: the OpenSRS relay is port 465 (implicit TLS) and libesmtp
only speaks STARTTLS.

Exclude nextcloud.log* from the sync and cap log_rotate_size. skudak-cloud
was running at loglevel 0 and had written a 64 GB log that was being rsynced
and pushed to S3; set it to 2 to match the home instance.

Stagger the timers (04:00 / 04:30) so both finish before the 05:00 TrueNAS
snapshot task, and bound TimeoutStartSec so a wedged rsync cannot leave the
unit activating forever and skip every subsequent trigger.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 16:03:18 -04:00
Bastian de Byl 1e1d53ecd8 feat(gitea-actions): install jq in the ESP-IDF CI image
esp32-stm32-vcu's scripts/release.sh uses jq to rewrite the protocol
manifest before publishing it to S3. The image did not have it, so the
release aborted at:

    ./scripts/release.sh: line 61: jq: command not found

The failure mode is nastier than a red job. jq is only reached *after*
the firmware .bin and version.json have already been uploaded, so
clients were served the new build while the git tag, the Gitea release
and the protocol manifest were never written — the repo still showed
the previous release as latest while production served a newer one.

Only jq was missing; aws and sha256sum are already present (verified in
the rebuilt image: jq-1.7, aws-ok, sha256sum-ok).

This surfaced when the skudak firmware jobs were moved into this image
(they previously ran on the default gitea-ci image, which has jq).
2026-07-25 12:01:16 -04:00
70 changed files with 3352 additions and 175 deletions
+2
View File
@@ -11,3 +11,5 @@
- role: github-actions
- role: graylog-config
tags: graylog-config
- role: ups
tags: ups
+6
View File
@@ -16,3 +16,9 @@
ansible.builtin.systemd:
name: fluent-bit
state: restarted
- name: restart_journald
become: true
ansible.builtin.systemd:
name: systemd-journald
state: restarted
+34
View File
@@ -0,0 +1,34 @@
---
# Outbound mail for system notifications (backup failures, cron output).
#
# msmtp rather than the esmtp already on the box: the OpenSRS relay uses port
# 465, which is IMPLICIT TLS, and libesmtp/esmtp only speaks STARTTLS. msmtp
# handles implicit TLS via `tls_starttls off` and mirrors the settings the
# TrueNAS box already uses.
# On Fedora the single `msmtp` package already ships /usr/bin/sendmail and
# /usr/lib/sendmail; there is no separate msmtp-sendmail package (Debian split).
- name: install msmtp
become: true
ansible.builtin.package:
name: msmtp
state: present
tags: mail
- name: configure msmtp
become: true
ansible.builtin.template:
src: msmtp/msmtprc.j2
dest: /etc/msmtprc
owner: root
group: root
mode: 0600
no_log: true
tags: mail
- name: point the mta alternative at msmtp
become: true
community.general.alternatives:
name: mta
path: /usr/bin/msmtp-sendmail
failed_when: false
tags: mail
+3
View File
@@ -3,6 +3,9 @@
- import_tasks: security.yml
- import_tasks: service.yml
- import_tasks: mail.yml
tags: mail
- import_tasks: fluent-bit.yml
tags: fluent-bit, graylog
+24
View File
@@ -1,4 +1,28 @@
---
# Bound the journal. Container stdout all lands here (log_driver=journald) and
# is shipped to Graylog by fluent-bit, so the local journal only needs to be a
# buffer -- see the template for the reasoning behind the size.
- name: ensure journald drop-in directory exists
become: true
ansible.builtin.file:
path: /etc/systemd/journald.conf.d
state: directory
owner: root
group: root
mode: 0755
tags: security, service, journald
- name: cap journald disk usage
become: true
ansible.builtin.template:
src: journald-size.conf.j2
dest: /etc/systemd/journald.conf.d/99-size.conf
owner: root
group: root
mode: 0644
notify: restart_journald
tags: security, service, journald
- name: ensure desired services are started and enabled
become: true
ansible.builtin.service:
@@ -0,0 +1,12 @@
### {{ ansible_managed }}
# Without a cap journald sizes itself at 10% of the filesystem, which on this
# 1.9 TB root means it will happily grow into the tens of gigabytes; it had
# reached 4 GB before this was set.
#
# Kept small on purpose: every container runs with log_driver=journald and
# fluent-bit drains the journal into Graylog continuously (systemd input,
# _COMM=conmon -- see templates/fluent-bit/fluent-bit.conf.j2), so Graylog is
# the system of record. What stays here is only the buffer that covers
# fluent-bit being down, and 500M is a long outage at this log rate.
[Journal]
SystemMaxUse={{ journald_max_use | default('500M') }}
@@ -0,0 +1,17 @@
# {{ ansible_managed }}
defaults
auth on
tls on
tls_trust_file /etc/pki/tls/certs/ca-bundle.crt
logfile /var/log/msmtp.log
account home
host {{ system_smtp_host | default('mail.b.hostedemail.com') }}
port {{ system_smtp_port | default(465) }}
# Port 465 is implicit TLS (SMTPS), not STARTTLS.
tls_starttls off
from {{ system_smtp_from | default('home@bdebyl.net') }}
user {{ system_smtp_user | default('home@bdebyl.net') }}
password {{ home_smtp }}
account default : home
@@ -1,16 +1,23 @@
# ESP-IDF firmware job image (managed by ansible: roles/gitea-actions).
# Adds node (required by actions/checkout and other JS actions), the AWS CLI
# (firmware artifacts ship to S3), and the common-yaml header generator's Python
# deps on top of the official Espressif toolchain.
# (firmware artifacts ship to S3), jq (the release script rewrites the protocol
# manifest with it), and the common-yaml header generator's Python deps on top
# of the official Espressif toolchain.
# IDF lives at /opt/esp/idf — firmware jobs source /opt/esp/idf/export.sh.
# python3-yaml + python3-jinja2 are installed as distro packages so the
# common-yaml generator runs with a plain `python3 generate.py` — no pip at job
# time (the base image's system Python is PEP 668 externally-managed) and no
# need to source the IDF venv just to generate headers.
#
# jq is required by esp32-stm32-vcu scripts/release.sh, which publishes the
# protocol manifest and BLE_API.md to S3 after the firmware upload. Without it
# the release aborts *after* the firmware and version.json are already live —
# clients get the new build while the tag, Gitea release and protocol manifest
# are never written. Keep it installed.
FROM espressif/idf:{{ esp_idf_version }}
RUN apt-get update && apt-get install -y --no-install-recommends \
curl ca-certificates unzip python3-yaml python3-jinja2 \
curl ca-certificates unzip jq python3-yaml python3-jinja2 \
&& curl -fsSL https://deb.nodesource.com/setup_20.x | bash - \
&& apt-get install -y --no-install-recommends nodejs \
&& rm -rf /var/lib/apt/lists/*
-5
View File
@@ -1,5 +0,0 @@
---
deps: [
php-sqlite,
php-fpm
]
-3
View File
@@ -1,3 +0,0 @@
---
dependencies:
- role: http
-11
View File
@@ -1,11 +0,0 @@
---
- name: install pi-hole-server
ansible.builtin.command: yay -S --noconfirm pi-hole-server
args:
creates: /bin/pihole
- name: install pi-hole-server dependencies
become: true
ansible.builtin.package:
name: "{{ deps }}"
state: present
-3
View File
@@ -1,3 +0,0 @@
---
- import_tasks: deps.yml
- import_tasks: php.yml
-12
View File
@@ -1,12 +0,0 @@
---
- name: replace pi.hole hostname
become: true
ansible.builtin.replace:
path: "{{ item }}"
regexp: "pi\\.hole"
replace: "pi.bdebyl.net"
loop:
- /srv/http/pihole/admin/scripts/pi-hole/php/auth.php
- /srv/http/pihole/pihole/index.php
tags:
- pihole
+143
View File
@@ -0,0 +1,143 @@
# podman role
Container orchestration for the home server. Containers are defined under
`tasks/containers/{base,home,skudak,debyltech}/` and wired from
`tasks/main.yml`, which is where every image tag is pinned.
## Backups
Every Nextcloud-style instance (Nextcloud, Gitea, BookStack, partsy) shares one
backup engine: `tasks/containers/cloud-backup.yml` plus
`templates/nextcloud/cloud-backup.sh.j2`. Each instance includes it with its own
vars, producing `/usr/local/bin/<name>-backup.sh` and a systemd timer.
Stages, in order — the ordering is deliberate, see the comments in the template:
1. **Database dump** inside the container (mariadb / mysql / postgres branches),
gzipped to `/var/backups/nextcloud/<name>/db/`. Promoted over yesterday's dump
only after passing `gzip -t` **and** a completion-trailer grep.
2. **SQLite snapshots** (`.backup`, then `pragma integrity_check`) where used.
3. **rsync** of the data tree, config, and db dumps to TrueNAS.
Failures raise `status=failed` on the `nextcloud-backup` syslog tag, which an
external Graylog rule matches to send mail. Do not rename that tag: it is shared
by every instance including Gitea, and renaming it here silently stops alerting
for all of them.
## Restore
**Untested backups are not a control.** Rehearse this into a scratch location
before you need it, and record the date you last did.
### 1. Database
Dumps live on the host at `/var/backups/nextcloud/<name>/db/<name>-YYYYMMDD.sql.gz`
and on TrueNAS at `<remote_path>/_backup/db/`. TrueNAS in turn cloud-syncs
`/mnt/glacier/skudakcloud` to Skudak's own iDrive e2 bucket, so a third copy
exists there — but restoring from it means going through the TrueNAS console,
not this host.
Verify the dump before trusting it:
```bash
gzip -t <name>-YYYYMMDD.sql.gz
gunzip -c <name>-YYYYMMDD.sql.gz | tail -c 512 # expect the completion trailer
```
Replay into the running database container. MariaDB/MySQL:
```bash
sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \
| podman exec -i <db_container> sh -c \
"exec env MYSQL_PWD=\$MYSQL_ROOT_PASSWORD mariadb -u root \$MYSQL_DATABASE"'
```
Postgres dumps are taken with `--clean --if-exists --no-owner`, so they replay
into an existing database:
```bash
sudo -H -u podman bash -c 'cd; gunzip -c /path/to/dump.sql.gz \
| podman exec -i <db_container> sh -c \
"exec env PGPASSWORD=\$POSTGRES_PASSWORD psql -U \$POSTGRES_USER \$POSTGRES_DB"'
```
### 2. Data tree
```bash
# from TrueNAS
rsync -az -e "ssh -i /etc/ssh/backup_keys/<name>" \
<ssh_user>@truenas.localdomain:<remote_path>/ <volumes>/<instance>/data/
```
Then fix ownership — the containers run as uid 33 inside a rootless userns:
```bash
sudo -H -u podman bash -c 'cd; podman unshare chown -R 33:33 <volumes>/<instance>/data'
```
### 3. Reconcile
```bash
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ maintenance:mode --on'
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ files:scan --all'
sudo -H -u podman bash -c 'cd; podman exec -u www-data <container> php occ maintenance:mode --off'
```
A DB snapshot slightly **older** than the files degrades to "files the app has
not indexed yet" and is repaired by `files:scan`. A DB snapshot **newer** than
the files references blobs that were never backed up, which surfaces as broken
shares and dead file entries — this is why the dump runs first.
### 4. LibreSign-specific
The signing CA lives in the data tree at
`data/appdata_*/libresign/pki/<instance>_<n>_openssl/`, so a data-tree restore
brings it back with everything else. After restoring, confirm it:
```bash
sudo -H -u podman bash -c 'cd; podman exec -u www-data skudak-cloud php occ libresign:configure:check'
```
Every check must report `success`. If `openssl-configure` reports an error, the
`certificate_engine` / `config_path` app config is pointing somewhere without a
CA — see the guarded generate task in `tasks/containers/skudak/cloud.yml`.
**Do not** simply re-run `libresign:configure:openssl` on a restored instance
without understanding why: it mints a *new* root CA and invalidates the trust
chain on every document already signed under the old one.
## LibreSign
Deployed on `skudak-cloud` only. LibreSign 14.1.0 requires Nextcloud server
`>=34.0.0,<35.0.0`, which the pinned `nextcloud:34.0.2-apache` satisfies. If the
Nextcloud tag is bumped to 35, LibreSign must be held or upgraded in step — the
two instances are pinned independently in `tasks/main.yml`, so `skudak-cloud`
can lag `cloud` if needed.
Dependency split, which drives what survives a container recreate:
| Component | Location | Survives recreate? |
|---|---|---|
| Java (JRE 21), PDFtk, jSignPdf | `data/appdata_*/libresign/` | **Yes** — persisted volume |
| Root CA / PKI | `data/appdata_*/libresign/pki/` | **Yes** — persisted volume |
| poppler-utils, ghostscript | `/usr` in the image | **No** — reinstalled by Ansible each run |
Certificate engine is **OpenSSL**, not CFSSL. CFSSL is the more common source of
LibreSign setup failures and buys nothing at this scale.
### Gotchas
- **Do not pass `--ou`** to `libresign:configure:openssl`. LibreSign appends its
own `libresign-ca-id:...` entry to the OU field, and the combined value
overruns the 64-character ASN.1 limit for `organizationalUnitName`, failing
with `string too long`.
- **A disabled app has no `occ` commands.** If `occ list | grep libresign`
returns nothing, the app is disabled, not missing — `occ app:list` will still
show it under `Disabled:`. This is what a Nextcloud major upgrade does to an
app it thinks is incompatible.
- `PHP_MEMORY_LIMIT` must be raised above the 512M image default; signing fails
opaquely mid-operation otherwise.
- `LC_ALL` / `LANG` must be set or the JVM comes up as `ANSI_X3.4-1968` and
LibreSign warns that accented characters in signer names will be mangled
(LibreSign issue #4872).
+51 -2
View File
@@ -1,4 +1,6 @@
---
# Where Nextcloud backup failure alerts are mailed (see containers/cloud-backup.yml).
backup_alert_email: bastian@debyl.io
bookstack_path: "{{ podman_volumes }}/bookstack"
cam2ip_path: "{{ podman_volumes }}/cam2ip"
cloud_path: "{{ podman_volumes }}/cloud"
@@ -22,7 +24,6 @@ gregtime_path: "{{ podman_volumes }}/gregtime"
hass_path: "{{ podman_volumes }}/hass"
# nginx_path: removed - nginx no longer used
# nosql_path: removed - nosql/redis no longer used
partkeepr_path: "{{ podman_volumes }}/partkeepr"
partsy_path: "{{ podman_volumes }}/partsy"
partsy_skudak_path: "{{ podman_volumes }}/partsy-skudak"
photos_path: "{{ podman_volumes }}/photos"
@@ -73,7 +74,6 @@ zomboid_maps:
modded: "Muldraugh, KY"
b42revamp: "AnruisiTown;Asakusa lake town;BananaBeachHouse;Blackstone;Cathaya Valley2.0;Cathaya Valley2.0 highway;Daisy County;DawnTown;Estate 39;Fort Waterfront B42;Greenleaf;GreenRiver;IrisEyot;Louisville_Quarantine_Zone;Louisville_River_Marina;Louisville_Riverboat;Maplewood;Meiya'sTown;Muldraugh_FireDept;muldraughmilitarybaseas24;PeaceTown;RaccoonCity;Raven Creek B42;map_distanciado;rvupdate;rv2;serenitybunker;Tikitown;Uncle Red's Bunker B42;Uncle Red's Bunker Redux B42;WestPoint-MilitaryTugBoat;Muldraugh, KY"
pihole_path: "{{ podman_volumes }}/pihole"
sshpass_cron_path: "{{ podman_volumes }}/sshpass_cron"
caddy_path: "{{ podman_volumes }}/caddy"
@@ -110,6 +110,25 @@ bookstack_server_name_new: wiki.skudak.com
cloud_skudak_server_name_new: cloud.skudak.com
gitea_skudak_server_name: git.skudak.com
# LibreSign root certificate authority identity for skudak-cloud. This is the
# issuer name that appears on every signed document, so it must match the
# entity's legal name -- it was previously generated as "Skudak Rennsport LLP",
# the pre-rename name. Changing these values does NOT re-issue the CA on its
# own; see the guarded generate task in containers/skudak/cloud.yml.
# Skudak brand palette, mirroring ~/src/skudak/skudak-site/src/styles/variables.css.
# --color-gray-900 for the mail header band and UI chrome; the white signature
# logo is legible on it. --color-accent is spent only on the CTA button, and
# lives in the skudakmail app rather than here since theming has no second
# colour slot.
theming_skudak_primary: "#0A0A0A"
libresign_skudak_cert_cn: Skudak LLP
libresign_skudak_cert_o: Skudak LLP
libresign_skudak_cert_c: US
libresign_skudak_cert_st: New Hampshire
libresign_skudak_cert_l: Newbury
# Legacy nginx/ModSecurity configuration removed - Caddy provides built-in security
# Web server configuration (Caddy is the default)
@@ -134,6 +153,16 @@ caddy_local_networks:
caddy_log_level: INFO
caddy_log_format: json
# Log rotation. Caddy rotates on implicit defaults (100MiB / keep 10 / 90d)
# without these, and those were demonstrably not holding: 20 rotated files per
# stream for the busiest logs and a 190-day-old .gz, for 1.3 GB across 16
# streams. Kept short deliberately -- fluent-bit tails every one of these into
# Graylog (see roles/common/templates/fluent-bit/fluent-bit.conf.j2), so
# Graylog is the system of record and the local files are only a buffer.
caddy_log_roll_size: 10MiB
caddy_log_roll_keep: 3
caddy_log_roll_keep_for: 168h
# Caddy performance tuning
caddy_max_request_body_mb: 500
@@ -175,3 +204,23 @@ geoip_path: "{{ graylog_path }}/geoip"
geoip_database_edition: GeoLite2-City
# geoip_maxmind_account_id: defined in vault
# geoip_maxmind_license_key: defined in vault
# Weekly podman prune (see templates/podman-prune.sh.j2). Both rootless users
# have their own image store; the git user runs the Gitea pods.
podman_prune_users:
- "{{ podman_user }}"
- "{{ git_user }}"
# Keep 30 days of unused images so a rollback needs no rebuild or re-pull.
podman_prune_until: 720h
podman_prune_oncalendar: "Sun *-*-* 02:00:00"
# Hardened CIFS options for the TrueNAS shares (see containers/home/photos.yml).
# x-systemd.automount is what makes a failed mount recoverable without a human.
cifs_mount_opts: >-
credentials=/etc/cifs/photos.creds,uid={{ podman_subuid.stdout }},gid={{ podman_subuid.stdout }},_netdev,nofail,soft,x-systemd.automount,x-systemd.mount-timeout=30,x-systemd.idle-timeout=600,vers=3.1.1,resilienthandles,echo_interval=10
# How often the watchdog checks mount health and heals immich.
cifs_watchdog_oncalendar: "*:0/5"
cifs_watchdog_mounts:
- "{{ photos_path }}/storage"
- "{{ photos_path }}/immich"
@@ -0,0 +1,44 @@
<?xml version="1.0"?>
<info xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:noNamespaceSchemaLocation="https://apps.nextcloud.com/schema/apps/info.xsd">
<id>skudakmail</id>
<name>Skudak Customisations</name>
<summary>Skudak-branded email templates and UI overrides for Nextcloud and LibreSign</summary>
<description><![CDATA[
Restyles outgoing Nextcloud and LibreSign mail to match the Skudak design
system at ~/src/skudak/skudak-site. Two supported extension points, no core
patch and no LibreSign fork:
1. `OCA\Skudakmail\Mail\SkudakEMailTemplate` extends Nextcloud's EMailTemplate
and is wired in via the `mail_template_class` system config value, which
Nextcloud checks in `lib/private/Mail/Mailer.php::createEMailTemplate()`.
It owns layout, typography, subject rewriting, button labels and the
footer LibreSign never adds.
2. `OCA\Skudakmail\Listener\SkudakMailListener` listens on
`OCP\Mail\Events\BeforeMessageSent` to embed the wordmark as an inline
(cid:) MIME part, so the logo survives the remote-image blocking that
Apple Mail, Gmail and Outlook apply by default. This cannot be done from
the template class, which has no reference to the message.
3. `OCA\Skudakmail\Listener\SkudakStyleListener` listens on
`OCP\AppFramework\Http\Events\BeforeTemplateRenderedEvent` and adds
css/libresign-mobile.css, which fixes the LibreSign public signing page
being clipped at the bottom on iOS Safari. Serving it from here rather
than patching LibreSign keeps the app's integrity signature intact and
survives app updates, which wipe the app directory.
The app has no routes, no UI, no settings and no database tables. The id
remains `skudakmail` for historical reasons -- it is referenced by the
`mail_template_class` system config -- but its scope is Skudak-wide
customisation, not mail alone.
]]></description>
<version>1.0.0</version>
<licence>agpl</licence>
<author>Skudak LLP</author>
<namespace>Skudakmail</namespace>
<category>customization</category>
<dependencies>
<nextcloud min-version="34" max-version="34"/>
</dependencies>
</info>
@@ -0,0 +1,40 @@
/*
* Mobile fix for the LibreSign public signing page.
*
* PROBLEM: src/ExternalApp.vue sets `height: 100vh` on `html body #content`
* and again on `#app-sidebar` under `@media (max-width: 512px)`. iOS Safari
* resolves 100vh against the LARGE viewport -- as though the browser chrome
* were hidden -- so the element extends behind the bottom toolbar and the
* signing action bar is clipped off-screen. The built `external` chunk uses
* 100vh seven times and dvh/svh/safe-area zero times.
*
* WHY NOT safe-area-inset: the page's viewport meta is
* `width=device-width, initial-scale=1.0, minimum-scale=1.0` with no
* `viewport-fit=cover`, so env(safe-area-inset-bottom) resolves to 0 here.
*
* WHY dvh: the dynamic viewport unit tracks the chrome as it shows and hides,
* which is exactly the behaviour wanted. Browsers without dvh support drop the
* declaration entirely and keep LibreSign's own 100vh -- so this degrades to
* today's behaviour rather than to something broken. No @supports needed.
*
* SCOPING IS LOad-BEARING. `#content` and `#app-sidebar` are Nextcloud-wide
* IDs used throughout the authenticated UI. Every rule below is scoped to
* `#body-public` + `.app-public`, which the public signing page sets:
* <body id="body-public" class="layout-base">
* <div id="content" class="app-public" role="main">
* Widening these selectors would restyle the whole instance.
*
* UPSTREAM: patched at source in src/ExternalApp.vue (lines 34 and 46) and
* submitted to LibreSign. Once that lands and this instance runs a release
* containing it, this file can be deleted.
*/
#body-public #content.app-public {
height: 100dvh;
}
@media (max-width: 512px) {
#body-public #app-sidebar {
height: 100dvh;
}
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 10 KiB

@@ -0,0 +1,41 @@
<?php
declare(strict_types=1);
namespace OCA\Skudakmail\AppInfo;
use OCA\Skudakmail\Listener\SkudakMailListener;
use OCA\Skudakmail\Listener\SkudakStyleListener;
use OCP\AppFramework\App;
use OCP\AppFramework\Bootstrap\IBootContext;
use OCP\AppFramework\Bootstrap\IBootstrap;
use OCP\AppFramework\Bootstrap\IRegistrationContext;
use OCP\AppFramework\Http\Events\BeforeTemplateRenderedEvent;
use OCP\Mail\Events\BeforeMessageSent;
class Application extends App implements IBootstrap {
public const APP_ID = 'skudakmail';
public function __construct(array $urlParams = []) {
parent::__construct(self::APP_ID, $urlParams);
}
public function register(IRegistrationContext $context): void {
// BeforeMessageSent fires in Mailer::send() (lib/private/Mail/Mailer.php:186),
// AFTER useTemplate() has flattened the template into subject/plain/html on
// the message, and BEFORE setRecipients() and the Symfony transport. That
// window is the only place an inline (cid:) logo can be attached -- see the
// listener for why the template class alone cannot do it.
$context->registerEventListener(BeforeMessageSent::class, SkudakMailListener::class);
// BeforeTemplateRenderedEvent is dispatched from
// lib/private/AppFramework/Middleware/AdditionalScriptsMiddleware.php:35 and
// lib/private/Template/TemplateManager.php:82 -- the latter covers public
// (unauthenticated) pages, which is the case that matters here since the
// LibreSign signing page is a #[PublicPage].
$context->registerEventListener(BeforeTemplateRenderedEvent::class, SkudakStyleListener::class);
}
public function boot(IBootContext $context): void {
}
}
@@ -0,0 +1,121 @@
<?php
declare(strict_types=1);
/**
* Embeds the Skudak wordmark as an inline (cid:) MIME part.
*
* WHY A LISTENER AND NOT THE TEMPLATE CLASS: Apple Mail, Gmail and Outlook all
* block remote images by default, and Apple Mail draws its own placeholder box
* rather than styled alt text -- so no amount of styling in the HTML rescues a
* remote <img>. The fix is a cid: reference backed by an inline MIME part, and
* that part must be attached to the MESSAGE. An IEMailTemplate subclass has no
* reference to the message, so it physically cannot do this; the template emits
* the <img>, this listener supplies the bytes and rewrites the src.
*
* BeforeMessageSent is the sanctioned hook -- "Emitted before a system mail is
* sent. It can be used to alter the message." (lib/public/Mail/Events/
* BeforeMessageSent.php). It fires at lib/private/Mail/Mailer.php:186, after
* useTemplate() has already rendered subject/plain/html onto the message and
* before setRecipients() and the transport, so a body rewrite here takes
* effect. No core patch, no LibreSign fork.
*
* FAILURE POSTURE: every step is defensive. If the asset is missing, the body
* is not ours, or anything throws, the listener leaves the message untouched
* and mail still goes out with a remote <img> -- degraded, never blocked. Mail
* that carries signature requests must not fail to send because branding
* broke.
*/
namespace OCA\Skudakmail\Listener;
use OC\Mail\Message;
use OCP\EventDispatcher\Event;
use OCP\EventDispatcher\IEventListener;
use OCP\Mail\Events\BeforeMessageSent;
use Psr\Log\LoggerInterface;
/** @template-implements IEventListener<BeforeMessageSent> */
class SkudakMailListener implements IEventListener {
/** Must match SkudakEMailTemplate::LOGO_PATH. */
private const LOGO_PATH_FRAGMENT = '/custom_apps/skudakmail/img/skudak-wordmark.png';
/** Content-ID. Symfony emits this as <skudak-wordmark.png>. */
private const CID = 'skudak-wordmark.png';
public function __construct(
private LoggerInterface $logger,
) {
}
public function handle(Event $event): void {
if (!$event instanceof BeforeMessageSent) {
return;
}
try {
$this->embedWordmark($event->getMessage());
} catch (\Throwable $e) {
// Never let branding break delivery of a signature request.
$this->logger->warning('skudakmail: inline logo embed skipped', [
'exception' => $e,
]);
}
}
private function embedWordmark(\OCP\Mail\IMessage $message): void {
// Mailer::send() guards `instanceof Message` before dispatching this
// event, so the concrete type is guaranteed -- but getSymfonyEmail()
// is not on the interface, so narrow explicitly rather than assume.
if (!$message instanceof Message) {
return;
}
$email = $message->getSymfonyEmail();
$html = $email->getHtmlBody();
if (!is_string($html) || $html === '') {
return;
}
// Only touch mail that actually renders our wordmark. Anything else --
// password resets, share notifications, other apps -- passes through.
if (!str_contains($html, self::LOGO_PATH_FRAGMENT)) {
return;
}
$asset = $this->assetPath();
if ($asset === null) {
return;
}
$bytes = @file_get_contents($asset);
if ($bytes === false || $bytes === '') {
return;
}
// Rewrite the absolute URL to a cid: reference. Matched on the path
// fragment with an optional query string so a cachebuster or a change
// of host still resolves.
$rewritten = preg_replace(
'#https?://[^"\']*' . preg_quote(self::LOGO_PATH_FRAGMENT, '#') . '(\?[^"\']*)?#',
'cid:' . self::CID,
$html,
);
if (!is_string($rewritten) || $rewritten === $html) {
return;
}
$email->embed($bytes, self::CID, 'image/png');
$message->setHtmlBody($rewritten);
}
/**
* Resolves img/skudak-wordmark.png relative to this file, so the app works
* from whatever apps directory Nextcloud has it in.
*/
private function assetPath(): ?string {
$path = dirname(__DIR__, 2) . '/img/skudak-wordmark.png';
return is_readable($path) ? $path : null;
}
}
@@ -0,0 +1,51 @@
<?php
declare(strict_types=1);
/**
* Injects Skudak's CSS overrides into rendered Nextcloud pages.
*
* Currently one override: the LibreSign public signing page clips its bottom
* action bar on iOS Safari, because ExternalApp.vue sizes #content to 100vh and
* Safari resolves that against the large viewport (chrome hidden). See
* css/libresign-mobile.css for the full reasoning.
*
* WHY A LISTENER RATHER THAN PATCHING LIBRESIGN: an app-store app carries
* appinfo/signature.json, so editing a single byte of it raises INVALID_HASH in
* the admin security check, and an app update wipes the directory outright
* (Installer::downloadApp() calls Files::rmdirr on it). A stylesheet served
* from our own app survives both, and survives Nextcloud upgrades.
*
* The stylesheet itself is tightly scoped to #body-public / .app-public. This
* listener is deliberately NOT scoped further -- adding a stylesheet is
* idempotent and cheap, and gating on which app is rendering would couple this
* to LibreSign's route structure for no benefit. The CSS decides where it
* applies; this only decides that it is available.
*/
namespace OCA\Skudakmail\Listener;
use OCA\Skudakmail\AppInfo\Application;
use OCP\AppFramework\Http\Events\BeforeTemplateRenderedEvent;
use OCP\EventDispatcher\Event;
use OCP\EventDispatcher\IEventListener;
use OCP\Util;
/** @template-implements IEventListener<BeforeTemplateRenderedEvent> */
class SkudakStyleListener implements IEventListener {
public function handle(Event $event): void {
if (!$event instanceof BeforeTemplateRenderedEvent) {
return;
}
// Never let a styling concern break page rendering. A signing page that
// loads unstyled is recoverable; one that 500s is not.
try {
Util::addStyle(Application::APP_ID, 'libresign-mobile');
} catch (\Throwable $e) {
// Intentionally swallowed -- no logger dependency is worth adding
// for a stylesheet, and a failure here has no user-visible effect
// beyond the override not applying.
}
}
}
@@ -0,0 +1,417 @@
<?php
declare(strict_types=1);
/**
* Skudak-branded email template.
*
* Wired in via the `mail_template_class` system config value, which Nextcloud
* checks in lib/private/Mail/Mailer.php::createEMailTemplate(). That is a
* supported extension point -- core is not patched, so Nextcloud upgrades do
* not clobber this.
*
* WHY THIS EXISTS AT ALL: LibreSign's outgoing mail is generic open-source
* boilerplate -- subject "LibreSign: There is a file for you to sign", heading
* "File to sign", button "Sign »filename«", and NO footer whatsoever (it never
* calls addFooter(); verified: zero hits for addFooter in custom_apps/libresign).
* That mail carries partnership instruments to signers, so it needs to read as
* an official Skudak LLP communication.
*
* DESIGN INTENT (mirrors ~/src/skudak/skudak-site/src/styles/variables.css):
* - Light ground, near-black text, NO coloured header band. skudak.com is
* --color-white #FAFAFA with --color-gray-900 #0A0A0A text; transactional
* mail from Stripe/Linear/DocuSign is likewise restrained. A band also
* leaves an ugly empty slab when the logo is blocked (see LOGO note).
* - --color-accent #2563EB on the CTA button only. That is the one place the
* site spends colour, so it is the one place this template does.
* - Inter, matching --font-sans, with the stock stack as fallback.
*
* LOGO: served from this app's own img/ directory rather than the theming app.
* Two reasons. (1) The theming logo is white-on-transparent because the web UI
* and login page are dark; a white mark is invisible on this template's white
* ground. (2) Decoupling means restyling mail can never disturb the web UI.
* /custom_apps/<app>/img/<file> is served publicly without auth (verified).
*
* Note that remote images are blocked by default in Apple Mail, Gmail and
* Outlook, and Apple Mail renders its own placeholder box rather than styled
* alt text -- so alt styling cannot rescue it. Surviving that requires a CID
* inline part via IMessage::attachInline(), which lives on the MESSAGE and is
* unreachable from a template subclass. Mitigated instead by dropping the
* band: a blocked logo now leaves plain white space, not a black slab.
*
* IMPLEMENTATION NOTE: font restyling is done by string-substitution against
* the PARENT's own markup rather than by redefining it. Those properties are
* large inline-CSS blobs with positional sprintf placeholders; copying them
* wholesale would mean re-auditing every placeholder on every upgrade, and a
* mismatch renders broken mail. Substitution degrades safely -- if upstream
* changes markup the replacements no-op and mail still sends, just unstyled.
* The header IS replaced wholesale, deliberately, because "no band" cannot be
* expressed as a substitution; its placeholder order is documented at its
* definition and must be kept in sync with upstream.
*/
namespace OCA\Skudakmail\Mail;
use OC\Mail\EMailTemplate;
class SkudakEMailTemplate extends EMailTemplate {
/** Stock Nextcloud font stack, replaced wholesale. Must match exactly. */
private const STOCK_FONTS = "-apple-system,BlinkMacSystemFont,'Segoe UI',Roboto,Oxygen-Sans,Ubuntu,Cantarell,'Helvetica Neue',Arial,sans-serif";
/** --font-sans, with the stock stack retained as fallback. */
private const SKUDAK_FONTS = "Inter,-apple-system,BlinkMacSystemFont,'Segoe UI',Roboto,Oxygen-Sans,Ubuntu,Cantarell,'Helvetica Neue',Arial,sans-serif";
private const ACCENT = '#2563EB'; // --color-accent
private const ON_ACCENT = '#FAFAFA'; // --color-white
private const INK = '#0A0A0A'; // --color-gray-900
private const MUTED = '#525252'; // --color-gray-500
private const FAINT = '#A3A3A3'; // --color-gray-300
private const RULE = '#E5E5E5'; // --color-gray-100
private const ENTITY = 'Skudak LLP';
private const SITE = 'https://skudak.com';
private const LOGO_PATH = '/custom_apps/skudakmail/img/skudak-wordmark.png';
/** Displayed width in px. The asset is 600px wide for retina. */
private const LOGO_DISPLAY_WIDTH = 190;
/**
* LibreSign's l10n wraps document names in German guillemets -- "Sign
* »contract«" -- regardless of locale. Mapped to US curly quotes, matching
* the ``...'' convention in the LaTeX document templates.
*/
private const QUOTE_MAP = ['»' => "\u{201C}", '«' => "\u{201D}"];
/**
* LibreSign subject -> Skudak subject. Keys are the exact English msgids
* from custom_apps/libresign/lib/Service/MailService.php (lines 51, 87,
* 121, 150, 172). Anything unmatched passes through untouched, so an
* upstream string change degrades to the original subject rather than a
* blank one.
*/
private const SUBJECT_MAP = [
'LibreSign: There is a file for you to sign' => 'Document for your signature',
'LibreSign: Changes into a file for you to sign' => 'Updated document for your signature',
'LibreSign: A file has been signed' => 'A document has been signed',
'LibreSign: A signature request has been canceled' => 'Signature request cancelled',
'LibreSign: Code to sign file' => 'Your signing verification code',
];
/**
* LibreSign heading -> Skudak heading. Exact English msgids from
* MailService.php lines 53/89, 123, 152.
*/
private const HEADING_MAP = [
'File to sign' => 'Review and sign',
'File signed' => 'Document signed',
'Signature request canceled' => 'Signature request cancelled',
];
/**
* LibreSign body copy -> Skudak body copy (MailService.php lines 60, 96,
* 174). Only the strings with NO %s interpolation are mapped; the two that
* carry a name or filename (lines 125, 154) arrive already substituted and
* so cannot be matched exactly -- they pass through unchanged.
*/
private const BODY_MAP = [
'There is a document for you to sign. Access the link below:'
=> 'Skudak LLP has sent you a document that requires your signature. Review it and sign using the link below.',
'Changes have been made in a file that you have to sign. Access the link below:'
=> 'A document awaiting your signature has been updated by Skudak LLP. Review the current version and sign using the link below.',
'Use this code to sign the document:'
=> 'Use this verification code to complete your signature:',
];
/**
* Template properties carrying the font stack. Listed explicitly rather
* than discovered reflectively so an upstream rename fails loudly in
* testing instead of silently skipping a block.
*/
private const STYLED_PARTS = [
'head', 'tail', 'heading', 'bodyBegin', 'bodyText',
'listBegin', 'listItem', 'listEnd', 'buttonGroup', 'button',
'bodyEnd', 'footer',
];
/**
* Own flag, deliberately NOT the parent's $footerAdded.
*
* Message::useTemplate() (lib/private/Mail/Message.php:289-296) calls
* renderText() at :291 BEFORE renderHtml() at :293, and renderText() sets
* $footerAdded = true. Guarding footer injection on !$footerAdded therefore
* never fires on the real send path -- the footer silently vanished from
* every mail while a renderHtml()-only test passed. Both renderers below
* call inject() and this flag makes the second call inert.
*/
private bool $skudakFooterInjected = false;
public function __construct(
\OCP\Defaults $themingDefaults,
\OCP\IURLGenerator $urlGenerator,
\OCP\L10N\IFactory $l10nFactory,
?int $logoWidth,
?int $logoHeight,
string $emailId,
array $data,
) {
$this->applySkudakStyling();
// Must run AFTER the substitutions: the parent constructor copies
// $this->head into $htmlBody as its first act, so restyling head
// afterwards would leave the already-emitted copy untouched.
parent::__construct(
$themingDefaults,
$urlGenerator,
$l10nFactory,
$logoWidth,
$logoHeight,
$emailId,
$data,
);
}
private function applySkudakStyling(): void {
foreach (self::STYLED_PARTS as $part) {
if (!property_exists($this, $part)) {
continue;
}
$this->$part = str_replace(self::STOCK_FONTS, self::SKUDAK_FONTS, $this->$part);
}
// Site headings are --font-weight-light with tightened tracking.
$this->heading = str_replace(
'font-size:24px;font-weight:400',
'font-size:26px;font-weight:300;letter-spacing:-0.02em',
$this->heading,
);
}
/**
* Rewrites LibreSign's subjects. Called by LibreSign on the TEMPLATE
* (MailService.php:51 etc.), not on the message, which is what makes this
* interceptable at all -- Message::useTemplate() later pulls the result via
* renderSubject(). Prefixed with the entity so the sender is unambiguous in
* an inbox list.
*/
public function setSubject(string $subject): void {
$mapped = self::SUBJECT_MAP[$subject] ?? null;
parent::setSubject(
$mapped === null ? $subject : self::ENTITY . ' — ' . $mapped,
);
}
/**
* Replaces the stock header wholesale: no coloured band, wordmark centred
* on white.
*
* Does NOT use the parent's $header property or its placeholder order --
* this is independent markup, so upstream changes to $header cannot break
* it (and equally cannot improve it). $logoWidth/$logoHeight from the
* Mailer are ignored on purpose: they are clamped to MAX_LOGO_SIZE = 105
* (lib/private/Mail/Mailer.php:60), which is too small for a wordmark to
* be legible.
*/
public function addHeader(): void {
if ($this->headerAdded) {
return;
}
$this->headerAdded = true;
$logoUrl = $this->urlGenerator->getAbsoluteURL(self::LOGO_PATH);
$alt = htmlspecialchars(self::ENTITY, ENT_QUOTES, 'UTF-8');
$w = self::LOGO_DISPLAY_WIDTH;
$fonts = self::SKUDAK_FONTS;
$ink = self::INK;
$this->htmlBody .= <<<HTML
<table align="center" style="border-collapse:collapse;border-spacing:0;margin:0 auto;padding:0;text-align:left;vertical-align:top;width:100%">
<tbody><tr style="padding:0;text-align:left;vertical-align:top">
<td align="center" style="border-collapse:collapse!important;margin:0;padding:40px 30px 28px 30px;text-align:center;vertical-align:top">
<img src="{$logoUrl}" alt="{$alt}" width="{$w}" style="-ms-interpolation-mode:bicubic;border:0;clear:both;display:block;margin:0 auto;outline:0;text-decoration:none;width:{$w}px;max-width:{$w}px;height:auto;color:{$ink};font-family:{$fonts};font-size:22px;font-weight:300;letter-spacing:-0.02em"/>
</td>
</tr></tbody>
</table>
HTML;
}
/**
* Both renderers inject the footer -- see $skudakFooterInjected.
*
* Mirrors the parent's own guard structure (renderHtml at
* lib/private/Mail/EMailTemplate.php:643, renderText at :656): close the
* body, append $tail, flip $footerAdded. The Skudak block goes in before
* $tail.
*/
public function renderHtml(): string {
$this->injectSkudakFooter();
return parent::renderHtml();
}
public function renderText(): string {
$this->injectSkudakFooter();
return parent::renderText();
}
private function injectSkudakFooter(): void {
if ($this->skudakFooterInjected || $this->footerAdded) {
return;
}
$this->skudakFooterInjected = true;
// Close the body ourselves so the footer lands INSIDE the layout
// rather than after it. The parent's render methods are then a no-op
// for body closing and only append $tail.
$this->ensureBodyIsClosed();
$this->htmlBody .= $this->skudakFooterHtml();
$this->plainBody .= $this->skudakFooterText();
}
private function skudakFooterHtml(): string {
$year = date('Y');
$entity = htmlspecialchars(self::ENTITY, ENT_QUOTES, 'UTF-8');
$fonts = self::SKUDAK_FONTS;
$site = self::SITE;
[$muted, $faint, $rule, $ink] = [self::MUTED, self::FAINT, self::RULE, self::INK];
// Table-based and fully inline-styled: <style> blocks, flex and grid
// are stripped or unsupported across Outlook and most webmail.
return <<<HTML
<table align="center" style="border-collapse:collapse;border-spacing:0;margin:0 auto;padding:0;text-align:left;vertical-align:top;width:100%">
<tbody><tr style="padding:0;text-align:left;vertical-align:top">
<td align="center" style="border-collapse:collapse!important;margin:0;padding:0 30px 44px 30px;text-align:center;vertical-align:top">
<table align="center" style="border-collapse:collapse;border-spacing:0;margin:0 auto;padding:0;text-align:center;width:100%;max-width:550px">
<tbody>
<tr><td style="border-collapse:collapse!important;border-top:1px solid {$rule};font-size:0;line-height:0;height:1px;margin:0;padding:0">&#xA0;</td></tr>
<tr><td align="center" style="border-collapse:collapse!important;color:{$muted};font-family:{$fonts};font-size:13px;font-weight:400;line-height:1.6;margin:0;padding:22px 0 0 0;text-align:center">
This is an official document-signing request from <strong style="color:{$ink};font-weight:600">{$entity}</strong>.<br/>
Nothing is signed unless you open the document and complete it yourself. If you were not expecting this, you can safely ignore it.
</td></tr>
<tr><td align="center" style="border-collapse:collapse!important;color:{$muted};font-family:{$fonts};font-size:13px;font-weight:400;line-height:1.6;margin:0;padding:16px 0 0 0;text-align:center">
<a href="{$site}/privacy" style="color:{$muted};text-decoration:underline">Privacy Policy</a>
&#160;&#183;&#160;
<a href="{$site}/terms" style="color:{$muted};text-decoration:underline">Terms of Use</a>
&#160;&#183;&#160;
<a href="{$site}" style="color:{$muted};text-decoration:underline">skudak.com</a>
</td></tr>
<tr><td align="center" style="border-collapse:collapse!important;color:{$faint};font-family:{$fonts};font-size:12px;font-weight:400;line-height:1.6;margin:0;padding:16px 0 0 0;text-align:center">
&copy; {$year} {$entity}. All rights reserved.<br/>
Automated message &mdash; please do not reply to this address.
</td></tr>
</tbody>
</table>
</td>
</tr></tbody>
</table>
HTML;
}
private function skudakFooterText(): string {
$year = date('Y');
$entity = self::ENTITY;
$site = self::SITE;
return <<<TEXT
--
This is an official document-signing request from {$entity}.
Nothing is signed unless you open the document and complete it yourself.
If you were not expecting this, you can safely ignore it.
Privacy Policy: {$site}/privacy
Terms of Use: {$site}/terms
© {$year} {$entity}. All rights reserved.
Automated message — please do not reply to this address.
TEXT;
}
private function tidyQuotes(string $text): string {
return strtr($text, self::QUOTE_MAP);
}
// Signatures below mirror the parent EXACTLY. $plainTitle/$plainText are
// deliberately untyped there (they accept string|bool -- false suppresses
// the plain-text variant), and narrowing a parameter type in an override
// is a fatal error in PHP.
public function addHeading(string $title, $plainTitle = ''): void {
$mapped = self::HEADING_MAP[$title] ?? $this->tidyQuotes($title);
parent::addHeading(
$mapped,
is_string($plainTitle) && $plainTitle !== ''
? (self::HEADING_MAP[$plainTitle] ?? $this->tidyQuotes($plainTitle))
: $plainTitle,
);
}
public function addBodyText(string $text, $plainText = ''): void {
$mapped = self::BODY_MAP[$text] ?? $this->tidyQuotes($text);
parent::addBodyText(
$mapped,
is_string($plainText) && $plainText !== ''
? (self::BODY_MAP[$plainText] ?? $this->tidyQuotes($plainText))
: $plainText,
);
}
/**
* Reimplemented for two reasons: the accent colour, and a fixed label.
*
* LibreSign builds "Sign »%s«" with the raw filename
* (MailService.php:64,100). Real documents here are named things like
* amendment-001-partner-compensation, which makes an ungainly button and
* leaks the document name to anyone who sees the inbox preview. Replaced
* with a fixed call to action; the document is identified on the landing
* page behind the link.
*
* Mirrors the parent's vsprintf argument order exactly:
* [$color, $color, $url, $color, $textColor, $textColor, $text].
* Kept in sync with parent::addBodyButton() -- if that changes upstream,
* this needs revisiting.
*/
public function addBodyButton(string $text, string $url, $plainText = ''): void {
if ($this->footerAdded) {
return;
}
$this->ensureBodyIsOpened();
$this->ensureBodyListClosed();
$label = $this->buttonLabelFor($text);
if ($plainText === '') {
$plainText = $label;
} elseif (is_string($plainText)) {
$plainText = $this->tidyQuotes($plainText);
}
$this->htmlBody .= vsprintf($this->button, [
self::ACCENT,
self::ACCENT,
$url,
self::ACCENT,
self::ON_ACCENT,
self::ON_ACCENT,
htmlspecialchars($label, ENT_QUOTES, 'UTF-8'),
]);
if ($plainText !== false) {
$this->plainBody .= $plainText . ': ';
}
$this->plainBody .= $url . PHP_EOL;
}
/**
* Maps LibreSign's filename-bearing labels onto fixed calls to action.
* Matched on the stable leading verb rather than the whole string, since
* the tail is a filename. Unknown labels pass through with quotes tidied.
*/
private function buttonLabelFor(string $text): string {
if (str_starts_with($text, 'Sign ')) {
return 'Review document';
}
if (str_starts_with($text, 'View signed file')) {
return 'View signed document';
}
return $this->tidyQuotes($text);
}
}
@@ -18,6 +18,18 @@
mode: 0600
setype: ssh_home_t
# A direct host-to-S3 stage was added here and then removed. Offsite already
# happens: the TrueNAS rsync below feeds /mnt/glacier/skudakcloud, and a
# TrueNAS cloud-sync task pushes that to Skudak's own iDrive e2 bucket. A
# second, direct push would have written the same data into the same bucket
# twice. If offsite is ever moved onto this host, it should REPLACE the rsync
# rather than run alongside it.
- name: remove obsolete backup S3 credentials
become: true
ansible.builtin.file:
path: "/etc/backup_s3/{{ backup_name }}"
state: absent
- name: template {{ backup_name }} backup script
become: true
ansible.builtin.template:
@@ -28,6 +40,27 @@
mode: 0755
setype: bin_t
# Shared by every backup instance. Rendered once per include; the second and
# later renders are no-ops.
- name: template nextcloud backup alert script
become: true
ansible.builtin.template:
src: nextcloud/nextcloud-backup-alert.sh.j2
dest: /usr/local/bin/nextcloud-backup-alert.sh
owner: root
group: root
mode: 0755
setype: bin_t
- name: template nextcloud backup failure handler unit
become: true
ansible.builtin.template:
src: nextcloud/nextcloud-backup-failed@.service.j2
dest: /etc/systemd/system/nextcloud-backup-failed@.service
owner: root
group: root
mode: 0644
- name: template {{ backup_name }} backup systemd service
become: true
ansible.builtin.template:
@@ -0,0 +1,40 @@
---
- name: template {{ cron_name }} cron script
become: true
ansible.builtin.template:
src: nextcloud/cloud-cron.sh.j2
dest: "{{ cron_script_path }}"
owner: root
group: root
mode: 0755
setype: bin_t
- name: template {{ cron_name }} cron systemd service
become: true
ansible.builtin.template:
src: nextcloud/cloud-cron.service.j2
dest: "/etc/systemd/system/{{ cron_name }}-cron.service"
owner: root
group: root
mode: 0644
vars:
instance_name: "{{ cron_name }}"
- name: template {{ cron_name }} cron systemd timer
become: true
ansible.builtin.template:
src: nextcloud/cloud-cron.timer.j2
dest: "/etc/systemd/system/{{ cron_name }}-cron.timer"
owner: root
group: root
mode: 0644
vars:
instance_name: "{{ cron_name }}"
- name: enable and start {{ cron_name }} cron timer
become: true
ansible.builtin.systemd:
name: "{{ cron_name }}-cron.timer"
enabled: true
state: started
daemon_reload: true
@@ -84,10 +84,44 @@
vars:
container_name: cloud
# Unbounded by default: nextcloud.log.1 had reached 1.12 GB and was being
# rsynced to TrueNAS and pushed to S3 on every run. Cap at 10 MiB.
- name: cap nextcloud log rotation size for cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data cloud
php occ config:system:set log_rotate_size --value 10485760 --type integer
register: cloud_log_rotate
changed_when: "'System config value log_rotate_size' in cloud_log_rotate.stdout"
failed_when: false
# Nextcloud's default ('auto') only expires trash when disk space is needed,
# so 66 GB of >30-day deletions sat untouched on a host with 1.3 TB free --
# the retention was effectively unbounded. 'auto, 30' makes the 30-day
# expiry unconditional while still purging early under space pressure.
- name: set nextcloud trashbin retention for cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data cloud
php occ config:system:set trashbin_retention_obligation --value "auto, 30"
register: cloud_trashbin_retention
changed_when: "'System config value trashbin_retention_obligation' in cloud_trashbin_retention.stdout"
failed_when: false
- include_tasks: containers/cloud-cron.yml
vars:
cron_name: cloud
cron_container: cloud
cron_script_path: /usr/local/bin/cloud-cron.sh
- include_tasks: containers/cloud-backup.yml
vars:
backup_name: cloud
data_path: "{{ cloud_path }}/data"
config_path: "{{ cloud_path }}/config"
db_container: cloud-db
ssh_key_path: /etc/ssh/backup_keys/cloud
ssh_key_content: "{{ cloud_backup_ssh_key }}"
ssh_user: cloud
@@ -1,59 +0,0 @@
---
# PartKeepr has been replaced by Partsy
# This playbook removes PartKeepr containers and services
# Keeping MySQL data volume for historical reference only
- name: stop and remove partkeepr container
become: true
become_user: "{{ podman_user }}"
containers.podman.podman_container:
name: partkeepr
state: absent
- name: stop and remove partkeepr-db container
become: true
become_user: "{{ podman_user }}"
containers.podman.podman_container:
name: partkeepr-db
state: absent
- name: remove systemd service for partkeepr
become: true
ansible.builtin.systemd:
name: "podman-partkeepr.service"
state: stopped
enabled: false
daemon_reload: true
ignore_errors: true
- name: remove systemd service for partkeepr-db
become: true
ansible.builtin.systemd:
name: "podman-partkeepr-db.service"
state: stopped
enabled: false
daemon_reload: true
ignore_errors: true
- name: remove systemd service files for partkeepr
become: true
ansible.builtin.file:
path: "{{ item }}"
state: absent
loop:
- "/etc/systemd/system/podman-partkeepr.service"
- "/etc/systemd/system/podman-partkeepr-db.service"
notify: systemd daemon-reload
- name: preserve partkeepr mysql data volume for history
become: true
ansible.builtin.file:
path: "{{ partkeepr_path }}/mysql"
state: directory
owner: "{{ podman_subuid.stdout }}"
group: "{{ podman_user }}"
mode: 0755
notify: restorecon podman
- name: flush handlers
ansible.builtin.meta: flush_handlers
@@ -17,13 +17,56 @@
- name: flush handlers
ansible.builtin.meta: flush_handlers
- name: create cifs credentials directory
become: true
ansible.builtin.file:
path: /etc/cifs
state: directory
owner: root
group: root
mode: 0700
# /etc/fstab is 0644 by design, so an inline `password=` is readable by every
# local user. Keep the credential in a 0600 file and reference it instead.
- name: deploy photos cifs credentials
become: true
ansible.builtin.template:
src: cifs-credentials.j2
dest: /etc/cifs/photos.creds
owner: root
group: root
mode: 0600
vars:
cifs_username: photos
cifs_password: "{{ photos_cifs_pass }}"
no_log: true
# These mounts previously failed permanently whenever TrueNAS was power-cycled:
# a systemd .mount unit does NOT retry after a failed attempt, so the share came
# back and nothing remounted, leaving immich serving an empty library while its
# database still referenced 14k assets.
#
# x-systemd.automount the actual fix -- any ACCESS to the path re-attempts the
# mount, so a failure stops being terminal
# _netdev / nofail order after network-online; a dead NAS must not block boot
# soft fail I/O with an error instead of blocking forever, so
# immich-server can be restarted during an outage rather
# than wedging in uninterruptible sleep. Accepted trade-off:
# a write interrupted mid-flight fails and is retried.
# idle-timeout unmount when unused, which clears stale handles instead
# of nursing a half-dead connection
# resilienthandles SMB3 rides out brief server blips transparently
#
# See also cifs-watchdog.sh.j2, which restarts immich-server when a mount that
# was unhealthy becomes healthy again -- the automount restores the FILESYSTEM,
# but the container still holds the old, empty view until it is bounced.
- name: mount photos cifs
become: true
ansible.posix.mount:
src: "{{ photos_cifs_src }}"
path: "{{ photos_path }}/storage"
fstype: cifs
opts: "username=photos,password={{ photos_cifs_pass }},uid={{ podman_subuid.stdout }},gid={{ podman_subuid.stdout }}"
opts: "{{ cifs_mount_opts }}"
state: mounted
- name: mount immich cifs
@@ -32,9 +75,23 @@
src: "{{ immich_cifs_src }}"
path: "{{ photos_path }}/immich"
fstype: cifs
opts: "username=photos,password={{ photos_cifs_pass }},uid={{ podman_subuid.stdout }},gid={{ podman_subuid.stdout }}"
opts: "{{ cifs_mount_opts }}"
state: mounted
# systemd-fstab-generator creates the .automount unit from x-systemd.automount,
# but ansible.posix.mount mounts the path directly and never starts it, leaving
# the on-access trigger INACTIVE -- so a dropped mount stayed dropped, which is
# the exact failure this work exists to fix. Enable it explicitly.
- name: enable cifs automount units
become: true
ansible.builtin.systemd:
name: "{{ item }}"
enabled: true
state: started
daemon_reload: true
loop: "{{ cifs_watchdog_mounts | map('regex_replace', '^/', '') | map('regex_replace', '/', '-') | map('regex_replace', '$', '.automount') | list }}"
failed_when: false
- import_tasks: podman/podman-check.yml
vars:
container_name: immich-machine-learning
@@ -25,15 +25,17 @@
- name: flush handlers
ansible.builtin.meta: flush_handlers
- name: copy skudak cloud libresign setup script
# The former libresign-setup.sh before-starting hook re-ran
# `occ libresign:install --java/--pdftk/--jsignpdf` on every container start,
# with every line ending in `|| echo`, so six months of failures logged
# nothing. Those binaries live under data/appdata_*/libresign, which IS a
# persisted volume, so they only ever needed installing once. Installation and
# verification are now explicit Ansible tasks below that actually fail.
- name: remove obsolete skudak cloud libresign setup hook
become: true
ansible.builtin.template:
src: nextcloud/libresign-setup.sh.j2
dest: "{{ cloud_skudak_path }}/scripts/libresign-setup.sh"
owner: "{{ podman_subuid.stdout }}"
group: "{{ podman_subuid.stdout }}"
mode: 0755
notify: restorecon podman
ansible.builtin.file:
path: "{{ cloud_skudak_path }}/scripts/libresign-setup.sh"
state: absent
- import_tasks: podman/podman-check.yml
vars:
@@ -63,6 +65,90 @@
vars:
container_name: skudak-cloud-db
# ---------------------------------------------------------------------------
# Redis: Nextcloud distributed cache + transactional file locking.
#
# Without it, memcache.locking is unset and Nextcloud falls back to
# DBLockingProvider (lib/private/Server.php:977) -- every file lock becomes a
# MariaDB write against oc_file_locks. A single directory PROPFIND takes dozens
# of locks and two desktop sync clients issue them continuously, which is the
# contention behind the intermittent multi-second stalls.
#
# Also fixes a second problem: memcache.local is APCu, which is PER-PROCESS,
# and Apache here runs mpm_prefork -- so every child holds its own cold cache.
# A distributed cache is shared across all of them.
#
# MUST be created before skudak-cloud below. The Nextcloud entrypoint writes
# its redis config on start; if the host does not resolve at that moment the
# instance comes up pointing at nothing.
- name: create skudak cloud redis config directory
become: true
ansible.builtin.file:
path: "{{ cloud_skudak_path }}/redis"
state: directory
owner: "{{ podman_subuid.stdout }}"
group: "{{ podman_subuid.stdout }}"
mode: 0755
notify: restorecon podman
- name: template skudak cloud redis config
become: true
ansible.builtin.template:
src: nextcloud/redis-skudak.conf.j2
dest: "{{ cloud_skudak_path }}/redis/redis.conf"
owner: "{{ podman_subuid.stdout }}"
group: "{{ podman_subuid.stdout }}"
mode: 0640
notify: restorecon podman
no_log: true
- name: flush handlers
ansible.builtin.meta: flush_handlers
# The redis:alpine image runs redis-server as uid 999 / gid 1000, NOT root, and
# the config is mounted :ro so the image's own entrypoint cannot chown it --
# it logs "cannot change owner ... Read-only file system" and then dies with
# "Fatal error, can't open config file: Permission denied", crash-looping.
# Nextcloud, already pointed at redis by then, answers HTTP 500.
#
# Same idiom as the `podman unshare chown -R 33:33` for www-data above: map the
# in-container uid through the rootless userns. 0640 owned by 999:1000 keeps
# the password unreadable to other users on the host while letting redis read
# it -- which is the entire reason for using a file over --requirepass.
- name: unshare chown the skudak redis config to the redis uid
become: true
become_user: "{{ podman_user }}"
changed_when: false
ansible.builtin.command: >
podman unshare chown 999:1000 {{ cloud_skudak_path }}/redis/redis.conf
- import_tasks: podman/podman-check.yml
vars:
container_name: skudak-cloud-redis
container_image: "{{ redis_image }}"
- name: create skudak-cloud-redis container
become: true
become_user: "{{ podman_user }}"
containers.podman.podman_container:
name: skudak-cloud-redis
image: "{{ redis_image }}"
restart_policy: on-failure:3
log_driver: journald
network:
- shared
# No `ports:` -- deliberately unpublished. Service discovery is by
# container name over `shared`, the same way MYSQL_HOST reaches
# skudak-cloud-db.
volumes:
- "{{ cloud_skudak_path }}/redis/redis.conf:/etc/redis/redis.conf:ro"
command: redis-server /etc/redis/redis.conf
- name: create systemd startup job for skudak-cloud-redis
include_tasks: podman/systemd-generate.yml
vars:
container_name: skudak-cloud-redis
- import_tasks: podman/podman-check.yml
vars:
container_name: skudak-cloud
@@ -83,11 +169,34 @@
MYSQL_DATABASE: skucloud
MYSQL_HOST: skudak-cloud-db
MYSQL_USER: skucloud
# LibreSign signs PDFs in-process; the image default of 512M is not
# enough and manifests as an opaque failure mid-signature.
PHP_MEMORY_LIMIT: 1024M
PHP_UPLOAD_LIMIT: 512M
# Without these the JVM comes up as ANSI_X3.4-1968 and LibreSign's
# config check warns that accented characters in signer names will be
# mangled. See LibreSign issue #4872.
LC_ALL: C.UTF-8
LANG: C.UTF-8
# These three env vars are the WHOLE redis wiring. The image ships
# config/redis.config.php, which -- when REDIS_HOST is set -- declares
# memcache.distributed, memcache.locking AND the connection block.
# Verified against the copy in this instance's persisted config volume.
#
# Do NOT also `occ config:system:set` those keys. occ writes config.php,
# but Nextcloud merges every *.config.php drop-in AFTER it, so the
# drop-in wins -- config.php would read as authoritative while being
# silently overridden. memcache.local stays APCu (apcu.config.php).
#
# REDIS_HOST_PASSWORD_FILE is NOT usable here: the drop-in in this
# volume predates that feature and reads only REDIS_HOST_PASSWORD.
REDIS_HOST: skudak-cloud-redis
REDIS_HOST_PORT: "6379"
REDIS_HOST_PASSWORD: "{{ cloud_skudak_redis_pass }}"
volumes:
- "{{ cloud_skudak_path }}/apps:/var/www/html/custom_apps"
- "{{ cloud_skudak_path }}/data:/var/www/html/data"
- "{{ cloud_skudak_path }}/config:/var/www/html/config"
- "{{ cloud_skudak_path }}/scripts/libresign-setup.sh:/docker-entrypoint-hooks.d/before-starting/libresign-setup.sh:ro"
ports:
- "8090:80"
@@ -96,20 +205,208 @@
vars:
container_name: skudak-cloud
# Install poppler-utils for pdfsig/pdfinfo (LibreSign handles java/pdftk/jsignpdf via occ)
# This needs to be reinstalled on each container recreation
- name: install poppler-utils in skudak-cloud
# ---------------------------------------------------------------------------
# LibreSign (e-signature for Skudak agreements)
#
# poppler-utils supplies pdfsig/pdfinfo; ghostscript is used for PDF
# normalisation. Both land in /usr, which is NOT a persisted volume, so they
# must be reinstalled after every container recreation. Java, PDFtk and
# jSignPdf are different -- LibreSign installs those under
# data/appdata_*/libresign, which IS persisted, so they survive.
- name: install libresign runtime dependencies in skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command:
cmd: >
podman exec -u 0 skudak-cloud
sh -c "apt-get update && apt-get install -y --no-install-recommends
poppler-utils && rm -rf /var/lib/apt/lists/*"
register: poppler_install
changed_when: "'is already the newest version' not in poppler_install.stdout"
poppler-utils ghostscript && rm -rf /var/lib/apt/lists/*"
register: libresign_deps
changed_when: "'is already the newest version' not in libresign_deps.stdout"
# When the container is recreated, the entrypoint re-extracts Nextcloud into
# the /var/www/html volume before Apache starts. Every occ call below races
# that: it fails with "Failed opening required .../lib/versioncheck.php" until
# the extraction completes. Poll until occ answers rather than sleeping a
# fixed interval, which would be both slower and still unreliable.
- name: wait for nextcloud to be ready in skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ status --output=json
register: skudak_occ_ready
until: skudak_occ_ready.rc == 0 and 'installed' in skudak_occ_ready.stdout
retries: 30
delay: 5
changed_when: false
# A disabled app deregisters every `occ libresign:*` command, which makes the
# app look uninstalled rather than switched off. Found disabled on 2026-07-31.
- name: ensure libresign app is enabled in skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ app:enable libresign
register: libresign_enable
changed_when: "'already enabled' not in libresign_enable.stdout"
# Ensure-installed: LibreSign no-ops when the binaries are already present
# under data/appdata_*/libresign (a persisted volume). It prints "Finished with
# success." either way and gives no signal distinguishing a fresh download from
# a no-op, so this never reports changed rather than reporting it every run.
- name: install libresign java/pdftk/jsignpdf binaries in skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ libresign:install --java --pdftk --jsignpdf
register: libresign_install
changed_when: false
failed_when: "'Finished with success' not in libresign_install.stdout"
- name: check whether libresign root certificate is configured
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ libresign:configure:check --certificate
register: libresign_cert_check
changed_when: false
failed_when: false
# Guarded deliberately. Running this unconditionally would mint a new root CA
# on every deploy and invalidate every certificate already issued to a signer,
# breaking the trust chain on documents that were already signed.
#
# Do NOT add --ou here: LibreSign appends its own `libresign-ca-id:...` entry
# to the OU field, and the combined value overruns the 64-character ASN.1
# limit for organizationalUnitName, failing with "string too long".
- name: generate libresign root certificate for skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ libresign:configure:openssl
--cn="{{ libresign_skudak_cert_cn }}"
-o "{{ libresign_skudak_cert_o }}"
-c "{{ libresign_skudak_cert_c }}"
-s "{{ libresign_skudak_cert_st }}"
-l "{{ libresign_skudak_cert_l }}"
when: "'error' in libresign_cert_check.stdout"
changed_when: true
# LibreSign defaults to requiring every signer to upload an identification
# document, which then needs approval by a member of `approval_group` before
# the sign action unlocks. For three partners signing their own partnership
# instruments that is pure friction -- the emailed invitation is the identity
# check. Without this, signers see "Upload file" and no way to sign.
- name: relax libresign identification-document gate in skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ config:app:set libresign identification_documents --value=0
register: libresign_ident
changed_when: "'is now set to' in libresign_ident.stdout"
# SIGNAME_AND_DESCRIPTION (the LibreSign default) typesets the signer's NAME as
# text and offers no drawing surface at all. GRAPHIC is the mode that asks for
# an actual signature graphic -- drawn, uploaded or typed -- and stamps ONLY
# that mark, with no description block.
#
# Deliberately GRAPHIC_ONLY rather than the LibreSign default of
# GRAPHIC_AND_DESCRIPTION. In the latter,
# SignatureTextService::getSignatureWidth() returns `$current / 2` whenever a
# text template is set, splitting the stamp into a graphic half and a text
# half. Our documents already typeset the signer's printed name and the date
# either side of the signature rule (\signatureblock in skudak-contract.cls),
# so LibreSign's own name/date block is both redundant and prone to colliding
# with the drawn mark. GRAPHIC_ONLY takes the early return in that method,
# using the full width to stamp the signature alone.
#
# The value MUST be exactly 'GRAPHIC_ONLY' -- see
# SignerElementsService::RENDER_MODE_GRAPHIC_ONLY (line 25). 'GRAPHIC' is NOT
# a valid constant; setting it writes a value the admin UI cannot match to any
# radio button, silently reverting the effective behaviour to the default.
- name: use signature-only stamp in skudak-cloud libresign
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ config:app:set libresign signature_render_mode --value=GRAPHIC_ONLY
register: libresign_render
changed_when: "'is now set to' in libresign_render.stdout"
# LibreSign stamps a validation footer onto EVERY page of a signed PDF
# (FooterHandler::getFooter(), default on). The QR block within it is a large
# square that lands in the same band as our own document footer -- the rule,
# "Page N of M" and the Skudak mark set by skudak-contract.cls -- and overlaps
# it.
#
# The QR is dropped; the "Digitally signed by ... Validate in <url>" TEXT is
# deliberately KEPT. That line is how a recipient independently verifies who
# signed, when, and under which certificate, which matters for instruments that
# may have to stand up in diligence. Only the redundant graphic goes -- the URL
# it encodes remains printed beside it.
#
# Must be written with --type=boolean: FooterHandler reads it via
# getValueBool() (line 158), and the typed appconfig API does not coerce a
# string "0" to false.
# Lets an email that belongs to an existing Nextcloud account be added as a
# LibreSign signer. Arbitrary external addresses already worked; ONLY
# account-owned ones failed, with a bare "No signers." and nothing logged.
#
# Root cause is in Nextcloud core, not LibreSign --
# lib/private/Collaboration/Collaborators/MailPlugin.php:128-164. On an exact
# email match against the local system address book, with
# shareeEnumerationFullMatch on (its default), the plugin adds a TYPE_USER
# result and returns false. LibreSign registers that plugin as
# MailByMailPlugin with shareType = TYPE_EMAIL, so the TYPE_USER branch is
# skipped, nothing is added, and the early return still fires -- never
# reaching line 243 where the free-form email result is synthesised.
#
# Safe here: shareapi_allow_share_dialog_user_enumeration is already at its
# default 'yes', so users are discoverable by partial search regardless. This
# changes how exact matches are handled, not who can be found.
#
# DO NOT set shareapi_restrict_user_enumeration_full_match_email to 'no'. That
# hits an early bail at MailPlugin.php:67-69 and disables email signer search
# ENTIRELY, including the arbitrary-address case that works today. The verify
# script asserts it has not been set that way.
- name: allow account-owned emails as libresign signers
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ config:app:set core
shareapi_restrict_user_enumeration_full_match --value=no
register: skudak_enum_fullmatch
changed_when: "'is now set to' in skudak_enum_fullmatch.stdout"
- name: drop libresign validation QR code from signed-PDF footer
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ config:app:set libresign write_qrcode_on_footer
--value=0 --type=boolean
register: libresign_qr
changed_when: "'is now set to' in libresign_qr.stdout"
# The whole point of this block. Previously every step ended in `|| echo`, so
# a broken LibreSign deployed clean and stayed broken for six months.
- name: verify libresign configuration in skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ libresign:configure:check
register: libresign_verify
changed_when: false
failed_when: libresign_verify.stdout is search('\berror\b')
- name: disable nextcloud signup link in config
become: true
ansible.builtin.lineinfile:
@@ -131,12 +428,189 @@
changed_when: "'System config value trusted_domains' in trusted_domain_result.stdout"
failed_when: false
# This instance was left at loglevel 0 (DEBUG) and had written a 64 GB
# nextcloud.log, almost entirely repeated deprecation notices. 2 = Warning,
# which is both the Nextcloud default and what the home instance already uses.
- name: set nextcloud loglevel for skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ config:system:set loglevel --value 2 --type integer
register: skudak_loglevel
changed_when: "'System config value loglevel' in skudak_loglevel.stdout"
failed_when: false
# Unbounded by default; see the equivalent task in containers/home/cloud.yml.
- name: cap nextcloud log rotation size for skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ config:system:set log_rotate_size --value 10485760 --type integer
register: skudak_log_rotate
changed_when: "'System config value log_rotate_size' in skudak_log_rotate.stdout"
failed_when: false
# ---------------------------------------------------------------------------
# Skudak mail branding
#
# custom_apps IS a persisted bind mount, so the app survives container
# recreation; only enabling it and the config values need reasserting.
- name: deploy skudakmail email-template app to skudak-cloud
become: true
ansible.builtin.copy:
src: skudakmail/
dest: "{{ cloud_skudak_path }}/apps/skudakmail/"
owner: "{{ podman_subuid.stdout }}"
group: "{{ podman_subuid.stdout }}"
mode: 0644
directory_mode: 0755
notify: restorecon podman
- name: unshare chown skudakmail app
become: true
become_user: "{{ podman_user }}"
changed_when: false
ansible.builtin.command: >
podman unshare chown -R 33:33 {{ cloud_skudak_path }}/apps/skudakmail
- name: enable skudakmail app in skudak-cloud
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud php occ app:enable skudakmail
register: skudakmail_enable
changed_when: "'already enabled' not in skudakmail_enable.stdout"
# Supported extension point -- Mailer::createEMailTemplate() checks this and
# instantiates the named class if it extends EMailTemplate. Not a core patch.
- name: point nextcloud at the skudak email template
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ config:system:set mail_template_class
--value={{ "OCA\\Skudakmail\\Mail\\SkudakEMailTemplate" }}
register: skudak_mail_class
changed_when: "'set to' in skudak_mail_class.stdout"
# Email asset URLs and every link LibreSign puts in a signature invitation are
# built from overwrite.cli.url when sending from a background job. It pointed
# at the pre-rename cloud.skudakrennsport.com, so invitations carried the old
# domain and the logo <img> resolved against it.
- name: set skudak-cloud canonical cli url
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud
php occ config:system:set overwrite.cli.url
--value=https://{{ cloud_skudak_server_name_new }}
register: skudak_cli_url
changed_when: "'set to' in skudak_cli_url.stdout"
# Theming that the email template reads. The logo MUST be a wide, tightly
# cropped image: Mailer clamps to MAX_LOGO_SIZE=105 preserving aspect, so a
# SQUARE logo renders as a 105x105 block in a coloured band -- which is
# exactly how an 8334x8334 upload turned the header into a giant blue blob.
- name: set skudak-cloud theming
become: true
become_user: "{{ podman_user }}"
ansible.builtin.command: >
podman exec -u www-data skudak-cloud php occ theming:config {{ item.k }} "{{ item.v }}"
loop:
- {k: name, v: "Skudak"}
- {k: slogan, v: "Aftermarket vintage car parts and restoration"}
- {k: url, v: "https://skudak.com"}
- {k: primary_color, v: "{{ theming_skudak_primary }}"}
- {k: background_color, v: "{{ theming_skudak_primary }}"}
register: skudak_theming
changed_when: "'Updated' in skudak_theming.stdout"
loop_control:
label: "{{ item.k }}"
# Branding rides on OC\Mail\EMailTemplate, which is Nextcloud's PRIVATE
# namespace with no API stability guarantee. A Nextcloud major upgrade disables
# the app (info.xml pins max-version), Mailer falls back to the stock template,
# and mail keeps sending -- unbranded and silent. This turns that silence into
# a failed play. Renders through Message::useTemplate(), the real path, and
# also re-asserts the LibreSign signing settings.
- name: template skudakmail verification script
become: true
ansible.builtin.template:
src: nextcloud/skudakmail-verify.php.j2
dest: "{{ cloud_skudak_path }}/scripts/skudakmail-verify.php"
owner: "{{ podman_subuid.stdout }}"
group: "{{ podman_subuid.stdout }}"
mode: 0644
notify: restorecon podman
- name: verify skudak mail branding is live
become: true
become_user: "{{ podman_user }}"
ansible.builtin.shell: >
set -o pipefail;
podman exec -i -u www-data skudak-cloud php
< {{ cloud_skudak_path }}/scripts/skudakmail-verify.php
args:
executable: /bin/bash
register: skudakmail_verify
changed_when: false
- include_tasks: containers/cloud-cron.yml
vars:
cron_name: skudak-cloud
cron_container: skudak-cloud
cron_script_path: /usr/local/bin/skudak-cloud-cron.sh
# This instance is BUSINESS data and deliberately backs up to TrueNAS ONLY.
#
# It used to reach personal cloud storage too: the TrueNAS "iDrive E2 Backup"
# cloud-sync task pushes /mnt/glacier to a personal iDrive e2 bucket, which
# swept skudakcloud/ along with it. That task now carries an explicit
# `/skudakcloud/**` exclude, and on 2026-07-30 the stranded copy was purged
# from the bucket -- business data does not belong in personal storage.
#
# The copy was also worthless as a backup: 30 objects against 20,802 files on
# TrueNAS (0.14%), stale since 2026-05-20. Worse, the bucket is VERSIONED and
# the sync runs in COPY mode (never deletes), so every daily run retained
# another ~60 GB version of the pre-cap nextcloud.log -- 56 of them, 3.46 TB,
# 99.3% of a 3.49 TB footprint. Deleting current objects alone reclaims
# nothing on a versioned bucket; the versions must be purged explicitly.
#
# Offsite is now BUSINESS-OWNED: Skudak's own iDrive e2 account, bucket
# `backup-all`, pushed by TrueNAS cloud-sync task "Skudak iDrive - Nextcloud"
# (id 8, /mnt/glacier/skudakcloud -> /skudakcloud, daily 06:00). That bucket
# has a 90-day NoncurrentVersionExpiration policy so the version bloat above
# cannot repeat. The personal task's `/skudakcloud/**` exclude is PERMANENT --
# it is what keeps business data out of personal storage, not a stopgap.
#
# A direct host-to-iDrive S3 stage was built here and then REMOVED on
# 2026-07-31. It would have written the same data into the same `backup-all`
# bucket that the TrueNAS cloud-sync task above already fills -- duplicate
# storage, two writers to one prefix, for no additional coverage. Offsite to
# business-owned storage was already solved by that cloud-sync task; the
# earlier note in this file proposed adding S3 *and then dropping the rsync*,
# i.e. replacement, and building both was a misreading of it.
#
# If offsite is ever moved onto this host, it must REPLACE the rsync below,
# not run beside it. The open question to settle first is whether the
# TrueNAS -> iDrive leg is independently verifiable; nobody has confirmed that
# task's run history end to end, and keeping this chain means trusting it.
- include_tasks: containers/cloud-backup.yml
vars:
backup_name: skudak-cloud
data_path: "{{ cloud_skudak_path }}/data"
config_path: "{{ cloud_skudak_path }}/config"
db_container: skudak-cloud-db
ssh_key_path: /etc/ssh/backup_keys/skudak-cloud
ssh_key_content: "{{ cloud_skudak_backup_ssh_key }}"
ssh_user: skucloud
remote_path: /mnt/glacier/skudakcloud
script_path: /usr/local/bin/skudak-cloud-backup.sh
# skudakcloud/data is mode 770, so the receiving side needs traversable
# dirs. This flag was hand-added on the host and was being silently
# reverted by every `make deploy TAGS=skudak-cloud`; it now lives in git.
backup_rsync_extra_args: "--chmod=Du=rwx,Dgo=rx"
# Staggered so both instances finish before the 05:00 TrueNAS snapshot.
backup_oncalendar: "*-*-* 04:30:00"
+1 -6
View File
@@ -15,12 +15,7 @@
- 443/tcp
# Gitea Skudak SSH
- 2222/tcp
# pihole (unused?)
- 53/tcp
- 53/udp
# nosql/redis
- 6379/tcp
# ???
# BookStack (wiki.skudak.com) -- container publishes 6875:8080
- 6875/tcp
# Satisfactory
- 7777/tcp
+137 -11
View File
@@ -2,6 +2,14 @@
- import_tasks: firewall.yml
- import_tasks: podman/podman.yml
- import_tasks: podman/podman-prune.yml
tags: podman-prune
# Heals the TrueNAS CIFS mounts and bounces immich when they return.
# Must be deployable independently of the photos containers.
- import_tasks: podman/cifs-watchdog.yml
tags: cifs-watchdog, photos
# WEB SERVER: Caddy is the default and only web server
# nginx has been completely replaced and removed
@@ -34,12 +42,6 @@
image: ghcr.io/home-assistant/home-assistant:2026.5.1
tags: hass
- import_tasks: containers/home/partkeepr.yml
vars:
db_image: docker.io/library/mariadb:10.0
image: docker.io/bdebyl/partkeepr:0.1.10
tags: partkeepr
- import_tasks: containers/home/partsy.yml
vars:
image: "git.debyl.io/debyltech/partsy:latest"
@@ -67,24 +69,25 @@
- import_tasks: containers/home/cloud.yml
vars:
db_image: docker.io/library/mariadb:10.6
image: docker.io/library/nextcloud:33.0.0-apache
image: docker.io/library/nextcloud:34.0.2-apache
tags: cloud
- import_tasks: containers/skudak/cloud.yml
vars:
db_image: docker.io/library/mariadb:10.6
image: docker.io/library/nextcloud:33.0.0-apache
redis_image: docker.io/redis:8.2-alpine
image: docker.io/library/nextcloud:34.0.2-apache
tags: skudak, skudak-cloud
- import_tasks: containers/debyltech/fulfillr.yml
vars:
image: git.debyl.io/debyltech/fulfillr:20260723.2044
image: git.debyl.io/debyltech/fulfillr:20260728.2155
tags: debyltech, fulfillr
# Staging back-office (fulfillr-dev.debyltech.com) — same image, staging Turso config.
- import_tasks: containers/debyltech/fulfillr-dev.yml
vars:
image: git.debyl.io/debyltech/fulfillr:20260723.2044
image: git.debyl.io/debyltech/fulfillr:20260728.2155
tags: debyltech, fulfillr-dev
- import_tasks: containers/debyltech/uptime-kuma.yml
@@ -109,7 +112,7 @@
- import_tasks: containers/home/gregtime.yml
vars:
image: localhost/greg-time-bot:3.9.25
image: localhost/greg-time-bot:3.10.0
tags: gregtime
- import_tasks: containers/home/zomboid.yml
@@ -117,3 +120,126 @@
image: docker.io/cm2network/steamcmd:root
tags: zomboid
# ---------------------------------------------------------- Gitea backups
# The Gitea pods themselves are owned by roles/git, but the backup machinery
# (containers/cloud-backup.yml plus templates/nextcloud/*) lives here, and an
# include_tasks reaching across roles would need a path outside the role. So
# the two Gitea backup instances are wired here alongside the container ones.
#
# Both run as ROOTLESS podman under "{{ git_user }}", not "{{ podman_user }}"
# -- hence backup_podman_user -- and use PostgreSQL rather than MariaDB.
#
# Scheduled ahead of the 04:00/04:30 Nextcloud runs so everything lands before
# the 05:00 TrueNAS snapshot.
# `apply` is required: tags on a dynamic include_tasks select whether the
# include runs, but do NOT propagate to the tasks inside it, so without this
# `make deploy TAGS=gitea-backup` includes the file and then filters out every
# task in it. The Nextcloud instances avoid this only because they are reached
# through a static import_tasks chain that tags their children at parse time.
- include_tasks:
file: containers/cloud-backup.yml
apply:
tags: gitea-backup
vars:
backup_name: gitea-debyl
backup_product: Gitea
backup_podman_user: "{{ git_user }}"
data_path: "{{ git_home }}/volumes/gitea/data"
db_container: gitea-debyl-postgres
backup_db_type: postgres
ssh_key_path: /etc/ssh/backup_keys/gitea
ssh_key_content: "{{ gitea_backup_ssh_key }}"
ssh_user: gitea
remote_path: /mnt/glacier/gitea
script_path: /usr/local/bin/gitea-backup.sh
# actions_log/artifacts are CI churn (510 MB and growing) and rebuildable;
# indexers/queues/tmp are derived state Gitea recreates on start. Note the
# default excludes are Nextcloud-specific, so this must be set explicitly.
backup_rsync_excludes: >-
--exclude '/gitea/actions_log' --exclude '/gitea/actions_artifacts'
--exclude '/gitea/tmp' --exclude '/gitea/indexers' --exclude '/gitea/queues'
backup_oncalendar: "*-*-* 03:30:00"
tags: gitea-backup
# ------------------------------------------------- Skudak app-data backups
# BookStack (wiki.skudak.com) and partsy-skudak are BUSINESS data. Both share
# one TrueNAS dataset (/mnt/glacier/skudakapps) so they need only one backup
# user, key and cloud-sync task between them; the personal "iDrive E2 Backup"
# task excludes /skudakapps/** and Skudak's own task pushes it to backup-all.
#
# Scheduled ahead of the 03:30+ Gitea/Nextcloud jobs and the 05:00 snapshot.
- include_tasks:
file: containers/cloud-backup.yml
apply:
tags: [skudak, skudak-apps-backup]
vars:
backup_name: bookstack
backup_product: BookStack
data_path: "{{ bookstack_path }}"
db_container: bookstack-db
# mysql:5.7 predates the mariadb-dump alias -- see cloud-backup.sh.j2.
backup_db_type: mysql
ssh_key_path: /etc/ssh/backup_keys/skudakapps
ssh_key_content: "{{ skudakapps_backup_ssh_key }}"
ssh_user: skudakapps
remote_path: /mnt/glacier/skudakapps/bookstack
script_path: /usr/local/bin/bookstack-backup.sh
# The wiki content is the DATABASE; /mysql is its raw datadir, which must
# not be rsynced live -- the dump above is the consistent copy. public/
# and storage/ hold the uploads and are the only file trees worth shipping.
backup_rsync_excludes: "--exclude '/mysql'"
backup_oncalendar: "*-*-* 03:00:00"
tags: skudak, skudak-apps-backup
- include_tasks:
file: containers/cloud-backup.yml
apply:
tags: [skudak, skudak-apps-backup]
vars:
backup_name: partsy-skudak
backup_product: Partsy
data_path: "{{ partsy_skudak_path }}"
# Live WAL-mode SQLite: snapshotted via `sqlite3 .backup` rather than
# rsynced, so the -wal/-shm sidecars are deliberately excluded from the
# file tree -- shipping them alongside a separately-taken snapshot would
# only invite a confusing restore.
backup_sqlite_dbs:
- "{{ partsy_skudak_path }}/data/partsy.db"
backup_rsync_excludes: "--exclude '*-wal' --exclude '*-shm'"
ssh_key_path: /etc/ssh/backup_keys/skudakapps
ssh_key_content: "{{ skudakapps_backup_ssh_key }}"
ssh_user: skudakapps
remote_path: /mnt/glacier/skudakapps/partsy-skudak
script_path: /usr/local/bin/partsy-skudak-backup.sh
backup_oncalendar: "*-*-* 03:10:00"
tags: skudak, skudak-apps-backup
# BUSINESS data. Rsynced to TrueNAS here, then pushed offsite to SKUDAK'S OWN
# iDrive e2 account (bucket `backup-all`) by TrueNAS cloud-sync task "Skudak
# iDrive - Gitea" (id 9, /mnt/glacier/skudakgit -> /skudakgit, daily 06:30).
#
# The personal "iDrive E2 Backup" task's /skudakgit/** exclude is PERMANENT:
# it is what keeps business data out of personal storage. Do not remove it --
# Skudak has its own task and bucket instead.
- include_tasks:
file: containers/cloud-backup.yml
apply:
tags: [skudak, gitea-backup-skudak]
vars:
backup_name: skudak-gitea
backup_product: Gitea
backup_podman_user: "{{ git_user }}"
data_path: "{{ git_home }}/volumes/gitea-skudak/data"
db_container: gitea-skudak-postgres
backup_db_type: postgres
ssh_key_path: /etc/ssh/backup_keys/skudak-gitea
ssh_key_content: "{{ skudakgit_backup_ssh_key }}"
ssh_user: skudakgit
remote_path: /mnt/glacier/skudakgit
script_path: /usr/local/bin/skudak-gitea-backup.sh
backup_rsync_excludes: >-
--exclude '/gitea/actions_log' --exclude '/gitea/actions_artifacts'
--exclude '/gitea/tmp' --exclude '/gitea/indexers' --exclude '/gitea/queues'
backup_oncalendar: "*-*-* 03:45:00"
tags: skudak, gitea-backup-skudak
@@ -0,0 +1,36 @@
---
- name: template cifs watchdog script
become: true
ansible.builtin.template:
src: cifs-watchdog.sh.j2
dest: /usr/local/bin/cifs-watchdog.sh
owner: root
group: root
mode: 0755
setype: bin_t
- name: template cifs watchdog systemd service
become: true
ansible.builtin.template:
src: cifs-watchdog.service.j2
dest: /etc/systemd/system/cifs-watchdog.service
owner: root
group: root
mode: 0644
- name: template cifs watchdog systemd timer
become: true
ansible.builtin.template:
src: cifs-watchdog.timer.j2
dest: /etc/systemd/system/cifs-watchdog.timer
owner: root
group: root
mode: 0644
- name: enable and start cifs watchdog timer
become: true
ansible.builtin.systemd:
name: cifs-watchdog.timer
enabled: true
state: started
daemon_reload: true
@@ -0,0 +1,36 @@
---
- name: template podman prune script
become: true
ansible.builtin.template:
src: podman-prune.sh.j2
dest: /usr/local/bin/podman-prune.sh
owner: root
group: root
mode: 0755
setype: bin_t
- name: template podman prune systemd service
become: true
ansible.builtin.template:
src: podman-prune.service.j2
dest: /etc/systemd/system/podman-prune.service
owner: root
group: root
mode: 0644
- name: template podman prune systemd timer
become: true
ansible.builtin.template:
src: podman-prune.timer.j2
dest: /etc/systemd/system/podman-prune.timer
owner: root
group: root
mode: 0644
- name: enable and start podman prune timer
become: true
ansible.builtin.systemd:
name: podman-prune.timer
enabled: true
state: started
daemon_reload: true
@@ -16,7 +16,11 @@
# Logging
log {
output file /var/log/caddy/caddy.log
output file /var/log/caddy/caddy.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format {{ caddy_log_format }}
level {{ caddy_log_level }}
}
@@ -73,7 +77,11 @@
reverse_proxy localhost:8088
log {
output file /var/log/caddy/photos.log
output file /var/log/caddy/photos.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -90,7 +98,11 @@
reverse_proxy localhost:6875
log {
output file /var/log/caddy/wiki.log
output file /var/log/caddy/wiki.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -122,7 +134,11 @@
}
log {
output file /var/log/caddy/assistant.log
output file /var/log/caddy/assistant.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -154,7 +170,11 @@
}
log {
output file /var/log/caddy/parts.log
output file /var/log/caddy/parts.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -164,7 +184,11 @@
import common_headers
reverse_proxy localhost:8082
log {
output file /var/log/caddy/parts-skudak.log
output file /var/log/caddy/parts-skudak.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -182,7 +206,11 @@
}
log {
output file /var/log/caddy/uptime-kuma.log
output file /var/log/caddy/uptime-kuma.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -200,7 +228,11 @@
}
log {
output file /var/log/caddy/uptime-kuma-personal.log
output file /var/log/caddy/uptime-kuma-personal.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -243,7 +275,11 @@
}
log {
output file /var/log/caddy/graylog.log
output file /var/log/caddy/graylog.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -281,7 +317,11 @@
redir /.well-known/caldav /remote.php/dav 301
log {
output file /var/log/caddy/cloud.log
output file /var/log/caddy/cloud.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -309,7 +349,11 @@
redir /.well-known/caldav /remote.php/dav 301
log {
output file /var/log/caddy/cloud-skudak.log
output file /var/log/caddy/cloud-skudak.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -323,7 +367,11 @@
}
log {
output file /var/log/caddy/gitea-debyl.log
output file /var/log/caddy/gitea-debyl.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -337,7 +385,11 @@
}
log {
output file /var/log/caddy/gitea-skudak.log
output file /var/log/caddy/gitea-skudak.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -384,7 +436,11 @@
}
log {
output file /var/log/caddy/fulfillr.log
output file /var/log/caddy/fulfillr.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -431,7 +487,11 @@
}
log {
output file /var/log/caddy/fulfillr-dev.log
output file /var/log/caddy/fulfillr-dev.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
}
}
@@ -453,7 +513,11 @@ test.debyl.io {
header Pragma "no-cache"
log {
output file /var/log/caddy/test.log
output file /var/log/caddy/test.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format json
level {{ caddy_log_level }}
}
@@ -0,0 +1,8 @@
# {{ ansible_managed }}
# CIFS credentials for the TrueNAS shares consumed by immich.
#
# This file exists so the password stops living in /etc/fstab, which is
# world-readable (0644) by design -- every local user could read the SMB
# credential straight out of it. Mounted via credentials= instead.
username={{ cifs_username }}
password={{ cifs_password }}
@@ -0,0 +1,13 @@
[Unit]
Description=Verify TrueNAS CIFS mounts and heal immich after recovery
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
ExecStart=/usr/local/bin/cifs-watchdog.sh
# Type=oneshot disables the start timeout by default. Every external call in
# the script is already wrapped in `timeout`, but bound the unit too so a
# wedged run cannot block every subsequent trigger.
TimeoutStartSec={{ cifs_watchdog_timeout | default('5m') }}
Nice=10
@@ -0,0 +1,97 @@
#!/bin/bash
# {{ ansible_managed }}
# Keep the TrueNAS CIFS mounts healthy and heal immich when they come back.
#
# Why this exists at all: the immich containers are systemd USER units under
# "{{ podman_user }}", while the mounts are SYSTEM units. A user unit cannot
# declare RequiresMountsFor= against a system mount, so there is no native way
# to say "restart immich when this mount returns". This bridges the two scopes.
#
# The failure it addresses: TrueNAS was power-cycled, the .mount units failed,
# and systemd never retried -- mount units are not restarted on failure. SMB
# came back, nothing remounted, and immich-server kept serving an empty library
# while its database still listed 14k assets.
#
# Deliberately NOT `set -e`: one unhealthy mount must not stop the others from
# being checked or healed. Same reasoning as nextcloud-backup-alert.sh.j2.
set -uo pipefail
TAG=cifs-watchdog
STATE_DIR=/var/lib/cifs-watchdog
log() { logger -t "$TAG" -p daemon.info -- "$*"; echo "$TAG: $*"; }
warn() { logger -t "$TAG" -p daemon.err -- "$*"; echo "$TAG: $*" >&2; }
install -d -m 0755 "$STATE_DIR"
healed=0
for mp in {{ cifs_watchdog_mounts | map('quote') | join(' ') }}; do
key="$STATE_DIR/$(systemd-escape -p "$mp")"
prev="unknown"
[ -f "$key" ] && prev="$(cat "$key" 2>/dev/null)"
# `mountpoint -q` alone is NOT sufficient: a CIFS mount whose server vanished
# stays "mounted" while every read returns EIO. The directory listing is the
# real health check. timeout guards against a hang despite `soft`.
if timeout 15 ls -1 "$mp" >/dev/null 2>&1 && mountpoint -q "$mp" 2>/dev/null; then
state=healthy
else
state=unhealthy
fi
recovered=0
if [ "$state" = unhealthy ]; then
warn "mount=$mp status=unhealthy action=recovering"
unit="$(systemd-escape -p --suffix=mount "$mp")"
automount="$(systemd-escape -p --suffix=automount "$mp")"
# A unit left in `failed` refuses to start again until it is reset.
systemctl reset-failed "$unit" "$automount" 2>/dev/null
# Start the .mount unit DIRECTLY -- do not try to trigger the automount by
# listing the path. An unmounted mount point is still an ordinary empty
# directory, so `ls` succeeds and tells us nothing; relying on it silently
# skipped recovery entirely in testing.
timeout 40 systemctl start "$unit" 2>/dev/null
# Re-arm the on-access trigger too, so a drop between watchdog ticks is
# repaired by the next process that touches the path rather than waiting.
systemctl start "$automount" 2>/dev/null
if timeout 15 ls -1 "$mp" >/dev/null 2>&1 && mountpoint -q "$mp" 2>/dev/null; then
state=healthy
recovered=1
log "mount=$mp status=recovered"
else
warn "mount=$mp status=still-unhealthy"
fi
fi
# Restart when the mount became healthy after being down -- either across
# ticks (prev=unhealthy) or WITHIN this run (recovered=1). The second case is
# not redundant: a drop that is detected and repaired by the same invocation
# leaves prev=healthy, so checking only the stored state silently skips the
# restart while the container still holds its stale, empty view.
#
# Still gated on actually reaching healthy, so a run that fails to recover
# does not bounce immich on every tick while the NAS is down.
if [ "$state" = healthy ] && { [ "$prev" = unhealthy ] || [ "$recovered" -eq 1 ]; }; then
healed=1
fi
echo "$state" > "$key"
done
if [ "$healed" -eq 1 ]; then
# ONLY immich-server: it is the sole consumer of both CIFS paths. Postgres,
# redis and machine-learning use local volumes and must not be bounced.
#
# Rootless podman under "{{ podman_user }}": -H for HOME, the `cd;` preamble
# is required (see CLAUDE.md), and XDG_RUNTIME_DIR for systemctl --user.
if sudo -H -u {{ podman_user }} bash -c \
'cd; export XDG_RUNTIME_DIR=/run/user/$(id -u)
systemctl --user restart {{ cifs_watchdog_restart_unit | default("immich-server.service") }}' 2>/dev/null; then
log "status=healed action=restarted unit={{ cifs_watchdog_restart_unit | default('immich-server.service') }}"
else
warn "status=healed action=restart-FAILED unit={{ cifs_watchdog_restart_unit | default('immich-server.service') }}"
fi
fi
exit 0
@@ -0,0 +1,13 @@
[Unit]
Description=Periodic TrueNAS CIFS mount health check
[Timer]
OnBootSec={{ cifs_watchdog_onbootsec | default('2m') }}
OnCalendar={{ cifs_watchdog_oncalendar }}
RandomizedDelaySec=30s
# Not Persistent: this is a liveness check, so a run missed while the host was
# off carries no meaning -- the next tick reflects current reality.
Persistent=false
[Install]
WantedBy=timers.target
@@ -1,6 +1,14 @@
[Unit]
Description=Nextcloud {{ instance_name }} backup to TrueNAS
Description={{ backup_product | default('Nextcloud') }} {{ instance_name }} backup to TrueNAS
After=network-online.target
Wants=network-online.target
OnFailure=nextcloud-backup-failed@%n.service
[Service]
Type=oneshot
ExecStart={{ script_path }}
# Type=oneshot disables the start timeout by default, so a wedged rsync would
# leave the unit "activating" forever and every subsequent daily trigger would
# be silently skipped. Bound it.
TimeoutStartSec={{ backup_timeout | default('4h') }}
Nice=10
@@ -1,4 +1,174 @@
#!/bin/bash
# {{ ansible_managed }}
# {{ backup_product | default('Nextcloud') }} "{{ backup_name }}" -> truenas.localdomain.
#
# Ordering is deliberate: the database is dumped BEFORE the file tree is
# synced. A DB snapshot slightly OLDER than the files degrades to "files
# the app has not indexed yet" and is repaired with a rescan. A DB snapshot
# NEWER than the files references blobs that never made it into the backup,
# which surfaces as broken shares and dead file entries on restore.
set -euo pipefail
rsync -az --exclude .ssh -e "ssh -i {{ ssh_key_path }} -o StrictHostKeyChecking=accept-new" \
{{ data_path }}/ {{ ssh_user }}@truenas.localdomain:{{ remote_path }}/
# Tag stays "nextcloud-backup" for EVERY instance, Gitea included: an external
# Graylog rule matches status=failed on this tag, and renaming it here would
# silently stop alerting for all of them. Misnomer retained deliberately.
TAG=nextcloud-backup
INSTANCE={{ backup_name }}
STAGE={{ backup_stage_path | default('/var/backups/nextcloud/' ~ backup_name) }}
KEEP={{ backup_db_keep | default(7) }}
DUMP="$STAGE/db/${INSTANCE}-$(date +%Y%m%d).sql.gz"
log() { logger -t "$TAG" -p daemon.info -- "instance=$INSTANCE $*"; echo "$TAG: $*"; }
fail() { logger -t "$TAG" -p daemon.err -- "instance=$INSTANCE status=failed $*"
echo "$TAG: FAILED: $*" >&2; exit 1; }
SSH="ssh -i {{ ssh_key_path }} -o StrictHostKeyChecking=accept-new -o ServerAliveInterval=30 -o ServerAliveCountMax=6"
DEST={{ ssh_user }}@truenas.localdomain
log "status=start"
{% if db_container | default('') %}
# ------------------------------------------------------------ 1. database
# The containers are ROOTLESS podman owned by
# "{{ backup_podman_user | default(podman_user) }}" -- Nextcloud runs under
# {{ podman_user }}, Gitea under its own git user -- but this script runs as
# root under systemd. Every podman call therefore goes through sudo:
# -H HOME becomes the podman user's home, so podman finds its rootless
# graph root under ~/.local/share/containers
# cd; required preamble (see CLAUDE.md) so the shell starts in that home
# XDG_RUNTIME_DIR the podman user's runtime dir. Lingering is enabled by
# roles/podman/tasks/podman/podman.yml so /run/user/<uid> exists;
# guarded anyway so podman falls back cleanly if it ever does not.
#
# No credential is stored in this file or placed on a host command line:
# $MYSQL_ROOT_PASSWORD / $MYSQL_DATABASE (or $POSTGRES_* for postgres) are
# expanded by the shell INSIDE the database container, which already carries
# them in its environment.
{% if backup_db_type | default('mariadb') != 'postgres' %}
#
# MariaDB-specific history, for the flag set below:
# --routines and --events initially aborted the dump here: mysql.proc read as
# corrupted (error 1728) and the event scheduler reported disabled (1577),
# both artefacts of an image bump without mariadb-upgrade. `mariadb-upgrade
# --force` has since been run against both instances and repaired the system
# tables, so the full flag set works and is kept for completeness.
#
# Note that upgrade still exits non-zero on these containers: it cannot
# install the `sys` schema because the datadir root (/var/lib/mysql) is owned
# by daemon rather than mysql, so mysqld may not create new top-level
# databases. `sys` is purely diagnostic and unused by Nextcloud, so this is
# cosmetic -- but it does mean creating a NEW database would fail too.
{% else %}
#
# --clean --if-exists makes the dump replayable into an existing database
# (`psql -U gitea gitea < dump.sql`) rather than only into an empty one, and
# --no-owner keeps it restorable when the target role names differ.
{% endif %}
pexec() {
sudo -H -u {{ backup_podman_user | default(podman_user) }} bash -c \
'cd; d=/run/user/$(id -u); [ -d "$d" ] && export XDG_RUNTIME_DIR="$d"
exec podman "$@"' _ "$@"
}
install -d -m 0700 "$STAGE" "$STAGE/db"
tmp="$DUMP.tmp"
rm -f "$tmp"
log "dumping {{ db_container }}"
set +e
{% if backup_db_type | default('mariadb') == 'postgres' %}
pexec exec {{ db_container }} sh -c '
exec env PGPASSWORD="$POSTGRES_PASSWORD" pg_dump -U "$POSTGRES_USER" \
--no-owner --clean --if-exists "$POSTGRES_DB"
' | gzip -6 > "$tmp"
{% elif backup_db_type | default('mariadb') == 'mysql' %}
{# Upstream mysql:5.7 ships mysqldump, NOT the mariadb-dump alias that only
appears in MariaDB 10.5+. Flags and completion trailer are identical to the
MariaDB branch; 5.7.21 was verified to accept --no-tablespaces. #}
pexec exec {{ db_container }} sh -c '
exec env MYSQL_PWD="$MYSQL_ROOT_PASSWORD" mysqldump -u root \
--single-transaction --quick --routines --events --triggers \
--no-tablespaces --default-character-set=utf8mb4 "$MYSQL_DATABASE"
' | gzip -6 > "$tmp"
{% else %}
pexec exec {{ db_container }} sh -c '
exec env MYSQL_PWD="$MYSQL_ROOT_PASSWORD" mariadb-dump -u root \
--single-transaction --quick --routines --events --triggers \
--no-tablespaces --default-character-set=utf8mb4 "$MYSQL_DATABASE"
' | gzip -6 > "$tmp"
{% endif %}
dump_rc=${PIPESTATUS[0]}
set -e
[ "$dump_rc" -eq 0 ] || fail "{% if backup_db_type | default('mariadb') == 'postgres' %}pg_dump{% elif backup_db_type | default('mariadb') == 'mysql' %}mysqldump{% else %}mariadb-dump{% endif %} {{ db_container }} exited $dump_rc"
# A new dump is promoted over yesterday's only after it proves complete: a
# valid gzip stream AND the trailer the dump tool writes only on a clean
# finish. The two engines word it differently, so the string is per-type --
# grepping for the MariaDB one against a pg_dump would fail every run.
# `mv` is atomic within the staging filesystem, so a failed or truncated run
# can never replace a good dump.
gzip -t "$tmp" || fail "dump is not a valid gzip stream"
gunzip -c "$tmp" | tail -c 512 \
| grep -q '{{ "PostgreSQL database dump complete" if backup_db_type | default("mariadb") == "postgres" else "Dump completed" }}' \
|| fail "dump is truncated (no completion trailer)"
mv -f "$tmp" "$DUMP"
log "db_dump=ok bytes=$(stat -c %s "$DUMP")"
# Local retention. The staging sync below mirrors with --delete, so remote
# retention follows the same window; deeper history comes from the TrueNAS
# periodic ZFS snapshots (see roles/podman/README.md).
ls -1t "$STAGE"/db/"$INSTANCE"-*.sql.gz | tail -n +$((KEEP + 1)) | xargs -r rm -f
{% endif %}
{% if backup_sqlite_dbs | default([]) %}
# ----------------------------------------------------- 1b. sqlite snapshots
# These databases are WAL-mode SQLite written by a LIVE process, so neither
# rsync option is correct: skipping the -wal loses committed transactions,
# and copying .db + -wal together can catch a checkpoint mid-write. Only
# `.backup` takes a consistent snapshot of a database that is being written.
#
# Run on the HOST, not in the container: the files live on the host
# filesystem and sqlite3 is installed there, so this needs no podman at all
# and works even when the owning container is stopped.
install -d -m 0700 "$STAGE" "$STAGE/db"
{% for db in backup_sqlite_dbs %}
sqlite_out="$STAGE/db/{{ db | basename | regex_replace('\\.db$', '') }}.sqlite"
if [ ! -f "{{ db }}" ]; then
fail "sqlite source missing: {{ db }}"
fi
sqlite3 "{{ db }}" ".backup '$sqlite_out.tmp'" \
|| fail "sqlite3 .backup failed for {{ db }}"
# Prove the snapshot is a loadable database before it replaces yesterday's,
# mirroring the gzip/trailer gate the SQL dumps get above.
sqlite3 "$sqlite_out.tmp" 'pragma integrity_check;' | grep -qx ok \
|| fail "sqlite snapshot failed integrity_check: {{ db }}"
mv -f "$sqlite_out.tmp" "$sqlite_out"
log "sqlite_dump=ok db={{ db | basename }} bytes=$(stat -c %s "$sqlite_out")"
{% endfor %}
{% endif %}
# ----------------------------------------------------------- 2. file tree
# --exclude .ssh is load-bearing: {{ remote_path }} IS {{ ssh_user }}'s home
# on TrueNAS and its authorized_keys lives there, so a `.ssh` directory
# appearing in the data tree must never be shipped. No --delete here: the
# data tree is append-mostly and a source-side mishap must not propagate.
log "syncing data"
rsync -az --timeout=1800 --exclude .ssh \
{{ backup_rsync_excludes | default("--exclude '/nextcloud.log*' --exclude '/updater.log'") }} \
{{ backup_rsync_extra_args | default('') }} \
-e "$SSH" {{ data_path }}/ "$DEST:{{ remote_path }}/"
# ------------------------------------------- 3. config/ and database dumps
{# --mkpath creates the nested _backup/<x>/ destination; rsync will not build
more than one missing level on its own. Requires rsync >= 3.2.3 on both
ends (galactica 3.4.1, truenas 3.2.7). #}
{% if config_path | default('') %}
log "syncing config"
rsync -az --timeout=600 --delete --mkpath {{ backup_rsync_extra_args | default('') }} \
-e "$SSH" {{ config_path }}/ "$DEST:{{ remote_path }}/_backup/config/"
{% endif %}
{% if db_container | default('') or backup_sqlite_dbs | default([]) %}
log "syncing db dumps"
rsync -az --timeout=600 --delete --mkpath {{ backup_rsync_extra_args | default('') }} \
-e "$SSH" "$STAGE/db/" "$DEST:{{ remote_path }}/_backup/db/"
{% endif %}
log "status=ok"
@@ -1,8 +1,9 @@
[Unit]
Description=Daily Nextcloud {{ instance_name }} backup
Description=Daily {{ backup_product | default('Nextcloud') }} {{ instance_name }} backup
[Timer]
OnCalendar=*-*-* 04:00:00
OnCalendar={{ backup_oncalendar | default('*-*-* 04:00:00') }}
RandomizedDelaySec={{ backup_randomized_delay | default('5m') }}
Persistent=true
[Install]
@@ -0,0 +1,13 @@
[Unit]
Description=Nextcloud {{ instance_name }} background jobs
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
ExecStart={{ cron_script_path }}
# Type=oneshot disables the start timeout by default, so a wedged cron.php
# would leave the unit "activating" forever and every subsequent 5-minute
# trigger would be silently skipped. Bound it.
TimeoutStartSec={{ cron_timeout | default('30m') }}
Nice=10
@@ -0,0 +1,51 @@
#!/bin/bash
# {{ ansible_managed }}
# Nextcloud "{{ cron_name }}" background jobs (cron.php).
#
# backgroundjobs_mode is "cron" on both instances, which means Nextcloud
# expects an external caller to run cron.php every ~5 minutes. Nothing was:
# the personal instance had not run a background job since 2026-05-14 and
# skudak since 2024-11-20. Without it Nextcloud never expires trash or file
# versions, never cleans stale chunked uploads, never sends calendar
# reminders, and -- easy to miss -- never rotates nextcloud.log, which makes
# the log_rotate_size cap set in containers/*/cloud.yml inert.
set -euo pipefail
TAG=nextcloud-cron
INSTANCE={{ cron_name }}
log() { logger -t "$TAG" -p daemon.info -- "instance=$INSTANCE $*"; }
fail() { logger -t "$TAG" -p daemon.err -- "instance=$INSTANCE status=failed $*"
echo "$TAG: FAILED: $*" >&2; exit 1; }
# The Nextcloud containers are ROOTLESS podman owned by "{{ podman_user }}",
# but this script runs as root under systemd. Same sudo/cd/XDG_RUNTIME_DIR
# preamble as cloud-backup.sh -- see CLAUDE.md for why `cd;` is required.
pexec() {
sudo -H -u {{ podman_user }} bash -c \
'cd; d=/run/user/$(id -u); [ -d "$d" ] && export XDG_RUNTIME_DIR="$d"
exec podman "$@"' _ "$@"
}
# A container that is down (deploy, image bump, host reboot) is not a failure
# worth flagging -- the next tick picks it up five minutes later.
if ! pexec container exists {{ cron_container }} 2>/dev/null; then
log "status=skipped reason=container-absent"
exit 0
fi
# occ and cron.php both refuse to do anything useful mid-upgrade. Skipping
# keeps a deploy window from parading as a run of failed units.
if pexec exec -u www-data {{ cron_container }} php occ status 2>/dev/null \
| grep -q 'maintenance: true'; then
log "status=skipped reason=maintenance"
exit 0
fi
set +e
pexec exec -u www-data {{ cron_container }} php -f /var/www/html/cron.php
rc=$?
set -e
[ "$rc" -eq 0 ] || fail "cron.php exited $rc"
log "status=ok"
@@ -0,0 +1,13 @@
[Unit]
Description=Nextcloud {{ instance_name }} background jobs every 5 minutes
[Timer]
OnBootSec={{ cron_onbootsec | default('5m') }}
OnUnitActiveSec={{ cron_interval | default('5m') }}
RandomizedDelaySec={{ cron_randomized_delay | default('30s') }}
# Deliberately NOT Persistent: this runs every 5 minutes, so replaying runs
# missed while the host was off buys nothing and just stampedes at boot.
Persistent=false
[Install]
WantedBy=timers.target
@@ -1,22 +0,0 @@
#!/bin/bash
# LibreSign dependency setup for Skudak Nextcloud
# Runs on container start via /docker-entrypoint-hooks.d/before-starting/
# Note: This runs as www-data, not root. poppler-utils is installed
# separately via Ansible using podman exec -u 0.
echo "=== LibreSign Setup: Installing dependencies ==="
# Install LibreSign-managed Java (required for PDFtk and jSignPdf)
# This downloads a specific Java version that LibreSign validates
echo "Installing Java..."
php /var/www/html/occ libresign:install --java || echo "Java install skipped or failed"
# Install PDFtk (requires Java)
echo "Installing PDFtk..."
php /var/www/html/occ libresign:install --pdftk || echo "PDFtk install skipped or failed"
# Install jSignPdf (requires Java)
echo "Installing jSignPdf..."
php /var/www/html/occ libresign:install --jsignpdf || echo "jSignPdf install skipped or failed"
echo "=== LibreSign Setup: Complete ==="
@@ -0,0 +1,69 @@
#!/bin/bash
# {{ ansible_managed }}
# OnFailure= handler for every backup unit (Nextcloud and Gitea). Invoked as:
# nextcloud-backup-alert.sh <failed-unit-name>
#
# One shared copy serves all instances -- cloud-backup.yml renders it once and
# later includes are no-ops -- so the wording here must not name one product.
# The failing UNIT name is what identifies the instance.
#
# Deliberately NOT `set -e`: an alert handler that dies partway through
# reports nothing, which is worse than a partial report. Same reasoning as
# roles/ups/templates/ups-restore.sh.j2.
set -uo pipefail
TAG=nextcloud-backup
UNIT="${1:-unknown}"
TO="{{ backup_alert_email | default('root') }}"
HOST="$(hostname -f 2>/dev/null || hostname)"
result="$(systemctl show -p Result --value "$UNIT" 2>/dev/null)"
code="$(systemctl show -p ExecMainStatus --value "$UNIT" 2>/dev/null)"
# Only claim a failure when the unit actually reports one. Starting this
# handler by hand (or any other spurious trigger) would otherwise mail out a
# subject line saying FAILED about a run that succeeded. Keeping status=failed
# exact also stops such triggers matching the Graylog alert rule.
if [ "${result:-success}" = "success" ]; then
state=spurious
prio=daemon.warning
headline="$(printf 'Backup alert handler was invoked on %s, but %s reports SUCCESS.\nThis is not a backup failure -- most likely the handler was started manually.' "$HOST" "$UNIT")"
subject="[$HOST] Backup alert (spurious, unit OK): $UNIT"
else
state=failed
prio=daemon.err
headline="Backup FAILED on $HOST"
subject="[$HOST] Backup FAILED: $UNIT"
fi
# One machine-parseable line for Graylog, then the context.
logger -t "$TAG" -p "$prio" -- \
"status=$state unit=$UNIT result=${result:-unknown} exit=${code:-unknown}"
body="$(printf '%s\n\nunit: %s\nresult: %s\nexit: %s\n\n--- last 40 journal lines ---\n' \
"$headline" "$UNIT" "${result:-unknown}" "${code:-unknown}")
$(journalctl -u "$UNIT" -n 40 --no-pager -o cat 2>/dev/null)"
echo "$body" | logger -t "$TAG" -p "$prio"
# Only genuine failures are worth an email. A spurious invocation carries no
# action for a human, and mailing it trains the reader to ignore the subject
# line -- which defeats the point of having the alert at all. The journald
# record above is kept either way, so spurious triggers stay greppable.
if [ "$state" != "failed" ]; then
logger -t "$TAG" -p daemon.info -- "alert_mail=skipped reason=$state"
exit 0
fi
# Mail is best-effort: if the MTA is not configured the journald record above
# is still the authoritative signal, so never fail the handler on this.
if command -v sendmail >/dev/null 2>&1; then
printf 'To: %s\nSubject: %s\nContent-Type: text/plain; charset=UTF-8\n\n%s\n' \
"$TO" "$subject" "$body" | sendmail -t \
&& logger -t "$TAG" -p daemon.info -- "alert_mail=sent to=$TO" \
|| logger -t "$TAG" -p daemon.err -- "alert_mail=failed to=$TO"
else
logger -t "$TAG" -p daemon.err -- "alert_mail=skipped reason=no-sendmail"
fi
exit 0
@@ -0,0 +1,6 @@
[Unit]
Description=Report failure of %i
[Service]
Type=oneshot
ExecStart=/usr/local/bin/nextcloud-backup-alert.sh %i
@@ -0,0 +1,40 @@
# {{ ansible_managed }}
#
# Redis for skudak-cloud: Nextcloud distributed cache + transactional file
# locking. Reachable only by container name on the `shared` podman network --
# no host port is published.
#
# The password lives HERE rather than on the command line as
# `redis-server --requirepass <pass>`. That is the existing house idiom (see
# the deleted container-nosql.yml in git history), but it leaks the secret into
# `podman inspect`, into the generated systemd unit under
# ~/.config/systemd/user/, and into `ps` for every user on the host. A 0640
# config file mounted read-only keeps it out of all three.
requirepass {{ cloud_skudak_redis_pass }}
# Bind to all interfaces WITHIN the container's network namespace. The
# container publishes no port, so this is reachable only from the `shared`
# podman network -- not from the host and not from the LAN.
bind 0.0.0.0
port 6379
protected-mode yes
# NO maxmemory / eviction policy, deliberately.
#
# Nextcloud puts BOTH the distributed cache and the transactional file locks in
# this instance. Cache entries are safely evictable; LOCKS ARE NOT. An
# `allkeys-lru` policy under memory pressure can evict a lock that a live
# request still believes it holds, which permits concurrent writers to the same
# file -- silent corruption rather than a visible error. With no maxmemory,
# Redis never evicts. The host has ~14 GiB free of 31 GiB and this instance
# holds a few hundred keys, so a cap buys nothing.
#
# If a cap is ever genuinely needed, use `maxmemory-policy noeviction` so Redis
# returns an error instead of silently discarding a lock.
# No persistence. Locks are ephemeral and TTL-bounded, and the cache is
# rebuildable -- there is nothing here worth surviving a restart. Persisting
# would be actively worse: a restored RDB could reinstate locks whose owning
# request died, blocking files until the TTL expired.
save ""
appendonly no
@@ -0,0 +1,186 @@
<?php
/**
* {{ ansible_managed }}
*
* Post-deploy assertion that Skudak mail branding is actually live.
*
* WHY THIS EXISTS: SkudakEMailTemplate extends OC\Mail\EMailTemplate, which is
* Nextcloud's PRIVATE namespace -- no API stability guarantee. Two things can
* silently switch the branding off:
*
* 1. A Nextcloud major upgrade. appinfo/info.xml pins max-version, so the app
* is auto-disabled as incompatible; Mailer::createEMailTemplate() then
* fails its class_exists() check and falls back to the stock template.
* Mail still sends -- unbranded. That is the right failure mode, but it is
* invisible without this check.
* 2. An upstream change to the private base class breaking an override.
*
* Renders through Message::useTemplate() -- the REAL path -- rather than
* calling renderHtml() directly. That distinction is not academic: renderText()
* runs first and flips the parent's footerAdded flag, and a renderHtml()-only
* test once passed green while live mail shipped with no footer at all.
*
* Exits non-zero with a diagnostic on any failure, so the Ansible task fails
* the play rather than reporting a clean deploy over broken branding.
*/
require_once '/var/www/html/lib/base.php';
$mailer = \OC::$server->get(\OCP\Mail\IMailer::class);
$dispatcher = \OC::$server->get(\OCP\EventDispatcher\IEventDispatcher::class);
// Mirrors MailService::notifyUnsignedUser() (custom_apps/libresign/lib/Service/MailService.php:85-116).
$template = $mailer->createEMailTemplate('settings.TestEmail');
$template->setSubject('LibreSign: There is a file for you to sign');
$template->addHeader();
$template->addHeading('File to sign', false);
$template->addBodyText('There is a document for you to sign. Access the link below:');
$template->addBodyButton('Sign »verify.pdf«', 'https://{{ cloud_skudak_server_name_new }}/verify');
$message = $mailer->createMessage();
$message->setTo(['verify@example.invalid' => 'Verify']);
$message->useTemplate($template);
// What Mailer::send() does at lib/private/Mail/Mailer.php:186. Nothing is sent.
$dispatcher->dispatchTyped(new \OCP\Mail\Events\BeforeMessageSent($message));
$html = $message->getSymfonyEmail()->getHtmlBody() ?? '';
$text = $message->getPlainBody();
$subject = $message->getSubject();
$inlineNames = [];
foreach ($message->getSymfonyEmail()->getAttachments() as $part) {
$inlineNames[] = (string)$part->getFilename();
}
$failures = [];
if (!$template instanceof \OCA\Skudakmail\Mail\SkudakEMailTemplate) {
$failures[] = 'template class is ' . get_class($template)
. ' -- expected SkudakEMailTemplate. Is the skudakmail app enabled, and does '
. 'appinfo/info.xml still allow this Nextcloud major?';
}
if (!str_starts_with($subject, 'Skudak LLP')) {
$failures[] = 'subject not rewritten: ' . $subject;
}
if (!str_contains($html, 'official document-signing request')) {
$failures[] = 'HTML footer missing (renderText/renderHtml ordering regression?)';
}
if (!str_contains($text, 'official document-signing request')) {
$failures[] = 'plain-text footer missing';
}
if (!str_contains($html, 'skudak.com/privacy') || !str_contains($html, 'skudak.com/terms')) {
$failures[] = 'privacy/terms links missing from footer';
}
if (preg_match('/[»«]/u', $html)) {
$failures[] = 'German guillemets survived into the body';
}
if (!str_contains($html, 'Review document')) {
$failures[] = 'button label not normalised to "Review document"';
}
if (!str_contains($html, 'cid:skudak-wordmark.png')) {
$failures[] = 'logo is not a cid: reference -- BeforeMessageSent listener did not fire';
}
if (!in_array('skudak-wordmark.png', $inlineNames, true)) {
$failures[] = 'inline logo MIME part absent (found: ' . (implode(', ', $inlineNames) ?: 'none') . ')';
}
// ---------------------------------------------------------------------------
// LibreSign signing settings. These live in oc_appconfig (the database), not on
// disk, so they survive container recreation -- but they are re-assertable and
// a stray click in the admin UI can change them silently. GRAPHIC in particular
// matters: any other mode makes SignatureTextService::getSignatureWidth()
// return $current / 2 and stamp a name/date block that duplicates -- and
// collides with -- the one our documents already typeset.
$appConfig = \OC::$server->get(\OCP\IAppConfig::class);
// Must be exactly GRAPHIC_ONLY -- SignerElementsService::RENDER_MODE_GRAPHIC_ONLY.
// The valid set is DESCRIPTION_ONLY / SIGNAME_AND_DESCRIPTION /
// GRAPHIC_AND_DESCRIPTION / GRAPHIC_ONLY. Anything outside it (a bare 'GRAPHIC',
// say) is accepted by occ but matches no radio in the admin UI and falls
// through to default behaviour, so this asserts membership, not just non-empty.
$renderMode = $appConfig->getValueString('libresign', 'signature_render_mode', '');
if ($renderMode !== 'GRAPHIC_ONLY') {
$failures[] = 'libresign signature_render_mode is "' . $renderMode
. '" -- expected GRAPHIC_ONLY (signature only). Any other mode halves the '
. 'stamp width and overlays a duplicate name/date block.';
}
// Read with getValueBool, exactly as FooterHandler:158 does -- asserting the
// string form would pass on a value the app itself reads as true.
if ($appConfig->getValueBool('libresign', 'write_qrcode_on_footer', true) !== false) {
$failures[] = 'libresign write_qrcode_on_footer is not false -- the validation '
. 'QR block will be stamped on every page and overlaps the document footer '
. 'set by skudak-contract.cls. (Was it written without --type=boolean?)';
}
// Signer search for account-owned emails. Both keys are asserted because the
// two failure modes are opposite and the second is the more dangerous:
// full_match = yes -> account-owned emails silently unselectable
// full_match_email = no -> email signer search disabled ENTIRELY
// Defaults are 'yes' for both (MailPlugin.php:50-55), so an unset
// full_match_email is correct and only an explicit 'no' is a problem.
if ($appConfig->getValueString('core', 'shareapi_restrict_user_enumeration_full_match', 'yes') !== 'no') {
$failures[] = 'core shareapi_restrict_user_enumeration_full_match is not "no" -- '
. 'emails belonging to an existing Nextcloud account cannot be added as '
. 'LibreSign signers (MailPlugin.php:163 aborts the search).';
}
if ($appConfig->getValueString('core', 'shareapi_restrict_user_enumeration_full_match_email', 'yes') === 'no') {
$failures[] = 'core shareapi_restrict_user_enumeration_full_match_email is "no" -- '
. 'this disables email signer search ENTIRELY (MailPlugin.php:67). It must be '
. 'unset or "yes"; it is NOT the knob for the account-owned-email problem.';
}
$identDocs = $appConfig->getValueString('libresign', 'identification_documents', '');
if ($identDocs !== '0') {
$failures[] = 'libresign identification_documents is "' . $identDocs
. '" -- expected 0. A non-zero value gates signing behind an ID upload '
. 'plus admin approval, and signers see no way to sign.';
}
// ---------------------------------------------------------------------------
// Redis: distributed cache + transactional file locking.
//
// These come from the image's config/redis.config.php drop-in, which only
// activates when REDIS_HOST is set on the container. If the env var is lost
// (a container recreated from a stale spec, say), Nextcloud silently reverts
// to DBLockingProvider and every file lock goes back to being a MariaDB write
// -- functional, but the stalls come back with no error anywhere.
$sysConfig = \OC::$server->get(\OCP\IConfig::class);
foreach (['memcache.locking', 'memcache.distributed'] as $key) {
$value = $sysConfig->getSystemValueString($key, '');
if ($value !== '\OC\Memcache\Redis') {
$failures[] = $key . ' is "' . $value . '" -- expected \\OC\\Memcache\\Redis. '
. 'Is REDIS_HOST still set on the skudak-cloud container?';
}
}
// Prove Redis is actually reachable and authenticating, not merely configured.
// A wrong password leaves the config looking perfect while every cache and
// lock operation fails at runtime.
try {
$cacheFactory = \OC::$server->get(\OCP\ICacheFactory::class);
if (!$cacheFactory->isAvailable()) {
$failures[] = 'distributed cache reports unavailable -- redis unreachable or auth failed';
} else {
$probe = $cacheFactory->createDistributed('skudakmail-verify');
$probe->set('probe', 'ok', 30);
if ($probe->get('probe') !== 'ok') {
$failures[] = 'distributed cache round-trip failed (set/get mismatch)';
}
$probe->remove('probe');
}
} catch (\Throwable $e) {
$failures[] = 'distributed cache threw: ' . $e->getMessage();
}
if ($failures !== []) {
fwrite(STDERR, "skudakmail branding verification FAILED:\n");
foreach ($failures as $f) {
fwrite(STDERR, " - $f\n");
}
exit(1);
}
echo "skudakmail branding OK (subject: $subject)\n";
@@ -0,0 +1,15 @@
[Unit]
Description=Prune unused podman images and volumes
After=network-online.target
[Service]
Type=oneshot
ExecStart=/usr/local/bin/podman-prune.sh
# Type=oneshot disables the start timeout by default, so a wedged prune would
# leave the unit "activating" forever and every later trigger would be
# silently skipped -- the same trap the backup units guard against.
TimeoutStartSec={{ podman_prune_timeout | default('1h') }}
# Pruning is I/O-heavy and competes with the 03:00-04:30 backups; keep it off
# the critical path.
Nice=15
IOSchedulingClass=idle
@@ -0,0 +1,53 @@
#!/bin/bash
# {{ ansible_managed }}
# Weekly reclaim of unused podman images and volumes.
#
# Every image bump leaves the previous tag behind and nothing ever removed
# them: this was written after finding 896 images totalling 59.6 GB, 75% of it
# unused -- 94 tags of greg-time-bot and 73 of fulfillr, one per deploy.
#
# --filter until={{ podman_prune_until }} keeps recent images so a rollback
# does not require a rebuild or re-pull. Anything older is unused AND stale.
#
# Volumes pruned here are podman's ANONYMOUS volumes, not the bind mounts
# under {{ podman_volumes }} that hold real service data -- those are
# directories on the host and podman does not know about them. The dangling
# ones seen in practice were 804 MB copies of Nextcloud's /var/www/html left
# by container recreations, which are image content, not data.
set -uo pipefail
TAG=podman-prune
log() { logger -t "$TAG" -p daemon.info -- "$*"; echo "$TAG: $*"; }
total_before=0
total_after=0
for u in {{ podman_prune_users | join(' ') }}; do
# Rootless podman: -H so HOME points at the user's store, and the `cd;`
# preamble is required (see CLAUDE.md) or podman cannot find its graph root.
run() {
sudo -H -u "$u" bash -c \
'cd; d=/run/user/$(id -u); [ -d "$d" ] && export XDG_RUNTIME_DIR="$d"
exec podman "$@"' _ "$@"
}
if ! id "$u" >/dev/null 2>&1; then
log "user=$u status=skipped reason=no-such-user"
continue
fi
before=$(run system df --format '{{ '{{' }}.Size{{ '}}' }}' 2>/dev/null | head -1)
# Deliberately NOT `set -e`: a prune failing for one user must not stop the
# other, and a busy image is a normal, non-fatal outcome.
img=$(run image prune -af --filter "until={{ podman_prune_until }}" 2>&1 | tail -1)
vol=$(run volume prune -f 2>&1 | tail -1)
after=$(run system df --format '{{ '{{' }}.Size{{ '}}' }}' 2>/dev/null | head -1)
log "user=$u images_before=$before images_after=$after"
log "user=$u image_prune=${img:-none} volume_prune=${vol:-none}"
done
log "status=ok"
@@ -0,0 +1,12 @@
[Unit]
Description=Weekly podman image and volume prune
[Timer]
OnCalendar={{ podman_prune_oncalendar | default('Sun *-*-* 02:00:00') }}
RandomizedDelaySec=15m
# Weekly, and growth is one tag per deploy, so a missed run is worth catching
# up on rather than skipping.
Persistent=true
[Install]
WantedBy=timers.target
+114
View File
@@ -0,0 +1,114 @@
# ups
UPS monitoring and staged shutdown for the home rack.
A CyberPower PR1500RT2U (`0764:0601`) is cabled by USB to `home.debyl.io` and
backs both that host and `truenas.localdomain` (Dell PowerEdge R415). This role
makes `home.debyl.io` the NUT server and gives it the ability to power TrueNAS
back on over IPMI.
## Outage sequence
| When | What happens | Driven by |
| --- | --- | --- |
| t+0 | UPS goes on battery, `ONBATT` logged to journald → Graylog | `upsmon` |
| t+2min | TrueNAS shuts itself down cleanly, shedding ~200 W | TrueNAS UPS service, Slave mode, `Shutdown Timer 120` |
| 10% charge | `home.debyl.io` shuts itself down and tells the UPS to cut its output | `upsmon` `SHUTDOWNCMD` + `/lib/systemd/system-shutdown/nutshutdown` |
| mains returns | UPS re-energizes; both machines power themselves back up | R415 `always-on` restore policy; `home.debyl.io` BIOS *After Power Loss → Power On* |
| mains back +5min | Best-effort IPMI power-on, if a dedicated iDRAC is ever fitted | `upssched``ups-restore.sh` |
| host boot | Same restore check, for the deep-drain case | `ups-restore.service` |
The 10% threshold is not a custom poller. CyberPower asserts its own low-battery
flag around 2035%, so `ups.conf` sets `ignorelb` plus
`override.battery.charge.low`, and stock `upsmon` fires at exactly the
configured percentage.
Likewise, the 2-minute TrueNAS shed is TrueNAS's own native "shutdown timer"
setting — no SSH key and no shutdown script from this side.
## Powering TrueNAS back on
Neither Wake-on-LAN nor IPMI works on this box today, so restore is done with
the chassis power restore policy instead.
**Wake-on-LAN is out.** `bce0` advertises no WOL capability (`ifconfig -m bce0`
has no `WOL_MAGIC`) and both NICs are bonded into an LACP `lagg0`.
**IPMI is out too, for now.** The iDRAC is in `shared / LOM1` mode and the
Enterprise card that would provide a dedicated management port is not fitted:
```
$ ipmitool sdr elist | grep -i idrac
iDRAC6 Ent Pres | 70h | ok | 7.1 | Absent
```
A shared-LOM iDRAC6 Express has no standby power. Measured directly: with the
chassis powered off the BMC does not even answer ARP, and it only reappears
~200s into POST, at the moment the host brings the NIC link up. That is a
hardware limitation, not a switch or BIOS problem.
**So restore works like this instead.** `ipmitool chassis policy always-on` is
set on the R415. On a deep outage `home.debyl.io` halts at 10% and NUT's
shutdown hook tells the UPS to cut its output; when mains returns the UPS
re-energizes, the R415 sees AC and boots itself.
The gap is the medium outage — mains returns after TrueNAS has shed but before
the battery reaches 10%. The UPS never cuts power, so TrueNAS stays off and
needs a manual power button press. Fitting a used iDRAC6 Enterprise card and
running `ipmitool delloem lan set dedicated` (plus a cable to the dedicated
port) closes that gap, and `ups-restore.sh` starts working with no code
changes — it is already deployed and simply logs and exits while the BMC is
unreachable.
## One-time setup outside Ansible
These are not managed by this role.
**iDRAC (already done, via `ipmitool` on TrueNAS):**
```
ipmitool lan set 1 ipsrc static
ipmitool lan set 1 ipaddr 192.168.1.12
ipmitool lan set 1 netmask 255.255.255.0
ipmitool lan set 1 defgw ipaddr 192.168.1.1
ipmitool lan set 1 access on # was disabled - nothing answers without this
ipmitool channel setaccess 1 2 callin=on ipmi=on link=on privilege=4
ipmitool user set password 2 '<idrac_password>'
ipmitool chassis policy always-on # this is what restores power after an outage
```
`idrac_password` is also the iDRAC web UI password for `root` - they share a
user database.
**TrueNAS UI → Services → UPS** (enable + start automatically):
| Field | Value |
| --- | --- |
| UPS Mode | Slave |
| Remote Host | `192.168.1.10` |
| Remote Port | `3493` |
| Identifier | `cyberpower` |
| Monitor User | `truenas` |
| Monitor Password | vault `nut_truenas_password` |
| Shutdown Mode | UPS goes on battery |
| Shutdown Timer | `120` |
| Shutdown Command | `/sbin/shutdown -p now` |
| Power Off UPS | unchecked |
**BIOS on `home.debyl.io`** (Lenovo 10MR0004US): Power → *After Power Loss*
**Power On**. On a full drain the NUT shutdown hook
(`/lib/systemd/system-shutdown/nutshutdown`) tells the UPS to cut its own
output; this BIOS setting is what brings the host back when mains returns.
## Vault keys
`nut_upsmon_password`, `nut_truenas_password`, `idrac_password`
## Operating
```
upsc cyberpower # full UPS status
upsc cyberpower battery.charge
sudo /usr/local/bin/truenas-power.sh status|on|soft|off
journalctl -t ups-restore -t ups-sched
```
+36
View File
@@ -0,0 +1,36 @@
---
# CyberPower PR1500LCDRT2U cabled by USB to this host (0764:0601).
# This host is the NUT server; truenas.localdomain is a NUT slave.
ups_name: cyberpower
# lsusb reports the product string as PR1500LCDRT2U, but the device itself
# reports device.model CP1500PFCRM2U. The latter is what it actually is.
ups_desc: CyberPower CP1500PFCRM2U
ups_vendorid: "0764"
ups_productid: "0601"
ups_listen_addr: 192.168.1.10
ups_listen_port: 3493
# truenas.localdomain - the only host allowed to reach upsd
ups_slave_ip: 192.168.1.11
# Battery charge at which THIS host shuts itself down. TrueNAS sheds much
# earlier via its own "shutdown timer" setting (see roles/ups/README.md).
ups_low_charge_pct: 10
# Seconds of stable mains after ONLINE before TrueNAS is powered back on.
ups_restore_stable_secs: 300
# R415 iDRAC6. Shares LOM1 with bce0.
idrac_host: 192.168.1.12
idrac_user: root
ups_deps:
[
ipmitool,
nut,
nut-client,
]
# Secrets live in ansible/vars/vault.yml (no vault_ prefix, per repo
# convention): nut_upsmon_password, nut_truenas_password, idrac_password
+35
View File
@@ -0,0 +1,35 @@
---
# Regenerates the nut-driver@<name> unit instances from ups.conf. Oneshot,
# so "restarted" just means "run it again".
- name: reload nut driver units
become: true
ansible.builtin.systemd:
name: nut-driver-enumerator.service
state: restarted
daemon_reload: true
# The enumerator only writes unit definitions; it will not pick up changed
# driver options in an already-running driver. Restart the instance itself.
- name: restart nut driver
become: true
ansible.builtin.systemd:
name: "nut-driver@{{ ups_name }}.service"
state: restarted
- name: restart nut server
become: true
ansible.builtin.systemd:
name: nut-server.service
state: restarted
- name: restart nut monitor
become: true
ansible.builtin.systemd:
name: nut-monitor.service
state: restarted
- name: restart firewalld
become: true
ansible.builtin.systemd:
name: firewalld
state: restarted
+7
View File
@@ -0,0 +1,7 @@
---
- name: install NUT and ipmitool
become: true
ansible.builtin.dnf:
name: "{{ ups_deps }}"
state: present
tags: ups
+12
View File
@@ -0,0 +1,12 @@
---
- name: allow NUT from the truenas slave only
become: true
ansible.posix.firewalld:
rich_rule: >-
rule family="ipv4" source address="{{ ups_slave_ip }}/32"
port port="{{ ups_listen_port }}" protocol="tcp" accept
permanent: true
immediate: true
state: enabled
notify: restart firewalld
tags: [ups, firewall]
+68
View File
@@ -0,0 +1,68 @@
---
# IPMI power control for truenas.localdomain (Dell R415, iDRAC6).
# WOL is not an option there: bce0 advertises no WOL capability and is an
# LACP lagg member, so IPMI is the only remote power-on path.
- name: deploy iDRAC credential
become: true
ansible.builtin.copy:
content: "{{ idrac_password }}\n"
dest: /etc/ups/idrac.pw
owner: root
group: nut
mode: '0640'
no_log: true
tags: ups
- name: deploy truenas power helper
become: true
ansible.builtin.template:
src: truenas-power.sh.j2
dest: /usr/local/bin/truenas-power.sh
owner: root
group: root
mode: '0755'
setype: bin_t
tags: ups
- name: deploy truenas restore script
become: true
ansible.builtin.template:
src: ups-restore.sh.j2
dest: /usr/local/bin/ups-restore.sh
owner: root
group: root
mode: '0755'
setype: bin_t
tags: ups
- name: deploy upssched command dispatcher
become: true
ansible.builtin.template:
src: ups-sched-cmd.sh.j2
dest: /usr/local/bin/ups-sched-cmd.sh
owner: root
group: root
mode: '0755'
setype: bin_t
tags: ups
# Covers the deep-drain case: if the battery ran out, this host was itself
# powered off when mains returned, so nothing was running to restore TrueNAS.
- name: deploy boot-time truenas restore unit
become: true
ansible.builtin.template:
src: ups-restore.service.j2
dest: /etc/systemd/system/ups-restore.service
owner: root
group: root
mode: '0644'
tags: ups
- name: enable boot-time truenas restore unit
become: true
ansible.builtin.systemd:
name: ups-restore.service
enabled: true
daemon_reload: true
tags: ups
+5
View File
@@ -0,0 +1,5 @@
---
- import_tasks: deps.yml
- import_tasks: nut.yml
- import_tasks: ipmi.yml
- import_tasks: firewall.yml
+129
View File
@@ -0,0 +1,129 @@
---
# NUT server. The UPS is on USB here, so this host drives it and serves
# status to truenas.localdomain over the network.
- name: set NUT run mode to netserver
become: true
ansible.builtin.template:
src: nut.conf.j2
dest: /etc/ups/nut.conf
owner: root
group: nut
mode: '0640'
notify:
- reload nut driver units
- restart nut driver
- restart nut server
- restart nut monitor
tags: ups
- name: deploy NUT driver configuration
become: true
ansible.builtin.template:
src: ups.conf.j2
dest: /etc/ups/ups.conf
owner: root
group: nut
mode: '0640'
notify:
- reload nut driver units
- restart nut driver
- restart nut server
tags: ups
- name: deploy upsd listener configuration
become: true
ansible.builtin.template:
src: upsd.conf.j2
dest: /etc/ups/upsd.conf
owner: root
group: nut
mode: '0640'
notify: restart nut server
tags: ups
- name: deploy upsd users
become: true
ansible.builtin.template:
src: upsd.users.j2
dest: /etc/ups/upsd.users
owner: root
group: nut
mode: '0640'
notify:
- restart nut server
- restart nut monitor
tags: ups
- name: deploy upsmon configuration
become: true
ansible.builtin.template:
src: upsmon.conf.j2
dest: /etc/ups/upsmon.conf
owner: root
group: nut
mode: '0640'
notify: restart nut monitor
tags: ups
- name: deploy upssched configuration
become: true
ansible.builtin.template:
src: upssched.conf.j2
dest: /etc/ups/upssched.conf
owner: root
group: nut
mode: '0640'
notify: restart nut monitor
tags: ups
# The nut package ships /usr/lib/udev/rules.d/62-nut-usbups.rules, which
# hands the UPS USB device to the nut user. Reload and retrigger so the
# driver can claim it without physically replugging the UPS.
- name: reload udev rules for NUT USB access
become: true
ansible.builtin.command:
cmd: udevadm control --reload-rules
changed_when: false
tags: ups
- name: retrigger UPS USB device
become: true
ansible.builtin.command:
cmd: >-
udevadm trigger --subsystem-match=usb
--attr-match=idVendor={{ ups_vendorid }}
changed_when: false
tags: ups
# nut-driver-enumerator reads ups.conf and generates nut-driver@{{ ups_name }}.
# The .service is oneshot (so never "started" for long) and is triggered by
# the .path unit watching ups.conf - enable both, but only start the .path.
- name: enable NUT driver enumerator
become: true
ansible.builtin.systemd:
name: nut-driver-enumerator.service
enabled: true
daemon_reload: true
tags: ups
- name: enable and start NUT driver enumerator path trigger
become: true
ansible.builtin.systemd:
name: nut-driver-enumerator.path
enabled: true
state: started
tags: ups
- name: enable and start NUT services
become: true
ansible.builtin.systemd:
name: "{{ item }}"
enabled: true
state: started
loop:
- "nut-driver@{{ ups_name }}.service"
- nut-server.service
- nut-monitor.service
- nut.target
tags: ups
+3
View File
@@ -0,0 +1,3 @@
# {{ ansible_managed }}
# netserver: run the driver + upsd locally and serve slaves over the network.
MODE=netserver
@@ -0,0 +1,34 @@
#!/bin/bash
# {{ ansible_managed }}
# Remote power control for truenas.localdomain via the R415 iDRAC6.
# Usage: truenas-power.sh {status|on|soft|off|cycle}
set -euo pipefail
PW_FILE=/etc/ups/idrac.pw
if [ ! -r "$PW_FILE" ]; then
echo "cannot read $PW_FILE" >&2
exit 1
fi
# -f keeps the password out of the process argument list.
ipmi() {
ipmitool -I lanplus -H {{ idrac_host }} -U {{ idrac_user }} \
-f "$PW_FILE" "$@"
}
case "${1:-status}" in
status) ipmi chassis power status ;;
on) ipmi chassis power on ;;
# ACPI soft-off. FreeBSD has hw.acpi.power_button_state=S5, so this is
# a clean TrueNAS shutdown. Normally unused: TrueNAS shuts itself down
# as a NUT slave. This is the manual escape hatch.
soft) ipmi chassis power soft ;;
# Hard cut, last resort only.
off) ipmi chassis power off ;;
cycle) ipmi chassis power cycle ;;
*)
echo "usage: $0 {status|on|soft|off|cycle}" >&2
exit 2
;;
esac
@@ -0,0 +1,13 @@
# {{ ansible_managed }}
[Unit]
Description=Restore TrueNAS power after an outage
After=network-online.target nut-server.service nut-monitor.service
Wants=network-online.target
Requires=nut-server.service
[Service]
Type=oneshot
ExecStart=/usr/local/bin/ups-restore.sh
[Install]
WantedBy=multi-user.target
@@ -0,0 +1,63 @@
#!/bin/bash
# {{ ansible_managed }}
# Power truenas.localdomain back on after an outage, but only when it is
# genuinely safe to do so. Called by upssched once mains has been stable,
# and once at boot by ups-restore.service (deep-drain recovery).
#
# Deliberately stateless: we ask the iDRAC whether the chassis is off
# rather than tracking whether we were the ones who shut it down. A NAS
# powered off by hand while on mains is safe, because no ONLINE event
# fires in that case and the boot path checks UPS status first.
#
# NOTE: as of now this is a best-effort secondary path. The R415 has no
# iDRAC6 Enterprise card ("iDRAC6 Ent Pres ... Absent"), so its shared-LOM
# BMC has no standby power and goes unreachable whenever the chassis is
# off - exactly when we would want it. An unreachable iDRAC is therefore
# the EXPECTED case here, and we exit quietly rather than alarming.
#
# The primary restore path needs no IPMI: the R415 power restore policy is
# set to always-on, so when the UPS cuts and then restores its output the
# server powers itself back up. Fit an iDRAC6 Enterprise card and switch it
# to dedicated mode and this script starts working with no changes.
set -uo pipefail
TAG=ups-restore
UPS={{ ups_name }}@localhost
log() { logger -t "$TAG" -- "$*"; echo "$TAG: $*"; }
# At boot, upsd may not be serving yet. Wait a bounded amount of time.
for _ in $(seq 1 30); do
if upsc "$UPS" ups.status >/dev/null 2>&1; then
break
fi
sleep 2
done
status=$(upsc "$UPS" ups.status 2>/dev/null || echo UNKNOWN)
if [[ "$status" != *OL* ]]; then
log "UPS status is '$status', not on line - refusing to power TrueNAS on"
exit 0
fi
power=$(/usr/local/bin/truenas-power.sh status 2>&1 || true)
case "$power" in
*"is off"*)
log "mains stable and chassis off - powering TrueNAS on"
if /usr/local/bin/truenas-power.sh on; then
log "power-on command accepted"
else
log "power-on command FAILED - check iDRAC at {{ idrac_host }}"
exit 1
fi
;;
*"is on"*)
log "TrueNAS already on, nothing to do"
;;
*)
# Expected while the box has no iDRAC6 Enterprise card: the BMC is
# simply not on the network with the chassis powered down.
log "iDRAC at {{ idrac_host }} unreachable - relying on the" \
"always-on power restore policy instead ($power)"
;;
esac
@@ -0,0 +1,14 @@
#!/bin/bash
# {{ ansible_managed }}
# upssched CMDSCRIPT. Runs as the unprivileged nut user.
set -uo pipefail
case "${1:-}" in
truenas-restore)
exec /usr/local/bin/ups-restore.sh
;;
*)
logger -t ups-sched -- "unknown timer '${1:-}'"
exit 1
;;
esac
+18
View File
@@ -0,0 +1,18 @@
# {{ ansible_managed }}
[{{ ups_name }}]
driver = usbhid-ups
port = auto
vendorid = {{ ups_vendorid }}
productid = {{ ups_productid }}
desc = "{{ ups_desc }}"
# CyberPower asserts its own low-battery flag around 20-35%, far too
# early for us. Ignore it and derive LB from our own threshold so
# upsmon fires SHUTDOWNCMD at exactly {{ ups_low_charge_pct }}%.
ignorelb
override.battery.charge.low = {{ ups_low_charge_pct }}
# Unlock driver.killpower. Without this the driver refuses to cut UPS
# output at shutdown, and the R415's always-on power restore policy
# would never see AC drop and return - i.e. nothing comes back after a
# deep outage. See README.md.
allow_killpower
+3
View File
@@ -0,0 +1,3 @@
# {{ ansible_managed }}
LISTEN 127.0.0.1 {{ ups_listen_port }}
LISTEN {{ ups_listen_addr }} {{ ups_listen_port }}
+11
View File
@@ -0,0 +1,11 @@
# {{ ansible_managed }}
# Local upsmon on this host.
[upsmon]
password = {{ nut_upsmon_password }}
upsmon master
# truenas.localdomain, running the TrueNAS UPS service in Slave mode.
[truenas]
password = {{ nut_truenas_password }}
upsmon slave
@@ -0,0 +1,40 @@
# {{ ansible_managed }}
MONITOR {{ ups_name }}@localhost 1 upsmon {{ nut_upsmon_password }} master
MINSUPPLIES 1
SHUTDOWNCMD "/usr/bin/systemctl poweroff"
NOTIFYCMD /usr/bin/upssched
# upsmon drops this file before halting; /lib/systemd/system-shutdown/nutshutdown
# reads it late in shutdown and, if present, tells the UPS to cut its output.
# That AC drop-and-return is what triggers the R415's always-on restore policy
# and this host's BIOS "After Power Loss: Power On". upsmon has NO compiled-in
# default for this - leave it unset and the UPS never powers down.
# Must be on tmpfs: a persistent path can go stale and make every ordinary
# reboot look like a forced shutdown.
POWERDOWNFLAG /run/nut/killpower
POLLFREQ 5
POLLFREQALERT 5
# Wait up to 30s for the truenas slave to disconnect before we halt.
HOSTSYNC 30
DEADTIME 15
RBWARNTIME 43200
NOCOMMWARNTIME 300
FINALDELAY 5
# SYSLOG puts every UPS event in the journal, which fluent-bit already
# forwards to Graylog (see roles/common/tasks/fluent-bit.yml).
# EXEC runs NOTIFYCMD, i.e. upssched, which drives the TrueNAS restore.
NOTIFYFLAG ONLINE SYSLOG+EXEC
NOTIFYFLAG ONBATT SYSLOG+EXEC
NOTIFYFLAG LOWBATT SYSLOG+EXEC
NOTIFYFLAG FSD SYSLOG+EXEC
NOTIFYFLAG COMMOK SYSLOG+EXEC
NOTIFYFLAG COMMBAD SYSLOG+EXEC
NOTIFYFLAG SHUTDOWN SYSLOG+EXEC
NOTIFYFLAG REPLBATT SYSLOG
NOTIFYFLAG NOCOMM SYSLOG
NOTIFYFLAG NOPARENT SYSLOG
@@ -0,0 +1,13 @@
# {{ ansible_managed }}
CMDSCRIPT /usr/local/bin/ups-sched-cmd.sh
PIPEFN /run/nut/upssched.pipe
LOCKFN /run/nut/upssched.lock
# Mains is back: wait for it to hold for {{ ups_restore_stable_secs }}s
# before powering TrueNAS back on, so we do not flap on unstable power.
AT ONLINE * CANCEL-TIMER truenas-restore
AT ONLINE * START-TIMER truenas-restore {{ ups_restore_stable_secs }}
# Power dropped again while the restore timer was pending - stand down.
AT ONBATT * CANCEL-TIMER truenas-restore
Binary file not shown.