back up Gitea + Skudak app data; drop PartKeepr and Pi-hole

Extends the Nextcloud backup machinery rather than adding a second
mechanism. cloud-backup.sh.j2 gains three guarded options, all no-ops for
the existing callers:

  backup_podman_user  Gitea runs rootless under `git`, not `podman`
  backup_db_type      postgres (Gitea) and mysql (BookStack) alongside
                      mariadb; each engine's completion trailer differs,
                      and grepping for the wrong one fails every run
  backup_sqlite_dbs   `sqlite3 .backup` for live WAL-mode SQLite, gated on
                      `pragma integrity_check` before promotion -- rsync
                      is either stale (no -wal) or torn (with it)

New instances: gitea-debyl, skudak-gitea, bookstack, partsy-skudak. The
alert handler is rendered once and shared, so its wording is now generic
rather than per-product; TAG stays nextcloud-backup because an external
Graylog rule matches on it.

`apply:` on the includes is load-bearing -- tags on a dynamic
include_tasks do not reach the tasks inside it.

Business data (skudak-gitea, bookstack, partsy-skudak) goes to TrueNAS
and on to Skudak's own iDrive account; the personal bucket's
/skudak*/** excludes are permanent, not a stopgap.

Removals: PartKeepr is superseded by Partsy, and its teardown never
finished -- it targeted /etc/systemd/system/podman-partkeepr*.service,
wrong prefix and wrong scope, leaving enabled user units in failed state.
Pi-hole's role was already orphaned (absent from deploy_home.yml); its
port 53 rule went with it after confirming nothing listens there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Bastian de Byl
2026-07-30 18:08:47 -04:00
parent a63bf5edec
commit bc110ce69e
15 changed files with 236 additions and 129 deletions
@@ -1,8 +1,12 @@
#!/bin/bash
# {{ ansible_managed }}
# OnFailure= handler for the Nextcloud backup units. Invoked as:
# OnFailure= handler for every backup unit (Nextcloud and Gitea). Invoked as:
# nextcloud-backup-alert.sh <failed-unit-name>
#
# One shared copy serves all instances -- cloud-backup.yml renders it once and
# later includes are no-ops -- so the wording here must not name one product.
# The failing UNIT name is what identifies the instance.
#
# Deliberately NOT `set -e`: an alert handler that dies partway through
# reports nothing, which is worse than a partial report. Same reasoning as
# roles/ups/templates/ups-restore.sh.j2.
@@ -23,13 +27,13 @@ code="$(systemctl show -p ExecMainStatus --value "$UNIT" 2>/dev/null)"
if [ "${result:-success}" = "success" ]; then
state=spurious
prio=daemon.warning
headline="$(printf 'Nextcloud backup alert handler was invoked on %s, but %s reports SUCCESS.\nThis is not a backup failure -- most likely the handler was started manually.' "$HOST" "$UNIT")"
subject="[$HOST] Nextcloud backup alert (spurious, unit OK): $UNIT"
headline="$(printf 'Backup alert handler was invoked on %s, but %s reports SUCCESS.\nThis is not a backup failure -- most likely the handler was started manually.' "$HOST" "$UNIT")"
subject="[$HOST] Backup alert (spurious, unit OK): $UNIT"
else
state=failed
prio=daemon.err
headline="Nextcloud backup FAILED on $HOST"
subject="[$HOST] Nextcloud backup FAILED: $UNIT"
headline="Backup FAILED on $HOST"
subject="[$HOST] Backup FAILED: $UNIT"
fi
# One machine-parseable line for Graylog, then the context.