Files
deploy_home/ansible/roles/podman/templates/podman-prune.sh.j2
T
Bastian de Byl f11391b28f bound log growth and reclaim ~52 GB of container disk
Caddy was rotating on implicit defaults (100MiB/keep 10/90d) that were not
holding -- 20 rotated files per stream and a 190-day-old .gz, 1.3 GB across
16 log streams. Made explicit at 10MiB/keep 3/7d.

Note roll_size et al are subdirectives of `output file`, NOT of `log`.
Getting that wrong does not degrade gracefully: Caddy refuses to start on a
bad config, so every site went down until it was corrected. Worth a
`caddy validate` gate before reload.

journald had no SystemMaxUse and had reached 4 GB, drifting toward its
10%-of-filesystem default (~190 GB on this root). Capped at 500M.

Both are safe to keep short because fluent-bit ships the journal and every
Caddy access log into Graylog -- though note its GELF output has been
erroring for days, which weakens that premise and wants investigating.

The larger find was unrelated to logs: 896 images totalling 59.6 GB with
75% unused (94 tags of greg-time-bot, 73 of fulfillr -- one per deploy) and
5.4 GB of dangling volumes, mostly 804 MB Nextcloud /var/www/html trees
orphaned by container recreations. Pruned to 22 images / 15.4 GB, and added
a weekly timer keeping 30 days so a rollback still needs no rebuild.

Also dropped the decommissioned 6379/tcp redis rule (nothing listening;
Immich's redis is on the shared podman network) and the orphaned nosql, s3
and searxng volume dirs.

Backup log exclusions turned out to be unnecessary: Gitea logs to console
so its log dirs are empty, Nextcloud already excludes its own, BookStack
mounts only uploads, and Caddy is not backed up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:13:35 -04:00

54 lines
2.0 KiB
Django/Jinja

#!/bin/bash
# {{ ansible_managed }}
# Weekly reclaim of unused podman images and volumes.
#
# Every image bump leaves the previous tag behind and nothing ever removed
# them: this was written after finding 896 images totalling 59.6 GB, 75% of it
# unused -- 94 tags of greg-time-bot and 73 of fulfillr, one per deploy.
#
# --filter until={{ podman_prune_until }} keeps recent images so a rollback
# does not require a rebuild or re-pull. Anything older is unused AND stale.
#
# Volumes pruned here are podman's ANONYMOUS volumes, not the bind mounts
# under {{ podman_volumes }} that hold real service data -- those are
# directories on the host and podman does not know about them. The dangling
# ones seen in practice were 804 MB copies of Nextcloud's /var/www/html left
# by container recreations, which are image content, not data.
set -uo pipefail
TAG=podman-prune
log() { logger -t "$TAG" -p daemon.info -- "$*"; echo "$TAG: $*"; }
total_before=0
total_after=0
for u in {{ podman_prune_users | join(' ') }}; do
# Rootless podman: -H so HOME points at the user's store, and the `cd;`
# preamble is required (see CLAUDE.md) or podman cannot find its graph root.
run() {
sudo -H -u "$u" bash -c \
'cd; d=/run/user/$(id -u); [ -d "$d" ] && export XDG_RUNTIME_DIR="$d"
exec podman "$@"' _ "$@"
}
if ! id "$u" >/dev/null 2>&1; then
log "user=$u status=skipped reason=no-such-user"
continue
fi
before=$(run system df --format '{{ '{{' }}.Size{{ '}}' }}' 2>/dev/null | head -1)
# Deliberately NOT `set -e`: a prune failing for one user must not stop the
# other, and a busy image is a normal, non-fatal outcome.
img=$(run image prune -af --filter "until={{ podman_prune_until }}" 2>&1 | tail -1)
vol=$(run volume prune -f 2>&1 | tail -1)
after=$(run system df --format '{{ '{{' }}.Size{{ '}}' }}' 2>/dev/null | head -1)
log "user=$u images_before=$before images_after=$after"
log "user=$u image_prune=${img:-none} volume_prune=${vol:-none}"
done
log "status=ok"