fix(podman-prune): reap the CI stores, which had reached 113 GB

podman_prune_users listed only the podman and git users, so the two stores that
turn over fastest were never touched. gitea-runner had reached 1205 images /
113.1 GB with 100% of it reclaimable, and actions-runner had 137 exited job
containers. That layer count is what makes overlayfs lookups -- and so CI
itself -- slow; the disk was the lesser problem.

Split into two policies, because the stores are not the same kind of thing:

  service users keep the 30-day rollback window, and their containers are
  deliberately NOT pruned. They are the live services, and reaping one that
  merely happens to be stopped would turn a transient crash into a unit that
  cannot start again until the next deploy.

  CI users get 48h and their exited job containers reaped too. Build layers
  carry no rollback value. Containers are reaped BEFORE images on purpose: an
  exited container pins the image it ran from, so pruning images first would
  leave those layers behind for another day.

Timer moved weekly -> daily; a week of CI turnover is what let the store reach
113 GB between runs. Persistent=true is kept so a missed run catches up.

First run reclaimed 134 GB: gitea-runner 113.1 -> 4.2 GB, actions-runner
7.6 GB -> 0, podman 25.7 -> 15.1 GB. Disk 449G -> 315G, all 26 containers up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Bastian de Byl
2026-08-28 11:39:45 -04:00
co-authored by Claude Opus 5
parent f21be79452
commit 2a6390c0ad
3 changed files with 70 additions and 25 deletions
+13 -1
View File
@@ -269,7 +269,19 @@ podman_prune_users:
- "{{ git_user }}"
# Keep 30 days of unused images so a rollback needs no rebuild or re-pull.
podman_prune_until: 720h
podman_prune_oncalendar: "Sun *-*-* 02:00:00"
# The CI runners were never pruned and had run away: gitea-runner reached 1205
# images / 113 GB with 100% reclaimable, actions-runner 137 exited job
# containers. Build layers carry no rollback value, so they keep a far shorter
# window than the service stores and their exited containers are reaped too.
podman_prune_ci_users:
- gitea-runner
- actions-runner
podman_prune_ci_until: 48h
# Daily rather than weekly: CI turns over many images a day, and a week of that
# is what let the store reach 113 GB between runs.
podman_prune_oncalendar: "*-*-* 02:00:00"
# Hardened CIFS options for the TrueNAS shares (see containers/home/photos.yml).
# x-systemd.automount is what makes a failed mount recoverable without a human.