podman_prune_users listed only the podman and git users, so the two stores that
turn over fastest were never touched. gitea-runner had reached 1205 images /
113.1 GB with 100% of it reclaimable, and actions-runner had 137 exited job
containers. That layer count is what makes overlayfs lookups -- and so CI
itself -- slow; the disk was the lesser problem.
Split into two policies, because the stores are not the same kind of thing:
service users keep the 30-day rollback window, and their containers are
deliberately NOT pruned. They are the live services, and reaping one that
merely happens to be stopped would turn a transient crash into a unit that
cannot start again until the next deploy.
CI users get 48h and their exited job containers reaped too. Build layers
carry no rollback value. Containers are reaped BEFORE images on purpose: an
exited container pins the image it ran from, so pruning images first would
leave those layers behind for another day.
Timer moved weekly -> daily; a week of CI turnover is what let the store reach
113 GB between runs. Persistent=true is kept so a missed run catches up.
First run reclaimed 134 GB: gitea-runner 113.1 -> 4.2 GB, actions-runner
7.6 GB -> 0, podman 25.7 -> 15.1 GB. Disk 449G -> 315G, all 26 containers up.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>