Files
deploy_home/ansible/roles/podman/templates/podman-prune.sh.j2
T
Bastian de BylandClaude Opus 5.5 0cc4c1460e feat(debyltech-cloud): limit customers to Files, Activity and signing
Nothing was group-restricted, so customers saw Dashboard, Photos, Office
and the rest, plus Nextcloud's first-run and promo apps.

- Disable for everyone: firstrunwizard, recommendations, related_resources,
  weather_status, survey_client, support, app_api, contactsinteraction,
  photos. A refused disable now fails the play (occ exits 0 on "can't be
  disabled").
- Restrict dashboard and office to staff. defaultapp=dashboard,files so
  staff land on the dashboard and customers fall through to Files.
- libresign groups_request_sign pinned to staff and asserted in verify.
  LibreSign itself is deliberately NOT group-restricted: that also blocks
  anonymous requests and would break public signing links.
- profile.enabled=false; lookup_server="" (lookup_server_connector
  can't be disabled).
- README: what customers can open.

Checked as a probe customer: apps=files,activity,libresign,text,viewer,
lands in Files, can't request signatures; staff land on Dashboard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 19:48:39 -04:00

100 lines
4.2 KiB
Django/Jinja

#!/bin/bash
# {{ ansible_managed }}
# Daily reclaim of unused podman images, volumes and exited job containers.
#
# Every image bump leaves the previous tag behind and nothing ever removed
# them: this was written after finding 896 images totalling 59.6 GB, 75% of it
# unused -- 94 tags of greg-time-bot and 73 of fulfillr, one per deploy.
#
# Two policies, because the stores serve different purposes:
#
# service users ({{ podman_prune_users | join(', ') }})
# until={{ podman_prune_until }} keeps recent images so a rollback does not
# require a rebuild or re-pull. Containers are deliberately NOT pruned here:
# they are the live services, and reaping one that merely happens to be
# stopped would turn a transient crash into a unit that cannot start again
# until the next deploy.
#
# CI users ({{ podman_prune_ci_users | join(', ') }})
# Build layers are throwaway and there is no rollback to protect, so these
# get a much shorter window ({{ podman_prune_ci_until }}) and their exited
# job containers are reaped too. They were never covered before: gitea-
# runner had reached 1205 images / 113 GB, 100% of it reclaimable, and it
# is the layer count that makes overlayfs lookups -- and so CI itself -- slow.
#
# Volumes pruned here are podman's ANONYMOUS volumes, not the bind mounts
# under {{ podman_volumes }} that hold real service data -- those are
# directories on the host and podman does not know about them. The dangling
# ones seen in practice were 804 MB copies of Nextcloud's /var/www/html left
# by container recreations, which are image content, not data.
set -uo pipefail
TAG=podman-prune
log() { logger -t "$TAG" -p daemon.info -- "$*"; echo "$TAG: $*"; }
# Rootless podman: -H so HOME points at the user's store, and the `cd;`
# preamble is required (see CLAUDE.md) or podman cannot find its graph root.
run() {
local u=$1
shift
sudo -H -u "$u" bash -c \
'cd; d=/run/user/$(id -u); [ -d "$d" ] && export XDG_RUNTIME_DIR="$d"
exec podman "$@"' _ "$@"
}
# prune_user <user> <until> <prune_containers: yes|no> [image prune filter...]
prune_user() {
local u=$1 keep=$2 do_containers=$3
shift 3
local before after img vol con
if ! id "$u" >/dev/null 2>&1; then
log "user=$u status=skipped reason=no-such-user"
return
fi
before=$(run "$u" system df --format '{{ '{{' }}.Size{{ '}}' }}' 2>/dev/null | head -1)
# Deliberately NOT `set -e`: a prune failing for one user must not stop the
# others, and a busy image is a normal, non-fatal outcome.
#
# Containers are reaped BEFORE images on purpose -- an exited container pins
# the image it ran from, so pruning images first would leave those layers
# behind for another day.
con=none
if [ "$do_containers" = yes ]; then
con=$(run "$u" container prune -f --filter "until=$keep" 2>&1 | tail -1)
fi
img=$(run "$u" image prune -af --filter "until=$keep" "$@" 2>&1 | tail -1)
vol=$(run "$u" volume prune -f 2>&1 | tail -1)
after=$(run "$u" system df --format '{{ '{{' }}.Size{{ '}}' }}' 2>/dev/null | head -1)
log "user=$u keep=$keep size_before=$before size_after=$after"
log "user=$u image_prune=${img:-none} volume_prune=${vol:-none} container_prune=${con:-none}"
}
for u in {{ podman_prune_users | join(' ') }}; do
prune_user "$u" "{{ podman_prune_until }}" no
done
for u in {{ podman_prune_ci_users | join(' ') }}; do
# CI base images (gitea-ci, -espidf, -platformio) carry the keep label: they
# are rebuilt or re-pulled from the registry only when missing, so pruning
# them just forces a multi-GB re-download on the next job.
prune_user "$u" "{{ podman_prune_ci_until }}" yes --filter "label!={{ podman_prune_ci_keep_label }}"
# ...but the label is inherited by every build of those images, including the
# one a rebuild supersedes. That copy loses its tag and becomes dangling, and
# the label filter above would keep it forever -- a multi-GB leak per weekly
# rebuild, ESP-IDF alone being several GB. Without -a, `image prune` removes
# only dangling images, so it can ignore the label without touching the live
# tagged ones.
dangling=$(run "$u" image prune -f --filter "until={{ podman_prune_ci_until }}" 2>&1 | tail -1)
log "user=$u dangling_prune=${dangling:-none}"
done
log "status=ok"