retire Graylog behind a flag, fix Caddy reloads, reap awsddns zombies

Graylog was the worst cost/benefit tenant on this 4-core box: two JVMs plus
MongoDB holding ~1.6 GB resident and ~3% CPU around the clock to store ~3k
messages a day -- about 28 MB across its four live indices. journald already
retains ~25 days of the same logs at its 500M cap, so this costs searchability,
not the logs.

The switch is `graylog_enabled` in inventory rather than a role default,
because three roles read it (common, podman, graylog-config). The disabled
path is an active teardown, not a skipped create: the containers already on
the host keep running and their systemd user units keep restarting them at
boot unless something stops and removes them. fluent-bit follows the same
flag -- with the GELF sink down it would spin retrying a dead 127.0.0.1:12202
and fill the journal it exists to drain -- but only the service state follows,
so re-enabling is a restart rather than a reinstall.

Caddy reloads were silently no-ops. The handler read /etc/caddy/Caddyfile,
which is a single-file bind mount, and podman binds those by inode; the
template module writes a temp file and renames it into place, so every deploy
gave the host file a new inode while the container kept seeing the one it was
created with. Config changes only ever landed when something recreated the
container. {{ caddy_path }}/config is also mounted, as a *directory*, and
directory mounts resolve names at open() time -- so /config/Caddyfile is
always the file Ansible just wrote.

awsddns and its four siblings had accumulated 12 zombies over 30 days of
uptime. The image's PID 1 is busybox crond, which only waitpid()s the job PIDs
it tracks and does no generic orphan reaping, so whenever the run-parts/sh
layer exited before the script it left a permanent <defunct>. init: true puts
catatonit at PID 1 to reap them, and the recreation clears the existing ones.

Also bumps fulfillr and greg-time-bot images.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Bastian de Byl
2026-08-25 23:40:10 -04:00
co-authored by Claude Opus 5
parent 773a2bbc9c
commit eea55def6c
11 changed files with 250 additions and 14 deletions
+16 -3
View File
@@ -81,13 +81,13 @@
- import_tasks: containers/debyltech/fulfillr.yml
vars:
image: git.debyl.io/debyltech/fulfillr:20260728.2155
image: git.debyl.io/debyltech/fulfillr:20260825.1909
tags: debyltech, fulfillr
# Staging back-office (fulfillr-dev.debyltech.com) — same image, staging Turso config.
- import_tasks: containers/debyltech/fulfillr-dev.yml
vars:
image: git.debyl.io/debyltech/fulfillr:20260728.2155
image: git.debyl.io/debyltech/fulfillr:20260825.1909
tags: debyltech, fulfillr-dev
- import_tasks: containers/debyltech/uptime-kuma.yml
@@ -100,7 +100,10 @@
image: docker.io/louislam/uptime-kuma:2.3.2
tags: home, uptime
# GeoIP is only consumed by Graylog's enrichment pipelines, so it follows the
# same switch -- see graylog_enabled in inventories/home/hosts.yml.
- import_tasks: data/geoip.yml
when: graylog_enabled | bool
tags: graylog, geoip
- import_tasks: containers/debyltech/graylog.yml
@@ -108,16 +111,26 @@
mongo_image: docker.io/mongo:7.0
opensearch_image: docker.io/opensearchproject/opensearch:2
image: docker.io/graylog/graylog:7.0.1
when: graylog_enabled | bool
tags: debyltech, graylog
# The disabled path is an active teardown, not just a skipped create: without it
# the containers already on the host keep running and their systemd user units
# keep restarting them at boot.
- import_tasks: containers/debyltech/graylog-teardown.yml
when: not (graylog_enabled | bool)
tags: debyltech, graylog
- import_tasks: containers/home/gregtime.yml
vars:
image: localhost/greg-time-bot:3.10.0
image: localhost/greg-time-bot:3.14.1
tags: gregtime
# Gated off by default — see zomboid_enabled in roles/podman/defaults/main.yml.
- import_tasks: containers/home/zomboid.yml
vars:
image: docker.io/cm2network/steamcmd:root
when: zomboid_enabled | bool
tags: zomboid
# ---------------------------------------------------------- Gitea backups