retire Graylog behind a flag, fix Caddy reloads, reap awsddns zombies

Graylog was the worst cost/benefit tenant on this 4-core box: two JVMs plus
MongoDB holding ~1.6 GB resident and ~3% CPU around the clock to store ~3k
messages a day -- about 28 MB across its four live indices. journald already
retains ~25 days of the same logs at its 500M cap, so this costs searchability,
not the logs.

The switch is `graylog_enabled` in inventory rather than a role default,
because three roles read it (common, podman, graylog-config). The disabled
path is an active teardown, not a skipped create: the containers already on
the host keep running and their systemd user units keep restarting them at
boot unless something stops and removes them. fluent-bit follows the same
flag -- with the GELF sink down it would spin retrying a dead 127.0.0.1:12202
and fill the journal it exists to drain -- but only the service state follows,
so re-enabling is a restart rather than a reinstall.

Caddy reloads were silently no-ops. The handler read /etc/caddy/Caddyfile,
which is a single-file bind mount, and podman binds those by inode; the
template module writes a temp file and renames it into place, so every deploy
gave the host file a new inode while the container kept seeing the one it was
created with. Config changes only ever landed when something recreated the
container. {{ caddy_path }}/config is also mounted, as a *directory*, and
directory mounts resolve names at open() time -- so /config/Caddyfile is
always the file Ansible just wrote.

awsddns and its four siblings had accumulated 12 zombies over 30 days of
uptime. The image's PID 1 is busybox crond, which only waitpid()s the job PIDs
it tracks and does no generic orphan reaping, so whenever the run-parts/sh
layer exited before the script it left a permanent <defunct>. init: true puts
catatonit at PID 1 to reap them, and the recreation clears the existing ones.

Also bumps fulfillr and greg-time-bot images.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Bastian de Byl
2026-08-25 23:40:10 -04:00
co-authored by Claude Opus 5
parent 773a2bbc9c
commit eea55def6c
11 changed files with 250 additions and 14 deletions
+4
View File
@@ -11,11 +11,15 @@
name: fail2ban
state: restarted
# Guarded: a config change still notifies this handler when graylog_enabled is
# false, and an unconditional restart would start the service the fluent-bit
# task just stopped.
- name: restart fluent-bit
become: true
ansible.builtin.systemd:
name: fluent-bit
state: restarted
when: graylog_enabled | bool
- name: restart_journald
become: true
+9 -3
View File
@@ -1,6 +1,12 @@
---
# Fluent Bit - Log forwarder from journald to Graylog GELF
# Deployed as systemd service (not container) for direct journal access
#
# Graylog's GELF input is fluent-bit's only output, so the two share a switch:
# with the stack down fluent-bit would just spin retrying a dead 127.0.0.1:12202
# and filling the journal it is meant to be draining. The package and config
# stay installed either way -- only the service follows graylog_enabled (see
# inventories/home/hosts.yml) -- so re-enabling is a restart, not a reinstall.
- name: install fluent-bit package
become: true
@@ -37,9 +43,9 @@
mode: '0644'
notify: restart fluent-bit
- name: enable and start fluent-bit service
- name: set fluent-bit service state to match graylog_enabled
become: true
ansible.builtin.systemd:
name: fluent-bit
enabled: true
state: started
enabled: "{{ graylog_enabled | bool }}"
state: "{{ 'started' if (graylog_enabled | bool) else 'stopped' }}"
@@ -3,10 +3,19 @@
# 1.9 TB root means it will happily grow into the tens of gigabytes; it had
# reached 4 GB before this was set.
#
# Kept small on purpose: every container runs with log_driver=journald and
# Every container runs with log_driver=journald, so this is where container
# stdout lands.
{% if graylog_enabled | bool %}
# fluent-bit drains the journal into Graylog continuously (systemd input,
# _COMM=conmon -- see templates/fluent-bit/fluent-bit.conf.j2), so Graylog is
# the system of record. What stays here is only the buffer that covers
# fluent-bit being down, and 500M is a long outage at this log rate.
# the system of record and what stays here is only the buffer that covers
# fluent-bit being down.
{% else %}
# Graylog is disabled (see graylog_enabled in inventories/home/hosts.yml), so
# the journal is now the only log store. No increase was needed for that: the
# cap is a size limit, not a time limit, and fluent-bit only ever read the
# journal rather than rotating it, so retention is unchanged at roughly 25 days
# at the current rate -- longer than Graylog's own four live indices covered.
{% endif %}
[Journal]
SystemMaxUse={{ journald_max_use | default('500M') }}