Files
deploy_home/ansible/roles/podman/tasks/containers/debyltech/graylog-teardown.yml
T
Bastian de BylandClaude Opus 5 eea55def6c retire Graylog behind a flag, fix Caddy reloads, reap awsddns zombies
Graylog was the worst cost/benefit tenant on this 4-core box: two JVMs plus
MongoDB holding ~1.6 GB resident and ~3% CPU around the clock to store ~3k
messages a day -- about 28 MB across its four live indices. journald already
retains ~25 days of the same logs at its 500M cap, so this costs searchability,
not the logs.

The switch is `graylog_enabled` in inventory rather than a role default,
because three roles read it (common, podman, graylog-config). The disabled
path is an active teardown, not a skipped create: the containers already on
the host keep running and their systemd user units keep restarting them at
boot unless something stops and removes them. fluent-bit follows the same
flag -- with the GELF sink down it would spin retrying a dead 127.0.0.1:12202
and fill the journal it exists to drain -- but only the service state follows,
so re-enabling is a restart rather than a reinstall.

Caddy reloads were silently no-ops. The handler read /etc/caddy/Caddyfile,
which is a single-file bind mount, and podman binds those by inode; the
template module writes a temp file and renames it into place, so every deploy
gave the host file a new inode while the container kept seeing the one it was
created with. Config changes only ever landed when something recreated the
container. {{ caddy_path }}/config is also mounted, as a *directory*, and
directory mounts resolve names at open() time -- so /config/Caddyfile is
always the file Ansible just wrote.

awsddns and its four siblings had accumulated 12 zombies over 30 days of
uptime. The image's PID 1 is busybox crond, which only waitpid()s the job PIDs
it tracks and does no generic orphan reaping, so whenever the run-parts/sh
layer exited before the script it left a permanent <defunct>. init: true puts
catatonit at PID 1 to reap them, and the recreation clears the existing ones.

Also bumps fulfillr and greg-time-bot images.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:40:10 -04:00

60 lines
1.9 KiB
YAML

---
# Runs in place of graylog.yml when graylog_enabled is false.
#
# A bare `when:` on the import would only stop Ansible from *creating* the
# stack; containers already on the host would keep running and their systemd
# user units would keep starting them at boot. So the disabled path has to be an
# active teardown: disable the units, then remove the containers.
#
# Volumes under {{ graylog_path }} are deliberately untouched -- indices, the
# Mongo database holding streams/pipelines/dashboards, and the node-id file all
# survive, so flipping graylog_enabled back to true resumes where this left off.
# graylog is stopped first: its unit `requires` the other two, and tearing down
# a dependency out from under it makes the shutdown noisy for no reason.
- name: stop and disable graylog stack systemd units
become: true
become_user: "{{ podman_user }}"
ansible.builtin.systemd:
name: "{{ item }}.service"
enabled: false
state: stopped
daemon_reload: true
scope: user
loop:
- graylog
- graylog-opensearch
- graylog-mongo
# Left over from the Elasticsearch-to-OpenSearch migration; the container is
# long gone but the enabled unit still tries to start it every boot.
- graylog-elastic
register: graylog_units
failed_when: false
tags: graylog
- name: remove graylog stack containers
become: true
become_user: "{{ podman_user }}"
containers.podman.podman_container:
name: "{{ item }}"
state: absent
loop:
- graylog
- graylog-opensearch
- graylog-mongo
tags: graylog
- name: remove graylog stack systemd unit files
become: true
become_user: "{{ podman_user }}"
ansible.builtin.file:
path: "{{ podman_home }}/.config/systemd/user/{{ item }}.service"
state: absent
loop:
- graylog
- graylog-opensearch
- graylog-mongo
- graylog-elastic
notify: reload podman systemd
tags: graylog