retire Graylog behind a flag, fix Caddy reloads, reap awsddns zombies
Graylog was the worst cost/benefit tenant on this 4-core box: two JVMs plus
MongoDB holding ~1.6 GB resident and ~3% CPU around the clock to store ~3k
messages a day -- about 28 MB across its four live indices. journald already
retains ~25 days of the same logs at its 500M cap, so this costs searchability,
not the logs.
The switch is `graylog_enabled` in inventory rather than a role default,
because three roles read it (common, podman, graylog-config). The disabled
path is an active teardown, not a skipped create: the containers already on
the host keep running and their systemd user units keep restarting them at
boot unless something stops and removes them. fluent-bit follows the same
flag -- with the GELF sink down it would spin retrying a dead 127.0.0.1:12202
and fill the journal it exists to drain -- but only the service state follows,
so re-enabling is a restart rather than a reinstall.
Caddy reloads were silently no-ops. The handler read /etc/caddy/Caddyfile,
which is a single-file bind mount, and podman binds those by inode; the
template module writes a temp file and renames it into place, so every deploy
gave the host file a new inode while the container kept seeing the one it was
created with. Config changes only ever landed when something recreated the
container. {{ caddy_path }}/config is also mounted, as a *directory*, and
directory mounts resolve names at open() time -- so /config/Caddyfile is
always the file Ansible just wrote.
awsddns and its four siblings had accumulated 12 zombies over 30 days of
uptime. The image's PID 1 is busybox crond, which only waitpid()s the job PIDs
it tracks and does no generic orphan reaping, so whenever the run-parts/sh
layer exited before the script it left a permanent <defunct>. init: true puts
catatonit at PID 1 to reap them, and the recreation clears the existing ones.
Also bumps fulfillr and greg-time-bot images.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
773a2bbc9c
commit
eea55def6c
@@ -1,5 +1,26 @@
|
||||
---
|
||||
all:
|
||||
vars:
|
||||
# Master switch for the Graylog logging stack (graylog + graylog-opensearch
|
||||
# + graylog-mongo), the fluent-bit journal shipper that feeds it, the GeoIP
|
||||
# database download it enriches with, and the graylog-config role that
|
||||
# provisions its streams/pipelines over the REST API.
|
||||
#
|
||||
# Lives in inventory rather than a role default because it is read by three
|
||||
# separate roles (common, podman, graylog-config) and role defaults are
|
||||
# scoped to their own role.
|
||||
#
|
||||
# Off because the stack is the worst cost/benefit tenant on this 4-core box:
|
||||
# two JVMs plus MongoDB hold ~1.6 GB resident and burn ~3% of the CPU around
|
||||
# the clock to store ~3k messages a day -- about 28 MB of actual log data
|
||||
# across its four live indices. journald already retains ~25 days of the
|
||||
# same logs at its 500M cap, so turning this off costs searchability, not
|
||||
# the logs themselves.
|
||||
#
|
||||
# Set true (here, or -e graylog_enabled=true) to bring it back. All data
|
||||
# under {{ graylog_path }} is left in place, so re-enabling resumes with the
|
||||
# existing indices, streams and pipelines intact.
|
||||
graylog_enabled: false
|
||||
hosts:
|
||||
home.debyl.io:
|
||||
ansible_user: fedora
|
||||
|
||||
Reference in New Issue
Block a user