retire Graylog behind a flag, fix Caddy reloads, reap awsddns zombies
Graylog was the worst cost/benefit tenant on this 4-core box: two JVMs plus
MongoDB holding ~1.6 GB resident and ~3% CPU around the clock to store ~3k
messages a day -- about 28 MB across its four live indices. journald already
retains ~25 days of the same logs at its 500M cap, so this costs searchability,
not the logs.
The switch is `graylog_enabled` in inventory rather than a role default,
because three roles read it (common, podman, graylog-config). The disabled
path is an active teardown, not a skipped create: the containers already on
the host keep running and their systemd user units keep restarting them at
boot unless something stops and removes them. fluent-bit follows the same
flag -- with the GELF sink down it would spin retrying a dead 127.0.0.1:12202
and fill the journal it exists to drain -- but only the service state follows,
so re-enabling is a restart rather than a reinstall.
Caddy reloads were silently no-ops. The handler read /etc/caddy/Caddyfile,
which is a single-file bind mount, and podman binds those by inode; the
template module writes a temp file and renames it into place, so every deploy
gave the host file a new inode while the container kept seeing the one it was
created with. Config changes only ever landed when something recreated the
container. {{ caddy_path }}/config is also mounted, as a *directory*, and
directory mounts resolve names at open() time -- so /config/Caddyfile is
always the file Ansible just wrote.
awsddns and its four siblings had accumulated 12 zombies over 30 days of
uptime. The image's PID 1 is busybox crond, which only waitpid()s the job PIDs
it tracks and does no generic orphan reaping, so whenever the run-parts/sh
layer exited before the script it left a permanent <defunct>. init: true puts
catatonit at PID 1 to reap them, and the recreation clears the existing ones.
Also bumps fulfillr and greg-time-bot images.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
773a2bbc9c
commit
eea55def6c
@@ -19,10 +19,15 @@ Stages, in order — the ordering is deliberate, see the comments in the templat
|
||||
2. **SQLite snapshots** (`.backup`, then `pragma integrity_check`) where used.
|
||||
3. **rsync** of the data tree, config, and db dumps to TrueNAS.
|
||||
|
||||
Failures raise `status=failed` on the `nextcloud-backup` syslog tag, which an
|
||||
external Graylog rule matches to send mail. Do not rename that tag: it is shared
|
||||
by every instance including Gitea, and renaming it here silently stops alerting
|
||||
for all of them.
|
||||
Failures raise `status=failed` on the `nextcloud-backup` syslog tag. The mail
|
||||
itself is sent by the unit's own `OnFailure=` handler
|
||||
(`templates/nextcloud/nextcloud-backup-alert.sh.j2`) straight through
|
||||
`sendmail`, so alerting does **not** depend on Graylog and is unaffected by
|
||||
`graylog_enabled` being off. A Graylog rule matching the same tag is a
|
||||
secondary, dashboard-side copy of that signal.
|
||||
|
||||
Do not rename that tag: it is shared by every instance including Gitea, and
|
||||
renaming it here silently stops the Graylog-side alerting for all of them.
|
||||
|
||||
|
||||
## Restore
|
||||
@@ -141,3 +146,55 @@ LibreSign setup failures and buys nothing at this scale.
|
||||
- `LC_ALL` / `LANG` must be set or the JVM comes up as `ANSI_X3.4-1968` and
|
||||
LibreSign warns that accented characters in signer names will be mangled
|
||||
(LibreSign issue #4872).
|
||||
|
||||
|
||||
## Logging
|
||||
|
||||
Every container runs with `log_driver=journald`, so container stdout lands in
|
||||
the host journal, capped at 500M by `roles/common/templates/journald-size.conf.j2`
|
||||
(about 25 days at the current rate). `journalctl CONTAINER_NAME=<name>` is the
|
||||
day-to-day way to read it.
|
||||
|
||||
On top of that sits an optional Graylog stack — `graylog`, `graylog-opensearch`,
|
||||
`graylog-mongo`, fed by a host `fluent-bit` service that tails the journal and
|
||||
ships GELF to `127.0.0.1:12202`, enriched with the MaxMind GeoIP database, and
|
||||
configured over the REST API by the separate `graylog-config` role.
|
||||
|
||||
**It is off.** The switch is `graylog_enabled` in
|
||||
`inventories/home/hosts.yml`, and it gates all five of those pieces at once:
|
||||
|
||||
| `graylog_enabled` | effect |
|
||||
| --- | --- |
|
||||
| `true` | stack deployed, fluent-bit shipping, GeoIP downloaded, `graylog-config` runs, `logs.debyl.io` proxies to the UI and `/gelf` |
|
||||
| `false` | containers removed and their systemd user units disabled and deleted, fluent-bit stopped and disabled, GeoIP and `graylog-config` skipped, `logs.debyl.io` answers 503 |
|
||||
|
||||
It was turned off because the cost/benefit is bad on a 4-core box: two JVMs plus
|
||||
MongoDB held ~1.6 GB resident and ~3% of the CPU continuously to store roughly
|
||||
3k messages a day — about 28 MB of real log data across four live indices — all
|
||||
of which journald already keeps for longer.
|
||||
|
||||
Turning it back on is `graylog_enabled: true` plus `make deploy TAGS=graylog`.
|
||||
Nothing is destroyed by the off path: the volumes under `{{ graylog_path }}`
|
||||
keep the indices, the Mongo database holding streams/pipelines/dashboards, and
|
||||
the node-id file, so the stack comes back with its configuration intact.
|
||||
|
||||
Two things genuinely stop while it is off, both by design:
|
||||
|
||||
- **Search and dashboards.** The logs still exist in journald; the query
|
||||
interface over them does not.
|
||||
- **External GELF ingest.** The AWS Lambda that POSTs to `logs.debyl.io/gelf`
|
||||
for the `debyltech-api` stream gets a 503. Those events are dropped, not
|
||||
queued — nothing else records them.
|
||||
|
||||
### Caddy config reloads
|
||||
|
||||
The `reload caddy` handler deliberately reads `/config/Caddyfile`, not the
|
||||
`/etc/caddy/Caddyfile` the container starts from, even though both are the same
|
||||
host file. `/etc/caddy/Caddyfile` is a **single-file** bind mount, which podman
|
||||
binds by inode, and Ansible's `template` module writes a temp file and renames
|
||||
it into place — so every deploy gives the host file a new inode while the
|
||||
container keeps seeing the one it was created with. Reloading from that path
|
||||
silently re-applied the previous config; changes only landed when something
|
||||
recreated the container. `{{ caddy_path }}/config` is also bind-mounted as a
|
||||
*directory* at `/config`, and directory mounts resolve names at `open()` time,
|
||||
so `/config/Caddyfile` is always the file Ansible just wrote.
|
||||
|
||||
Reference in New Issue
Block a user