Commit Graph
5 Commits
Author SHA1 Message Date
Bastian de BylandClaude Opus 5 eea55def6c retire Graylog behind a flag, fix Caddy reloads, reap awsddns zombies
Graylog was the worst cost/benefit tenant on this 4-core box: two JVMs plus
MongoDB holding ~1.6 GB resident and ~3% CPU around the clock to store ~3k
messages a day -- about 28 MB across its four live indices. journald already
retains ~25 days of the same logs at its 500M cap, so this costs searchability,
not the logs.

The switch is `graylog_enabled` in inventory rather than a role default,
because three roles read it (common, podman, graylog-config). The disabled
path is an active teardown, not a skipped create: the containers already on
the host keep running and their systemd user units keep restarting them at
boot unless something stops and removes them. fluent-bit follows the same
flag -- with the GELF sink down it would spin retrying a dead 127.0.0.1:12202
and fill the journal it exists to drain -- but only the service state follows,
so re-enabling is a restart rather than a reinstall.

Caddy reloads were silently no-ops. The handler read /etc/caddy/Caddyfile,
which is a single-file bind mount, and podman binds those by inode; the
template module writes a temp file and renames it into place, so every deploy
gave the host file a new inode while the container kept seeing the one it was
created with. Config changes only ever landed when something recreated the
container. {{ caddy_path }}/config is also mounted, as a *directory*, and
directory mounts resolve names at open() time -- so /config/Caddyfile is
always the file Ansible just wrote.

awsddns and its four siblings had accumulated 12 zombies over 30 days of
uptime. The image's PID 1 is busybox crond, which only waitpid()s the job PIDs
it tracks and does no generic orphan reaping, so whenever the run-parts/sh
layer exited before the script it left a permanent <defunct>. init: true puts
catatonit at PID 1 to reap them, and the recreation clears the existing ones.

Also bumps fulfillr and greg-time-bot images.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:40:10 -04:00
Bastian de BylandClaude Opus 5 f11391b28f bound log growth and reclaim ~52 GB of container disk
Caddy was rotating on implicit defaults (100MiB/keep 10/90d) that were not
holding -- 20 rotated files per stream and a 190-day-old .gz, 1.3 GB across
16 log streams. Made explicit at 10MiB/keep 3/7d.

Note roll_size et al are subdirectives of `output file`, NOT of `log`.
Getting that wrong does not degrade gracefully: Caddy refuses to start on a
bad config, so every site went down until it was corrected. Worth a
`caddy validate` gate before reload.

journald had no SystemMaxUse and had reached 4 GB, drifting toward its
10%-of-filesystem default (~190 GB on this root). Capped at 500M.

Both are safe to keep short because fluent-bit ships the journal and every
Caddy access log into Graylog -- though note its GELF output has been
erroring for days, which weakens that premise and wants investigating.

The larger find was unrelated to logs: 896 images totalling 59.6 GB with
75% unused (94 tags of greg-time-bot, 73 of fulfillr -- one per deploy) and
5.4 GB of dangling volumes, mostly 804 MB Nextcloud /var/www/html trees
orphaned by container recreations. Pruned to 22 images / 15.4 GB, and added
a weekly timer keeping 30 days so a rollback still needs no rebuild.

Also dropped the decommissioned 6379/tcp redis rule (nothing listening;
Immich's redis is on the shared podman network) and the orphaned nosql, s3
and searxng volume dirs.

Backup log exclusions turned out to be unnecessary: Gitea logs to console
so its log dirs are empty, Nextcloud already excludes its own, BookStack
mounts only uploads, and Caddy is not backed up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:13:35 -04:00
Bastian de BylandClaude Opus 5 0ab423ca55 harden nextcloud backups: db dumps, alerting, drift fix
The data-only rsync left no way to restore a working instance: mysql/ and
config/ were never backed up, so a recovery would have files but no shares,
users or metadata. Dump the database before syncing files (a DB older than
the files is repairable with occ files:scan; a newer one references blobs
that never made it into the backup) and ship config/ alongside it.

Capture the --chmod=Du=rwx,Dgo=rx flag that had been hand-added to the
deployed skudak-cloud script. It was outside git, so every deploy silently
reverted it. It now lives in backup_rsync_extra_args.

Add OnFailure= alerting. The units failed silently before, which is how an
iDrive sync failure sat unnoticed since May. msmtp rather than the esmtp
already installed: the OpenSRS relay is port 465 (implicit TLS) and libesmtp
only speaks STARTTLS.

Exclude nextcloud.log* from the sync and cap log_rotate_size. skudak-cloud
was running at loglevel 0 and had written a 64 GB log that was being rsynced
and pushed to S3; set it to 2 to match the home instance.

Stagger the timers (04:00 / 04:30) so both finish before the 05:00 TrueNAS
snapshot task, and bound TimeoutStartSec so a wedged rsync cannot leave the
unit activating forever and skip every subsequent trigger.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 16:03:18 -04:00
Bastian de BylandClaude Opus 4.5 61692b36a2 refactor: reorganize fluent-bit and geoip out of containers
- Move fluent-bit to common role (systemd service, not a container)
- Move geoip to podman/tasks/data/ (data prep, not a container)
- Remove debyltech tag from geoip (not a debyltech service)
- Fix check_mode for fetch subuid task to enable dry-run mode

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-28 12:34:43 -05:00
Bastian de Byl 38561cb968 gitea, zomboid updates, ssh key fixes 2025-12-19 10:39:56 -05:00