Three findings from investigating sustained I/O pressure on the root SSD.
Grouped because the logging and storage edits land in the same task file.
rsyslog was writing a second complete copy of the journal to
/var/log/messages: 6.7 GB of rotated copies, two weekly files of which were
2.8 GB and 2.5 GB. It loads imjournal, so it reads the journal directly and
ForwardToSyslog=no alone does not stop it -- the unit itself has to go.
Verified nothing consumes those files first: fail2ban runs backend=systemd and
matches on the journal ("No file is currently monitored"), and lsof showed only
rsyslogd holding them. Measured afterwards: writes 46 -> 23 GB/day, /var/log
6.7 GB -> 655 MB, journald still capturing container stdout.
The SSD was on bfq, which fedora's stock 60-block-scheduler.rules picks for any
rotational=0 disk. bfq is built for spinning disks and desktop interactivity:
it costs CPU per request, lets reads queue behind write bursts, and hard-caps
nr_requests at 64. With 26 containers and two CI runners writing at once that
is the wrong trade. mq-deadline rather than none because this is SATA with a
32-deep NCQ queue, not NVMe -- the merging and the read-expiry deadline both
earn their place.
Writeback was at the stock percent-of-RAM ratios, so on 31 GB the kernel would
sit on 3.1 GB before starting writeback and 6.2 GB before blocking writers.
Flushing that to a QLC drive that falls to ~80-160 MB/s once its SLC cache is
spent takes tens of seconds with everything stalled behind it. Capped in
absolute bytes instead: one long stall traded for frequent short ones.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Caddy was rotating on implicit defaults (100MiB/keep 10/90d) that were not
holding -- 20 rotated files per stream and a 190-day-old .gz, 1.3 GB across
16 log streams. Made explicit at 10MiB/keep 3/7d.
Note roll_size et al are subdirectives of `output file`, NOT of `log`.
Getting that wrong does not degrade gracefully: Caddy refuses to start on a
bad config, so every site went down until it was corrected. Worth a
`caddy validate` gate before reload.
journald had no SystemMaxUse and had reached 4 GB, drifting toward its
10%-of-filesystem default (~190 GB on this root). Capped at 500M.
Both are safe to keep short because fluent-bit ships the journal and every
Caddy access log into Graylog -- though note its GELF output has been
erroring for days, which weakens that premise and wants investigating.
The larger find was unrelated to logs: 896 images totalling 59.6 GB with
75% unused (94 tags of greg-time-bot, 73 of fulfillr -- one per deploy) and
5.4 GB of dangling volumes, mostly 804 MB Nextcloud /var/www/html trees
orphaned by container recreations. Pruned to 22 images / 15.4 GB, and added
a weekly timer keeping 30 days so a rollback still needs no rebuild.
Also dropped the decommissioned 6379/tcp redis rule (nothing listening;
Immich's redis is on the shared podman network) and the orphaned nosql, s3
and searxng volume dirs.
Backup log exclusions turned out to be unnecessary: Gitea logs to console
so its log dirs are empty, Nextcloud already excludes its own, BookStack
mounts only uploads, and Caddy is not backed up.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>