Three findings from investigating sustained I/O pressure on the root SSD.
Grouped because the logging and storage edits land in the same task file.
rsyslog was writing a second complete copy of the journal to
/var/log/messages: 6.7 GB of rotated copies, two weekly files of which were
2.8 GB and 2.5 GB. It loads imjournal, so it reads the journal directly and
ForwardToSyslog=no alone does not stop it -- the unit itself has to go.
Verified nothing consumes those files first: fail2ban runs backend=systemd and
matches on the journal ("No file is currently monitored"), and lsof showed only
rsyslogd holding them. Measured afterwards: writes 46 -> 23 GB/day, /var/log
6.7 GB -> 655 MB, journald still capturing container stdout.
The SSD was on bfq, which fedora's stock 60-block-scheduler.rules picks for any
rotational=0 disk. bfq is built for spinning disks and desktop interactivity:
it costs CPU per request, lets reads queue behind write bursts, and hard-caps
nr_requests at 64. With 26 containers and two CI runners writing at once that
is the wrong trade. mq-deadline rather than none because this is SATA with a
32-deep NCQ queue, not NVMe -- the merging and the read-expiry deadline both
earn their place.
Writeback was at the stock percent-of-RAM ratios, so on 31 GB the kernel would
sit on 3.1 GB before starting writeback and 6.2 GB before blocking writers.
Flushing that to a QLC drive that falls to ~80-160 MB/s once its SLC cache is
spent takes tens of seconds with everything stalled behind it. Capped in
absolute bytes instead: one long stall traded for frequent short ones.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
29 lines
1.4 KiB
Django/Jinja
29 lines
1.4 KiB
Django/Jinja
### {{ ansible_managed }}
|
|
# Without a cap journald sizes itself at 10% of the filesystem, which on this
|
|
# 1.9 TB root means it will happily grow into the tens of gigabytes; it had
|
|
# reached 4 GB before this was set.
|
|
#
|
|
# Container stdout lands here: every container runs with log_driver=journald
|
|
# except zomboid, which is chatty enough to evict everything else and so gets
|
|
# its own rotating k8s-file log instead.
|
|
{% if graylog_enabled | bool %}
|
|
# fluent-bit drains the journal into Graylog continuously (systemd input,
|
|
# _COMM=conmon -- see templates/fluent-bit/fluent-bit.conf.j2), so Graylog is
|
|
# the system of record and what stays here is only the buffer that covers
|
|
# fluent-bit being down.
|
|
{% else %}
|
|
# Graylog is disabled (see graylog_enabled in inventories/home/hosts.yml), so
|
|
# the journal is now the only log store. No increase was needed for that: the
|
|
# cap is a size limit, not a time limit, and fluent-bit only ever read the
|
|
# journal rather than rotating it, so retention is unchanged at roughly 25 days
|
|
# at the current rate -- longer than Graylog's own four live indices covered.
|
|
{% endif %}
|
|
[Journal]
|
|
SystemMaxUse={{ journald_max_use | default('500M') }}
|
|
|
|
# Nothing should be forwarded to syslog: rsyslog is disabled (see
|
|
# tasks/service.yml) because it duplicated the whole journal into
|
|
# /var/log/messages. This also closes the imuxsock path so anything that
|
|
# logs via logger(1) still lands in the journal and nowhere else.
|
|
ForwardToSyslog=no
|