Three findings from investigating sustained I/O pressure on the root SSD.
Grouped because the logging and storage edits land in the same task file.
rsyslog was writing a second complete copy of the journal to
/var/log/messages: 6.7 GB of rotated copies, two weekly files of which were
2.8 GB and 2.5 GB. It loads imjournal, so it reads the journal directly and
ForwardToSyslog=no alone does not stop it -- the unit itself has to go.
Verified nothing consumes those files first: fail2ban runs backend=systemd and
matches on the journal ("No file is currently monitored"), and lsof showed only
rsyslogd holding them. Measured afterwards: writes 46 -> 23 GB/day, /var/log
6.7 GB -> 655 MB, journald still capturing container stdout.
The SSD was on bfq, which fedora's stock 60-block-scheduler.rules picks for any
rotational=0 disk. bfq is built for spinning disks and desktop interactivity:
it costs CPU per request, lets reads queue behind write bursts, and hard-caps
nr_requests at 64. With 26 containers and two CI runners writing at once that
is the wrong trade. mq-deadline rather than none because this is SATA with a
32-deep NCQ queue, not NVMe -- the merging and the read-expiry deadline both
earn their place.
Writeback was at the stock percent-of-RAM ratios, so on 31 GB the kernel would
sit on 3.1 GB before starting writeback and 6.2 GB before blocking writers.
Flushing that to a QLC drive that falls to ~80-160 MB/s once its SLC cache is
spent takes tens of seconds with everything stalled behind it. Capped in
absolute bytes instead: one long stall traded for frequent short ones.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
19 lines
1.0 KiB
Django/Jinja
19 lines
1.0 KiB
Django/Jinja
# {{ ansible_managed }}
|
|
# Fedora's /usr/lib/udev/rules.d/60-block-scheduler.rules picks bfq for every
|
|
# rotational=0 SATA disk. bfq is a fairness scheduler built for spinning disks
|
|
# and desktop interactivity: it costs real CPU per request and, with 26
|
|
# containers plus two CI runners all writing at once, it lets reads queue
|
|
# behind write bursts. That is the shape of the stalls seen here -- I/O
|
|
# pressure spiking to 76% while the CPU sat nearly idle.
|
|
#
|
|
# mq-deadline instead of none: this is a SATA SSD with a 32-deep NCBQ queue,
|
|
# not an NVMe device with its own deep queues, so the request merging and the
|
|
# read-expiry deadline are both worth having. The deadline is what stops reads
|
|
# starving behind a QLC write burst.
|
|
#
|
|
# bfq also hard-caps nr_requests at 64; mq-deadline allows a deeper queue,
|
|
# which is what lets concurrent container and CI I/O actually overlap.
|
|
ACTION=="add|change", KERNEL=="sd[a-z]", ATTR{queue/rotational}=="0", \
|
|
ATTR{queue/scheduler}="{{ ssd_io_scheduler }}", \
|
|
ATTR{queue/nr_requests}="{{ ssd_nr_requests }}"
|