Files
deploy_home/ansible/roles/common/defaults/main.yml
T
Bastian de BylandClaude Opus 5 f21be79452 perf(host): drop the duplicate syslog copy and fix the SSD I/O path
Three findings from investigating sustained I/O pressure on the root SSD.
Grouped because the logging and storage edits land in the same task file.

rsyslog was writing a second complete copy of the journal to
/var/log/messages: 6.7 GB of rotated copies, two weekly files of which were
2.8 GB and 2.5 GB. It loads imjournal, so it reads the journal directly and
ForwardToSyslog=no alone does not stop it -- the unit itself has to go.
Verified nothing consumes those files first: fail2ban runs backend=systemd and
matches on the journal ("No file is currently monitored"), and lsof showed only
rsyslogd holding them. Measured afterwards: writes 46 -> 23 GB/day, /var/log
6.7 GB -> 655 MB, journald still capturing container stdout.

The SSD was on bfq, which fedora's stock 60-block-scheduler.rules picks for any
rotational=0 disk. bfq is built for spinning disks and desktop interactivity:
it costs CPU per request, lets reads queue behind write bursts, and hard-caps
nr_requests at 64. With 26 containers and two CI runners writing at once that
is the wrong trade. mq-deadline rather than none because this is SATA with a
32-deep NCQ queue, not NVMe -- the merging and the read-expiry deadline both
earn their place.

Writeback was at the stock percent-of-RAM ratios, so on 31 GB the kernel would
sit on 3.1 GB before starting writeback and 6.2 GB before blocking writers.
Flushing that to a QLC drive that falls to ~80-160 MB/s once its SLC cache is
spent takes tens of seconds with everything stalled behind it. Capped in
absolute bytes instead: one long stall traded for frequent short ones.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:39:05 -04:00

33 lines
712 B
YAML

---
deps:
[
cockpit-podman,
cronie,
fail2ban,
fail2ban-selinux,
git,
logrotate,
podman,
podman-docker,
python-docker,
]
fail2ban_jails: [sshd.local, zomboid.local]
fail2ban_filters: [zomboid.conf]
services:
- crond
- podman.socket
- podman
- fail2ban
- systemd-timesyncd
# Storage tuning for the SATA SSD (see tasks/service.yml and the udev template).
ssd_io_scheduler: mq-deadline
ssd_nr_requests: "256"
# 256 MB before background writeback starts, 1 GB before writers block. Absolute
# bytes rather than the default percent-of-RAM ratios, which scale to multi-GB
# stalls on a 31 GB host.
vm_dirty_background_bytes: "268435456"
vm_dirty_bytes: "1073741824"