Commit Graph
2 Commits
Author SHA1 Message Date
Bastian de BylandClaude Opus 5 9954d774e7 fix(ups): start upsd only once the network is actually online
After the host rebooted, truenas began mailing NOCOMM continuously. The UPS
driver was fine throughout -- nut-driver@cyberpower stayed running and the
CyberPower is still on USB here. What died was upsd, which publishes its state
to truenas over the network:

  upsd: not listening on 192.168.1.10 port 3493
  upsd: Fatal error: some listening interfaces were not available
  nut-server.service: Start request repeated too quickly

upsd binds an explicit address, but the packaged unit only orders itself
After=network.target, which is satisfied when networking STARTS rather than
when an address exists. It tried to bind 11 seconds into boot, before
NetworkManager had assigned the address, then burned all five default restart
attempts inside one second, tripped the start limit and stayed dead. The
shipped unit carries these two lines commented out, because upstream knows the
case.

network-online.target is the correct ordering; NetworkManager-wait-online is
enabled here so it genuinely waits for addresses. StartLimitIntervalSec=0 and
RestartSec are belt and braces: a slow address can no longer exhaust the
attempts, and retries are spaced instead of hammered.

truenas needs no change -- it is a correctly configured SLAVE that reconnected
on its own, and its NUT config is regenerated from the middleware database
anyway. Verified: upsd listening on both addresses, UPS OL at 100%, and truenas
querying it again with zero failures since.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:40:19 -04:00
Bastian de BylandClaude Opus 5 5776dbe1bf UPS monitoring, Nextcloud cron, container image bumps
ups role: NUT server on home.debyl.io for the CyberPower PR1500RT2U that backs
both it and truenas.localdomain, with staged shutdown (TrueNAS sheds at t+2min,
host at 10% charge) and best-effort IPMI power-on when mains returns. The 10%
threshold leans on ignorelb + override.battery.charge.low rather than a custom
poller, because CyberPower asserts its own low-battery flag far too early.
Credentials come from vault vars; nothing sensitive is templated in the clear.

Nextcloud background jobs: both instances have backgroundjobs_mode "cron", which
expects an external caller every ~5 minutes, and nothing was calling. The
personal instance had not run a background job since 2026-05-14 and skudak since
2024-11-20. Consequently trash and file versions never expired, stale chunked
uploads accumulated, calendar reminders never fired, and nextcloud.log was never
rotated -- which quietly made the existing log_rotate_size cap inert. Added a
systemd timer per instance, skipping cleanly when the container is down or in
maintenance so deploy windows don't show up as failed units.

Trash retention on the personal instance: the default "auto" only expires when
disk space demands it, so 66 GB of >30-day deletions sat on a host with 1.3 TB
free -- effectively unbounded. "auto, 30" makes the 30-day expiry unconditional
while still purging early under pressure.

Image bumps:
  nextcloud       33.0.0  -> 34.0.2   (both cloud and skudak-cloud)
  greg-time-bot   3.9.25  -> 3.10.0
  fulfillr        20260723.2044 -> 20260728.2155  (prod and dev)

The fulfillr bump records what is already deployed: both containers were rolled
to that image on 2026-07-29 for SCRUM-156 (digital product releases + customer
update campaign). Committing it keeps the repo from claiming an older tag than
the host is actually running, which would otherwise roll fulfillr backwards on
the next clean-checkout deploy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 13:00:51 -04:00