After the host rebooted, truenas began mailing NOCOMM continuously. The UPS driver was fine throughout -- nut-driver@cyberpower stayed running and the CyberPower is still on USB here. What died was upsd, which publishes its state to truenas over the network: upsd: not listening on 192.168.1.10 port 3493 upsd: Fatal error: some listening interfaces were not available nut-server.service: Start request repeated too quickly upsd binds an explicit address, but the packaged unit only orders itself After=network.target, which is satisfied when networking STARTS rather than when an address exists. It tried to bind 11 seconds into boot, before NetworkManager had assigned the address, then burned all five default restart attempts inside one second, tripped the start limit and stayed dead. The shipped unit carries these two lines commented out, because upstream knows the case. network-online.target is the correct ordering; NetworkManager-wait-online is enabled here so it genuinely waits for addresses. StartLimitIntervalSec=0 and RestartSec are belt and braces: a slow address can no longer exhaust the attempts, and retries are spaced instead of hammered. truenas needs no change -- it is a correctly configured SLAVE that reconnected on its own, and its NUT config is regenerated from the middleware database anyway. Verified: upsd listening on both addresses, UPS OL at 100%, and truenas querying it again with zero failures since. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
31 lines
1.5 KiB
Django/Jinja
31 lines
1.5 KiB
Django/Jinja
# {{ ansible_managed }}
|
|
# upsd binds explicit addresses (see upsd.conf), but the packaged unit only
|
|
# orders itself After=network.target -- which is satisfied when networking
|
|
# STARTS, not when an address actually exists. The shipped unit even carries
|
|
# these two lines commented out, because upstream knows the case.
|
|
#
|
|
# On 2026-08-28 the host came up and upsd tried to bind 11 seconds into boot,
|
|
# before NetworkManager had assigned the address:
|
|
# upsd: not listening on {{ ups_listen_addr }} port {{ ups_listen_port }}
|
|
# upsd: Fatal error: some listening interfaces were not available
|
|
# It then burned all five of the default restart attempts inside one second,
|
|
# tripped the start limit, and stayed dead. truenas monitors this host as a
|
|
# SLAVE (MONITOR cyberpower@{{ ups_listen_addr }}:{{ ups_listen_port }}), so
|
|
# it alarmed NOCOMM continuously until someone noticed. The UPS driver itself
|
|
# was fine throughout -- only the server that publishes its state was gone.
|
|
#
|
|
# NetworkManager-wait-online is enabled on this host, so network-online.target
|
|
# genuinely waits for addresses rather than just for the service to start.
|
|
#
|
|
# The restart settings are belt and braces: StartLimitIntervalSec=0 removes the
|
|
# rate limit so a slow address can never exhaust the attempts, and RestartSec
|
|
# spaces the retries out instead of hammering five times in a second.
|
|
[Unit]
|
|
Wants=network-online.target
|
|
After=network-online.target
|
|
StartLimitIntervalSec=0
|
|
|
|
[Service]
|
|
Restart=on-failure
|
|
RestartSec={{ ups_server_restart_sec }}
|