Files
deploy_home/ansible/roles/ups/handlers/main.yml
T
Bastian de BylandClaude Opus 5 9954d774e7 fix(ups): start upsd only once the network is actually online
After the host rebooted, truenas began mailing NOCOMM continuously. The UPS
driver was fine throughout -- nut-driver@cyberpower stayed running and the
CyberPower is still on USB here. What died was upsd, which publishes its state
to truenas over the network:

  upsd: not listening on 192.168.1.10 port 3493
  upsd: Fatal error: some listening interfaces were not available
  nut-server.service: Start request repeated too quickly

upsd binds an explicit address, but the packaged unit only orders itself
After=network.target, which is satisfied when networking STARTS rather than
when an address exists. It tried to bind 11 seconds into boot, before
NetworkManager had assigned the address, then burned all five default restart
attempts inside one second, tripped the start limit and stayed dead. The
shipped unit carries these two lines commented out, because upstream knows the
case.

network-online.target is the correct ordering; NetworkManager-wait-online is
enabled here so it genuinely waits for addresses. StartLimitIntervalSec=0 and
RestartSec are belt and braces: a slow address can no longer exhaust the
attempts, and retries are spaced instead of hammered.

truenas needs no change -- it is a correctly configured SLAVE that reconnected
on its own, and its NUT config is regenerated from the middleware database
anyway. Verified: upsd listening on both addresses, UPS OL at 100%, and truenas
querying it again with zero failures since.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:40:19 -04:00

41 lines
1.2 KiB
YAML

---
# Regenerates the nut-driver@<name> unit instances from ups.conf. Oneshot,
# so "restarted" just means "run it again".
- name: reload nut driver units
become: true
ansible.builtin.systemd:
name: nut-driver-enumerator.service
state: restarted
daemon_reload: true
# The enumerator only writes unit definitions; it will not pick up changed
# driver options in an already-running driver. Restart the instance itself.
- name: restart nut driver
become: true
ansible.builtin.systemd:
name: "nut-driver@{{ ups_name }}.service"
state: restarted
# daemon_reload so the drop-in under nut-server.service.d is picked up. The
# reload also clears the start-limit counter, which matters because a unit that
# has latched into "start request repeated too quickly" refuses a plain restart
# until that state is reset.
- name: restart nut server
become: true
ansible.builtin.systemd:
name: nut-server.service
state: restarted
daemon_reload: true
- name: restart nut monitor
become: true
ansible.builtin.systemd:
name: nut-monitor.service
state: restarted
- name: restart firewalld
become: true
ansible.builtin.systemd:
name: firewalld
state: restarted