UPS monitoring, Nextcloud cron, container image bumps

ups role: NUT server on home.debyl.io for the CyberPower PR1500RT2U that backs
both it and truenas.localdomain, with staged shutdown (TrueNAS sheds at t+2min,
host at 10% charge) and best-effort IPMI power-on when mains returns. The 10%
threshold leans on ignorelb + override.battery.charge.low rather than a custom
poller, because CyberPower asserts its own low-battery flag far too early.
Credentials come from vault vars; nothing sensitive is templated in the clear.

Nextcloud background jobs: both instances have backgroundjobs_mode "cron", which
expects an external caller every ~5 minutes, and nothing was calling. The
personal instance had not run a background job since 2026-05-14 and skudak since
2024-11-20. Consequently trash and file versions never expired, stale chunked
uploads accumulated, calendar reminders never fired, and nextcloud.log was never
rotated -- which quietly made the existing log_rotate_size cap inert. Added a
systemd timer per instance, skipping cleanly when the container is down or in
maintenance so deploy windows don't show up as failed units.

Trash retention on the personal instance: the default "auto" only expires when
disk space demands it, so 66 GB of >30-day deletions sat on a host with 1.3 TB
free -- effectively unbounded. "auto, 30" makes the 30-day expiry unconditional
while still purging early under pressure.

Image bumps:
  nextcloud       33.0.0  -> 34.0.2   (both cloud and skudak-cloud)
  greg-time-bot   3.9.25  -> 3.10.0
  fulfillr        20260723.2044 -> 20260728.2155  (prod and dev)

The fulfillr bump records what is already deployed: both containers were rolled
to that image on 2026-07-29 for SCRUM-156 (digital product releases + customer
update campaign). Committing it keeps the repo from claiming an older tag than
the host is actually running, which would otherwise roll fulfillr backwards on
the next clean-checkout deploy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Bastian de Byl
2026-07-30 13:00:51 -04:00
parent 16145fb6bd
commit 5776dbe1bf
27 changed files with 768 additions and 5 deletions
+114
View File
@@ -0,0 +1,114 @@
# ups
UPS monitoring and staged shutdown for the home rack.
A CyberPower PR1500RT2U (`0764:0601`) is cabled by USB to `home.debyl.io` and
backs both that host and `truenas.localdomain` (Dell PowerEdge R415). This role
makes `home.debyl.io` the NUT server and gives it the ability to power TrueNAS
back on over IPMI.
## Outage sequence
| When | What happens | Driven by |
| --- | --- | --- |
| t+0 | UPS goes on battery, `ONBATT` logged to journald → Graylog | `upsmon` |
| t+2min | TrueNAS shuts itself down cleanly, shedding ~200 W | TrueNAS UPS service, Slave mode, `Shutdown Timer 120` |
| 10% charge | `home.debyl.io` shuts itself down and tells the UPS to cut its output | `upsmon` `SHUTDOWNCMD` + `/lib/systemd/system-shutdown/nutshutdown` |
| mains returns | UPS re-energizes; both machines power themselves back up | R415 `always-on` restore policy; `home.debyl.io` BIOS *After Power Loss → Power On* |
| mains back +5min | Best-effort IPMI power-on, if a dedicated iDRAC is ever fitted | `upssched``ups-restore.sh` |
| host boot | Same restore check, for the deep-drain case | `ups-restore.service` |
The 10% threshold is not a custom poller. CyberPower asserts its own low-battery
flag around 2035%, so `ups.conf` sets `ignorelb` plus
`override.battery.charge.low`, and stock `upsmon` fires at exactly the
configured percentage.
Likewise, the 2-minute TrueNAS shed is TrueNAS's own native "shutdown timer"
setting — no SSH key and no shutdown script from this side.
## Powering TrueNAS back on
Neither Wake-on-LAN nor IPMI works on this box today, so restore is done with
the chassis power restore policy instead.
**Wake-on-LAN is out.** `bce0` advertises no WOL capability (`ifconfig -m bce0`
has no `WOL_MAGIC`) and both NICs are bonded into an LACP `lagg0`.
**IPMI is out too, for now.** The iDRAC is in `shared / LOM1` mode and the
Enterprise card that would provide a dedicated management port is not fitted:
```
$ ipmitool sdr elist | grep -i idrac
iDRAC6 Ent Pres | 70h | ok | 7.1 | Absent
```
A shared-LOM iDRAC6 Express has no standby power. Measured directly: with the
chassis powered off the BMC does not even answer ARP, and it only reappears
~200s into POST, at the moment the host brings the NIC link up. That is a
hardware limitation, not a switch or BIOS problem.
**So restore works like this instead.** `ipmitool chassis policy always-on` is
set on the R415. On a deep outage `home.debyl.io` halts at 10% and NUT's
shutdown hook tells the UPS to cut its output; when mains returns the UPS
re-energizes, the R415 sees AC and boots itself.
The gap is the medium outage — mains returns after TrueNAS has shed but before
the battery reaches 10%. The UPS never cuts power, so TrueNAS stays off and
needs a manual power button press. Fitting a used iDRAC6 Enterprise card and
running `ipmitool delloem lan set dedicated` (plus a cable to the dedicated
port) closes that gap, and `ups-restore.sh` starts working with no code
changes — it is already deployed and simply logs and exits while the BMC is
unreachable.
## One-time setup outside Ansible
These are not managed by this role.
**iDRAC (already done, via `ipmitool` on TrueNAS):**
```
ipmitool lan set 1 ipsrc static
ipmitool lan set 1 ipaddr 192.168.1.12
ipmitool lan set 1 netmask 255.255.255.0
ipmitool lan set 1 defgw ipaddr 192.168.1.1
ipmitool lan set 1 access on # was disabled - nothing answers without this
ipmitool channel setaccess 1 2 callin=on ipmi=on link=on privilege=4
ipmitool user set password 2 '<idrac_password>'
ipmitool chassis policy always-on # this is what restores power after an outage
```
`idrac_password` is also the iDRAC web UI password for `root` - they share a
user database.
**TrueNAS UI → Services → UPS** (enable + start automatically):
| Field | Value |
| --- | --- |
| UPS Mode | Slave |
| Remote Host | `192.168.1.10` |
| Remote Port | `3493` |
| Identifier | `cyberpower` |
| Monitor User | `truenas` |
| Monitor Password | vault `nut_truenas_password` |
| Shutdown Mode | UPS goes on battery |
| Shutdown Timer | `120` |
| Shutdown Command | `/sbin/shutdown -p now` |
| Power Off UPS | unchecked |
**BIOS on `home.debyl.io`** (Lenovo 10MR0004US): Power → *After Power Loss*
**Power On**. On a full drain the NUT shutdown hook
(`/lib/systemd/system-shutdown/nutshutdown`) tells the UPS to cut its own
output; this BIOS setting is what brings the host back when mains returns.
## Vault keys
`nut_upsmon_password`, `nut_truenas_password`, `idrac_password`
## Operating
```
upsc cyberpower # full UPS status
upsc cyberpower battery.charge
sudo /usr/local/bin/truenas-power.sh status|on|soft|off
journalctl -t ups-restore -t ups-sched
```