5776dbe1bf
ups role: NUT server on home.debyl.io for the CyberPower PR1500RT2U that backs both it and truenas.localdomain, with staged shutdown (TrueNAS sheds at t+2min, host at 10% charge) and best-effort IPMI power-on when mains returns. The 10% threshold leans on ignorelb + override.battery.charge.low rather than a custom poller, because CyberPower asserts its own low-battery flag far too early. Credentials come from vault vars; nothing sensitive is templated in the clear. Nextcloud background jobs: both instances have backgroundjobs_mode "cron", which expects an external caller every ~5 minutes, and nothing was calling. The personal instance had not run a background job since 2026-05-14 and skudak since 2024-11-20. Consequently trash and file versions never expired, stale chunked uploads accumulated, calendar reminders never fired, and nextcloud.log was never rotated -- which quietly made the existing log_rotate_size cap inert. Added a systemd timer per instance, skipping cleanly when the container is down or in maintenance so deploy windows don't show up as failed units. Trash retention on the personal instance: the default "auto" only expires when disk space demands it, so 66 GB of >30-day deletions sat on a host with 1.3 TB free -- effectively unbounded. "auto, 30" makes the 30-day expiry unconditional while still purging early under pressure. Image bumps: nextcloud 33.0.0 -> 34.0.2 (both cloud and skudak-cloud) greg-time-bot 3.9.25 -> 3.10.0 fulfillr 20260723.2044 -> 20260728.2155 (prod and dev) The fulfillr bump records what is already deployed: both containers were rolled to that image on 2026-07-29 for SCRUM-156 (digital product releases + customer update campaign). Committing it keeps the repo from claiming an older tag than the host is actually running, which would otherwise roll fulfillr backwards on the next clean-checkout deploy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
64 lines
2.4 KiB
Django/Jinja
64 lines
2.4 KiB
Django/Jinja
#!/bin/bash
|
|
# {{ ansible_managed }}
|
|
# Power truenas.localdomain back on after an outage, but only when it is
|
|
# genuinely safe to do so. Called by upssched once mains has been stable,
|
|
# and once at boot by ups-restore.service (deep-drain recovery).
|
|
#
|
|
# Deliberately stateless: we ask the iDRAC whether the chassis is off
|
|
# rather than tracking whether we were the ones who shut it down. A NAS
|
|
# powered off by hand while on mains is safe, because no ONLINE event
|
|
# fires in that case and the boot path checks UPS status first.
|
|
#
|
|
# NOTE: as of now this is a best-effort secondary path. The R415 has no
|
|
# iDRAC6 Enterprise card ("iDRAC6 Ent Pres ... Absent"), so its shared-LOM
|
|
# BMC has no standby power and goes unreachable whenever the chassis is
|
|
# off - exactly when we would want it. An unreachable iDRAC is therefore
|
|
# the EXPECTED case here, and we exit quietly rather than alarming.
|
|
#
|
|
# The primary restore path needs no IPMI: the R415 power restore policy is
|
|
# set to always-on, so when the UPS cuts and then restores its output the
|
|
# server powers itself back up. Fit an iDRAC6 Enterprise card and switch it
|
|
# to dedicated mode and this script starts working with no changes.
|
|
set -uo pipefail
|
|
|
|
TAG=ups-restore
|
|
UPS={{ ups_name }}@localhost
|
|
|
|
log() { logger -t "$TAG" -- "$*"; echo "$TAG: $*"; }
|
|
|
|
# At boot, upsd may not be serving yet. Wait a bounded amount of time.
|
|
for _ in $(seq 1 30); do
|
|
if upsc "$UPS" ups.status >/dev/null 2>&1; then
|
|
break
|
|
fi
|
|
sleep 2
|
|
done
|
|
|
|
status=$(upsc "$UPS" ups.status 2>/dev/null || echo UNKNOWN)
|
|
if [[ "$status" != *OL* ]]; then
|
|
log "UPS status is '$status', not on line - refusing to power TrueNAS on"
|
|
exit 0
|
|
fi
|
|
|
|
power=$(/usr/local/bin/truenas-power.sh status 2>&1 || true)
|
|
case "$power" in
|
|
*"is off"*)
|
|
log "mains stable and chassis off - powering TrueNAS on"
|
|
if /usr/local/bin/truenas-power.sh on; then
|
|
log "power-on command accepted"
|
|
else
|
|
log "power-on command FAILED - check iDRAC at {{ idrac_host }}"
|
|
exit 1
|
|
fi
|
|
;;
|
|
*"is on"*)
|
|
log "TrueNAS already on, nothing to do"
|
|
;;
|
|
*)
|
|
# Expected while the box has no iDRAC6 Enterprise card: the BMC is
|
|
# simply not on the network with the chassis powered down.
|
|
log "iDRAC at {{ idrac_host }} unreachable - relying on the" \
|
|
"always-on power restore policy instead ($power)"
|
|
;;
|
|
esac
|