Commit Graph
100 Commits
Author SHA1 Message Date
Bastian de Byl f7f4d903a0 Merge remote-tracking branch 'origin/master' into feat/ci-images-registry 2026-09-14 19:57:33 -04:00
Bastian de BylandClaude Opus 5 75e257077e SCRUM-196: fulfillr 20260914.2149 (embedded tzdata for GA4)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 17:55:15 -04:00
Bastian de BylandClaude Opus 5 80edc84586 SCRUM-196: fulfillr 20260914.2103 (funnel + GA4 traffic endpoints)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 17:10:08 -04:00
Bastian de BylandClaude Opus 5 42e6f5271d fix(gitea-actions): serve CI images from the Gitea registry
The CI job images only existed under localhost/ in the gitea-runner store,
and the nightly CI prune deletes any image older than 48h that no container
holds. After every idle stretch CI failed in 0-1s pulling
localhost/gitea-ci:latest until the role was re-run and the images rebuilt.

- Build under git.debyl.io/gitbot/..., push after every run, and pull from the
  registry instead of rebuilding when the Containerfile is unchanged.
- Log gitea-runner in via ~/.docker/config.json, which both act_runner (job
  image pulls) and podman read.
- Label the base images io.debyl.ci-base and skip that label in the CI prune;
  its `until` counts from build time, so a re-pulled image would otherwise be
  deleted again the next night.

Workflows pinning `container: image: localhost/gitea-ci-*` must move to the
registry paths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 16:40:57 -04:00
Bastian de BylandClaude Opus 5 e174e259eb SCRUM-196: Read GA4 key from fulfillr_ga4_credentials_json
Match the vault variable name holding the service-account key file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 16:21:04 -04:00
Bastian de BylandClaude Opus 5 e681a46b78 SCRUM-196: GA4 analytics config for fulfillr Traffic & funnel tab
Render an analytics block (property 353859448 + service-account key) into the
fulfillr dev and prod configs once fulfillr_ga4_credentials is in the vault.
Without the vault var the block is omitted and the portal reports GA as not
connected. Remember to restart the container after deploy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 14:02:38 -04:00
Bastian de BylandClaude Opus 5 d9ab55f05d chore(rsvp): bump to 1.0.4
Compacts the admin invite table so it fits without clipping: short token link
with Copy/Msg, status as an emoji, Qty, allergens/notes behind popovers, and
relative update times.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 10:16:39 -04:00
Bastian de BylandClaude Opus 5 1c5d3b020a chore(rsvp): bump to 1.0.3
Fixes the admin invite table clipping its Edit/Revoke column, and adds
default-headcount people estimates to the invited/awaiting summary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:36:08 -04:00
Bastian de BylandClaude Opus 5 1b9fccd779 chore(rsvp): bump to 1.0.2
Adds the anonymized "Who's coming" card for guests who have answered, and a
message-template copy button on the admin invite table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:20:57 -04:00
Bastian de BylandClaude Opus 5 f59ade748b feat(rsvp): deploy rsvp.debyl.io, the invite-only party RSVP app
A single Go binary with SQLite, built and loaded as localhost/rsvpd:<VERSION>
by make deploy-remote in ~/src/rsvp-debylio. Both of its listeners are
published on 127.0.0.1 only: Caddy proxies the public one to everyone and the
admin one (/admin, no login) only to caddy_local_networks. A Caddyfile mistake
alone cannot expose admin, and neither can a port mistake alone.

Guests' invite links are the credential and they sit in the URL path, which
shapes the vhost:

- It does not import common_headers. That snippet sets Referrer-Policy
  same-origin, which would replace the app's no-referrer and let a token leak
  in a Referer header.
- Its access log rewrites request>uri to /i/REDACTED and drops the Location
  response header, since every POST 303s back to /i/<token>.
- Caddy's error logger is separate from the site's and wrote the raw URI to
  caddy.log when the upstream was down. The global log now excludes
  http.log.error.rsvp and a filtered rsvp-errors logger takes it instead.
  Verified with zero token occurrences in both logs, locally and live.

The data directory is owned directly by the host uid of the container's uid
10001 (subuid + 10000). Setting it to the podman user and chowning back each run
flipped ownership on every deploy and briefly locked the app out of its
database.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:46 -04:00
Bastian de BylandClaude Opus 5 a9f51b77e1 fix(podman): pull a new image before removing the running container
podman-check deleted the old container as soon as the pinned image differed,
and the create task pulled afterwards. A tag that did not exist, or a registry
that was down, therefore left the service with no container at all. The pull
now happens first, so that failure stops the play with the old container still
running. localhost/ images are built and loaded by hand and are never pulled.

The pull is skipped when the container does not exist yet: there is nothing to
protect, and containers[0] is not there to compare against. Without that guard
the first deploy of any new service failed on the conditional -- rsvp was the
first to hit it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:46 -04:00
Bastian de BylandClaude Opus 5 73e50bb19d chore(gregtime): bump to 3.17.3
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:45 -04:00
Bastian de BylandClaude Opus 5 6cd4d56de1 feat(labelprint): 4x6 label print proxy on a Raspberry Pi
A Pi 3B+ (stickah.local) shares a Phomemo PM246 to the LAN as a plain CUPS
queue, so any machine can print 4x6 labels -- fulfillr-site's shipping labels
in particular -- without installing the vendor driver, which is x86-64 only.
The role builds the TSPL CUPS driver from source instead.

It is Debian, not Fedora, so it lives in its own inventory and playbook
(make deploy-labelprint / check-labelprint) and the home.debyl.io roles can
never run against it. make bootfs renders its cloud-init first-boot files onto
a freshly imaged SD card from the same templates the role uses. The Wi-Fi
credentials for the home and rescue networks are in the vault.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:45 -04:00
Bastian de BylandClaude Opus 5 be0d02b938 feat(zomboid): restore the world from one of PZ's own backups
Rolling a world back meant hand-work over ssh: stop the service, move the live
save aside, unzip the right archive, chown into the container's subuid range,
relabel, start. That is the wrong shape of task to do by hand, and it is always
done under time pressure -- by construction, because the archive you want is
being deleted while you work.

PZ keeps BackupsCount=10 per set and writes one every BackupsPeriod=30 minutes,
so a periodic backup is reachable for about five hours and then gone. On
2026-09-05 the snapshot the admins asked for (05:16, four minutes before the
incident) had about 90 minutes of life left when the request came in.

Same shape as the wipe: the Discord bot writes a trigger file into its own rw
volume, zomboid-restore.path notices it, and zomboid-restore.service runs the
script as the podman user. The bot gets no ssh, no systemd, and keeps only its
existing read-only mount of the Zomboid volume.

Two details carry most of the correctness.

Resolution is by mtime, not by index. The rotation renames the files -- today's
backup_7.zip is backup_8.zip half an hour from now, and a new backup_7.zip holds
a different world -- so an index is valid only while the listing is fresh, which
is not long enough to survive a human reading a confirmation prompt. The trigger
names a set and an mtime; the script resolves the path itself, whitelists the
filename, and refuses if nothing matches. It never accepts a path.

Everything that can fail is checked before the server is touched. A rotated-out
target, an archive with no debbzoid world in it, a bad action, a traversal
attempt in the set name: each aborts with the server still running and writes a
result file the bot reports back. The live world is moved aside rather than
deleted, so a restore is undoable and the last three are kept.

One thing PZ does not advertise: its backups do not cover the whole save
directory. blam/, a mod's own state, is in none of them -- not the 05:16 archive
and not the newest one. Restoring only what the archive holds therefore lands
the world slightly *behind* the target rather than on it, so anything present in
the displaced world and absent from the archive is carried across.

The gregtime tag moves to 3.17.0 for the bot half of this -- `backups`,
`restore <n>`, `restore confirm`, `restore undo`, gated to the same two admins
as the wipe. That image is built and running on the host; its source is not
committed yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
2026-09-05 09:12:09 -04:00
Bastian de BylandClaude Opus 5 ba9c4f2bfe fix(zomboid): re-sync the world settings from the live server
The vendored Sophie preset and the running debbzoid world had drifted, and the
repo only held the preset. That is not a restore point: force-pushing it would
have reverted the admins' in-game tuning rather than recovering it, which is
exactly how a Sophie world quietly became an Apocalypse one once already -- 139
values reverted, loot from 0.35 back to 0.9, CharacterFreePoints 0 to 60.

files/zomboid/sophie/SandboxVars.lua is now a snapshot of the live world taken
2026-08-31, not the preset as shipped. The modlist is untouched and still
upstream, which is why zomboid_preset_version now names the two halves and their
separate dates. server.ini.j2 carries the eight keys that had drifted:

  PlayerSafehouse              false -> true
  SafehouseAllowNonResidential false -> true   (the diner/gas-station case)
  SafehouseAllowRespawn        false -> true
  SafehouseAllowLoot           true  -> false
  SafehouseAllowFire           true  -> false
  TrashDeleteAll               false -> true
  MapRemotePlayerVisibility    1     -> 4
  ResetID                      6953472 -> 826046

ResetID is in that list on purpose, and matters most. It is the world's
soft-reset token: a file value that differs from the one the live world was
created with tells every connected client to roll a new character. Carrying the
live value makes a deliberate force-push a no-op instead of a server-wide wipe
prompt.

Spawn config gets its own switch, zomboid_spawn_force. spawnregions.lua and
spawnpoints.lua are the only config a running world re-reads -- at every server
start, where SandboxVars is read once, when the world is created -- so a spawn
edit is deployable on the live world without a wipe. Sharing zomboid_config_force
between them would have meant force-pushing the whole preset to land a one-line
spawn edit, rewriting the INI (hence the ResetID hazard above) and the world's
SandboxVars along with it.

config-template/ now tracks the repo unconditionally. Nothing on the host writes
that directory and the server cannot see it, so it has no hand edits to protect;
if it does not track the repo it is not a restore point, just an older world's
settings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
2026-09-05 09:11:31 -04:00
Bastian de BylandClaude Opus 5 9954d774e7 fix(ups): start upsd only once the network is actually online
After the host rebooted, truenas began mailing NOCOMM continuously. The UPS
driver was fine throughout -- nut-driver@cyberpower stayed running and the
CyberPower is still on USB here. What died was upsd, which publishes its state
to truenas over the network:

  upsd: not listening on 192.168.1.10 port 3493
  upsd: Fatal error: some listening interfaces were not available
  nut-server.service: Start request repeated too quickly

upsd binds an explicit address, but the packaged unit only orders itself
After=network.target, which is satisfied when networking STARTS rather than
when an address exists. It tried to bind 11 seconds into boot, before
NetworkManager had assigned the address, then burned all five default restart
attempts inside one second, tripped the start limit and stayed dead. The
shipped unit carries these two lines commented out, because upstream knows the
case.

network-online.target is the correct ordering; NetworkManager-wait-online is
enabled here so it genuinely waits for addresses. StartLimitIntervalSec=0 and
RestartSec are belt and braces: a slow address can no longer exhaust the
attempts, and retries are spaced instead of hammered.

truenas needs no change -- it is a correctly configured SLAVE that reconnected
on its own, and its NUT config is regenerated from the middleware database
anyway. Verified: upsd listening on both addresses, UPS OL at 100%, and truenas
querying it again with zero failures since.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:40:19 -04:00
Bastian de BylandClaude Opus 5 fb09a01c88 fix(podman): stop restorecon walking the TrueNAS CIFS mounts
restorecon -Frv over the podman volumes tree descends into two SMB shares
mounted inside it -- volumes/photos/immich and volumes/photos/storage, both
from truenas -- so it was relabelling the entire remote photo library over the
network, on a filesystem that cannot store SELinux xattrs at all.

On 2026-08-28 that pinned CPU#0 at 100% system time and the kernel logged six
escalating soft lockups:

  watchdog: BUG: soft lockup - CPU#0 stuck for 1423s! [restorecon:132780]

New process creation starved, so sshd accepted connections and then hung during
session setup, and the host needed a hard reboot. An earlier run the same day
had already been SIGKILLed at 49s, which was the same bug surfacing quietly.

-x keeps it on the local filesystem. Measured: 1,279,231 local files walked and
relabelled in 29.2s, against never finishing before. The mounts are also
x-systemd.automount, so merely walking into them triggers a mount -- there was
never anything there to relabel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:40:07 -04:00
Bastian de BylandClaude Opus 5 ae99ba415e fix(immich): keep the ML models resident, and pin the ML URL explicitly
Smart search, face detection and OCR were all failing with

  Machine learning request to "http://immich-machine-learning:3003" failed

while the container still reported Up. Podman only sees PID 1: gunicorn's
master was alive and holding the listening socket, but its worker had died at a
WORKER TIMEOUT and was never respawned -- hence a connect timeout rather than a
refusal. The dead worker was a zombie whose remaining thread was stuck in
uninterruptible sleep in exit_mmap, so it survived SIGKILL, podman rm -f and
rm -f -t 0, and kept the container name and network alias until the host was
rebooted.

MACHINE_LEARNING_MODEL_TTL=0 addresses the cause rather than the symptom. The
default unloads models after 300s idle, so every search following a gap
reloaded four of them (CLIP, buffalo_l detection + recognition, PP-OCRv5) on a
CPU-only 4-core box and then tore those mappings back down -- and that teardown
is what wedged. Keeping them resident costs ~1-2 GB and removes the path.

IMMICH_MACHINE_LEARNING_URL is now set explicitly instead of relying on
immich's implicit default, so the ML container can be renamed without silently
losing search, and the name is a variable so a replacement can be stood up
beside a broken one without editing tasks.

Verified after deploy: zero ML failures, "in-memory cache with unloading
disabled", ping 200 from immich-server, 26/26 containers healthy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:39:57 -04:00
Bastian de BylandClaude Opus 5 2a6390c0ad fix(podman-prune): reap the CI stores, which had reached 113 GB
podman_prune_users listed only the podman and git users, so the two stores that
turn over fastest were never touched. gitea-runner had reached 1205 images /
113.1 GB with 100% of it reclaimable, and actions-runner had 137 exited job
containers. That layer count is what makes overlayfs lookups -- and so CI
itself -- slow; the disk was the lesser problem.

Split into two policies, because the stores are not the same kind of thing:

  service users keep the 30-day rollback window, and their containers are
  deliberately NOT pruned. They are the live services, and reaping one that
  merely happens to be stopped would turn a transient crash into a unit that
  cannot start again until the next deploy.

  CI users get 48h and their exited job containers reaped too. Build layers
  carry no rollback value. Containers are reaped BEFORE images on purpose: an
  exited container pins the image it ran from, so pruning images first would
  leave those layers behind for another day.

Timer moved weekly -> daily; a week of CI turnover is what let the store reach
113 GB between runs. Persistent=true is kept so a missed run catches up.

First run reclaimed 134 GB: gitea-runner 113.1 -> 4.2 GB, actions-runner
7.6 GB -> 0, podman 25.7 -> 15.1 GB. Disk 449G -> 315G, all 26 containers up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:39:45 -04:00
Bastian de BylandClaude Opus 5 f21be79452 perf(host): drop the duplicate syslog copy and fix the SSD I/O path
Three findings from investigating sustained I/O pressure on the root SSD.
Grouped because the logging and storage edits land in the same task file.

rsyslog was writing a second complete copy of the journal to
/var/log/messages: 6.7 GB of rotated copies, two weekly files of which were
2.8 GB and 2.5 GB. It loads imjournal, so it reads the journal directly and
ForwardToSyslog=no alone does not stop it -- the unit itself has to go.
Verified nothing consumes those files first: fail2ban runs backend=systemd and
matches on the journal ("No file is currently monitored"), and lsof showed only
rsyslogd holding them. Measured afterwards: writes 46 -> 23 GB/day, /var/log
6.7 GB -> 655 MB, journald still capturing container stdout.

The SSD was on bfq, which fedora's stock 60-block-scheduler.rules picks for any
rotational=0 disk. bfq is built for spinning disks and desktop interactivity:
it costs CPU per request, lets reads queue behind write bursts, and hard-caps
nr_requests at 64. With 26 containers and two CI runners writing at once that
is the wrong trade. mq-deadline rather than none because this is SATA with a
32-deep NCQ queue, not NVMe -- the merging and the read-expiry deadline both
earn their place.

Writeback was at the stock percent-of-RAM ratios, so on 31 GB the kernel would
sit on 3.1 GB before starting writeback and 6.2 GB before blocking writers.
Flushing that to a QLC drive that falls to ~80-160 MB/s once its SLC cache is
spent takes tens of seconds with everything stalled behind it. Capped in
absolute bytes instead: one long stall traded for frequent short ones.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:39:05 -04:00
Bastian de BylandClaude Opus 5 9e136ce903 fix(hass): trust pasta's link-local address, and bump to 2026.8.3
assistant.debyl.io returned 400 on every request. With use_x_forwarded_for
set, home assistant hard-fails any request carrying X-Forwarded-For from an
address outside trusted_proxies:

  Received X-Forwarded-For header from an untrusted proxy 169.254.2.1

caddy runs --network host and proxies to localhost:8123, so the source address
is whatever pasta presents inside the container's netns. Rootless podman
switched from slirp4netns to pasta in v5 (this host runs 5.7.1), which moved
that address from 10.0.2.x -- covered by the existing 10.0.0.0/8 entry -- to a
link-local tap0 address that nothing in the list matched. The container has no
10.x address at all any more.

Latent since the pasta migration; it only surfaced when the site was next
opened. Reproduced directly:

  curl -H 'X-Forwarded-For: 1.2.3.4' http://localhost:8123/  ->  400

Bumped to 2026.8.3 in the same pass (identical digest to the stable tag).
The config directory was snapshotted first: the 2026.5.1 -> 2026.8.3 database
migration is not reversible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:38:50 -04:00
Bastian de BylandClaude Opus 5 87a0332a75 chore: deploy fulfillr 20260827.1453 to prod (SCRUM-182)
Ships the SES MessageId capture on ticket email timeline events.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 11:00:15 -04:00
Bastian de BylandClaude Opus 5 98706bf418 chore: deploy greg-time-bot 3.16.5
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:19:38 -04:00
Bastian de BylandClaude Opus 5 60d5ec4ae3 fix: leave world settings alone on a wipe, and snapshot them first
Reverts the restore-from-template behaviour. A wipe now resets the world and the
player database only; Server/<name>_{SandboxVars,spawnregions,spawnpoints}.lua
carry over untouched, so hand tuning survives it. Verified by checksum: all three
files byte-identical either side of a wipe.

That gives up the guarantee the restore bought. PZ writes the running world's
settings back over those files, so if a world is ever created with defaults the
file inherits them and later wipes regenerate from them -- which is how a Sophie
world became an Apocalypse one. Nothing corrects that automatically now, so the
wipe takes a snapshot of the settings before it starts, keeping the last ten
under config-backup/. Greg's edits were lost once because the only record of them
was a file the server had since overwritten; that is the hole this fills.
config-template/ still holds the pristine Sophie preset to copy back from.

The snapshot is best-effort throughout. The first version of it created the
directory with plain mkdir, the parent belongs to the container's subuid rather
than the podman user, and set -e turned that into a failed wipe -- the same shape
as the tee that broke this script before. Ansible owns the directory now and
every step of the snapshot tolerates failure, because a backup that cannot be
written is not a reason to refuse to wipe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:10:37 -04:00
Bastian de BylandClaude Opus 5 1c467b1a76 fix: stop the world wipe taking the world settings with it
A wipe was quietly turning a Sophie world into an Apocalypse one. 139 sandbox
values had reverted -- loot rates from 0.35 back to 0.9, ranged weapons and ammo
to 2.0, CharacterFreePoints from 0 to 60 -- and ten of Sophie's mod-added
options had vanished entirely.

The reset script was not deleting the settings. PZ reads
Server/<name>_SandboxVars.lua only when it creates a world, and then writes the
running world's settings back over that same file. The file is an output, not an
input. So once any world came up with defaults, the server stamped those defaults
into the file, and every wipe afterwards regenerated from them -- inheriting the
corruption rather than causing it, and with no way back out on its own. The same
shape as the admin-password deadlock.

The intended settings now live in config-template/, which the server has no
reason to touch, and a wipe restores Server/<name>_{SandboxVars,spawnregions,
spawnpoints}.lua from there before restarting. That is also the file to hand-edit
when changing the world: edit the template, wipe, and the new world has the edits.

Verified by wiping: 139 differences before, 0 after. The server still rewrites
the live file on boot -- 996 keys become 1061 as it adds newer options -- but
every Sophie value survives that now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:55:06 -04:00
Bastian de BylandClaude Opus 5 05fc3b5c65 fix: drop VehicleDoorsHotkey, which never loaded
A Build 41 mod (versionMin=41.60) carried into Sophie's B42 list. It indexes
the ModOptions framework at VehicleDoorsHotkey_Options.lua:4, that framework is
not in the list, and so it dies there every boot:

  attempted index: ModOptions of non-table: null

Twenty-one exceptions a boot for a hotkey that has never once worked. Excluded
rather than fixed by adding ModOptions, since nothing else in the list wants it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 19:33:48 -04:00
Bastian de BylandClaude Opus 5 a47551724b fix: make the Discord world reset actually delete the world
The reset script died on its very first line. log() piped through tee into
zomboid/logs/world-reset.log, that directory was owned by the container's
subuid, and the script runs as the podman user -- so tee returned EACCES and
set -e killed the run before it stopped the server or touched a save. Every
`@bot` reset since had been a no-op that reported nothing.

Nothing mounts zomboid/logs into a container; it only holds output from
host-side helpers running as the podman user, so it is now owned by that user
rather than by the subuid the container volumes need.

log() no longer treats the file as load-bearing either. stdout is already
captured by the journal, so an unwritable log is worth continuing past rather
than aborting a wipe over.

Verified end to end by writing the trigger exactly as the bot does: server
stopped, saves and player database deleted, service restarted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:33:57 -04:00
Bastian de BylandClaude Opus 5 ea98d1b556 fix: strip the server-only log stack script that breaks client world load
The dedicated server (app 380870) and the client are packaged separately and
have drifted. Base.Log_Stack_01 was removed from the client in 42.16, but the
server package still ships entity_logstack.txt.

The server registers every script it can see into a new world's
WorldDictionary, so the entry lands in the save, and every client then fails at
world load with:

  WorldDictionaryException: [SpriteConfigs] Missing dictionary script on
  client: Base.Log_Stack_01

Nothing in multiplayer can legitimately reference a script that no client has,
so dropping it server-side costs nothing.

The strip runs after the SteamCMD update rather than once at install time on
purpose: validate restores the file on every boot, so it has to be removed on
every boot too.

Note this does not repair a world whose dictionary already records the entry --
that save still needs regenerating.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 14:36:36 -04:00
Bastian de BylandClaude Opus 5 62c4015410 fix: let systemd actually notice when the Zomboid container exits
`@bot restart` saves, sends RCON quit, and relies on the unit's Restart=always
to bring the server back. It never came back. The container sat at exited(0)
while the unit reported active/running with NRestarts=0, so the bot waited for
a startup that was never going to happen and the server stayed down until
someone noticed.

podman generate systemd emits Type=forking with ExecStart=podman start and a
PIDFile pointing at conmon. podman start returns immediately, so the process
systemd was told to supervise was never its child -- it warns about exactly
this in the journal, once per poll, and then does not notice the exit:

  zomboid.service: Supervising process 2431312 which is not our child.
  We'll most likely not notice when it exits.

That PIDFile is also stale by design: it embeds the container ID, so every
deploy that recreates the container leaves it pointing at nothing.

Type=simple with `podman start -a` keeps podman in the foreground as systemd's
own child, so the exit is seen and Restart=always does what it always claimed
to. stdout and stderr are discarded deliberately -- attaching re-emits the
container's output for systemd to capture straight back into the journal, which
is the flood the k8s-file driver exists to prevent. podman logs still has it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:43:53 -04:00
Bastian de BylandClaude Opus 5 32748d1786 fix: stop the query filter banning every player who tries to connect
The anti-scanner rules dropped 53-byte Steam query packets above 1/hour with a
burst of 2, per source IP, on the theory that only scanners send queries and
that real players would already have been marked verified by the priority-2
rule when they sent something larger.

That premise is inverted. A client's first contact with the server IS a 53-byte
query, so nobody can be verified before querying, and nobody can query more
than twice an hour without being dropped. The counters said so plainly: five
packets had ever matched the verified-accept rule, against 30,563 drops.
fail2ban then banned each dropped player for a week -- 32 live bans, 550 total,
firing every ten minutes, every one of them a residential address.

Keeps the shape of the protection and moves the threshold somewhere no real
client reaches: opening the server browser or retrying a connection is a
handful of queries, a flood is thousands. Removes the old rules first so hosts
carrying them converge instead of stacking a second copy.

The bans themselves were also hooked into INPUT for tcp only, so they never
blocked the UDP game traffic they were meant to -- the drop rule was doing all
the damage on its own. Cleared the outstanding 32.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:26:45 -04:00
Bastian de BylandClaude Opus 5 a4bb78585b chore: bump skudak Nextcloud to 34.0.3 and BookStack to 26.5.4
Nextcloud follows the personal instance, which is now verified on 34.0.3:
`occ upgrade` completed, needsDbUpgrade is false, and cloud.debyl.io reports
34.0.3.2 out of maintenance mode.

BookStack skips 26.4 -- 26.3.4 is simply where this instance sat. Note the tag
is 26.5.4, not 26.05.4: solidnerd drops the month's leading zero.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:18:43 -04:00
Bastian de BylandClaude Opus 5 ad8a5392d6 chore: bump personal Immich to v3.1.0 and personal Nextcloud to 34.0.3
Immich moves the server and machine-learning images only; the Postgres and
Redis images stay on their pinned digests, which are tied to the vectorchord
extension version rather than the app release. v3.1.0 is a quality-of-life
release -- its one breaking change is dropping iOS 14 in the mobile client,
which does not touch the server.

Nextcloud 34.0.2 -> 34.0.3 is a patch release; the official image's entrypoint
runs `occ upgrade` itself on first start with a newer version.

Skudak's Nextcloud is deliberately left on 34.0.2 until this one is verified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:08:22 -04:00
Bastian de BylandClaude Opus 5 91d9f6b419 fix: actually bound the Zomboid log archive, which logrotate could not
The logrotate config added with the Sophie overhaul claimed to bound PZ's own
logs. It did not. data/Logs had reached 18 GB across 112 entries and 146k
files, and rotation had never once fired there.

Two reasons, both wrong assumptions on my part. PZ rolls its logs into dated
directories -- logs_2025-12-14/ through logs_2026-08-26/, up to 279 MB each --
so the Logs/*.txt glob only ever matched a handful of loose files at the top.
And it rotates by size: not one file in that tree exceeds 100M, because the
growth is in the number of files, not the size of any of them.

Age-based pruning is the right tool, so this adds a daily zomboid-log-prune
timer keeping 14 days and deleting the emptied directories behind it. It runs
under podman unshare, since the files belong to the container's UID.

logrotate keeps server-console.txt, which is a single ever-growing file and
genuinely is what it is good at. Narrowed its scope to say so rather than
implying coverage it never had.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 02:34:30 -04:00
Bastian de BylandClaude Opus 5 a0421226a2 fix: correct four stale mod IDs in the Sophie preset
On 42.20.3 the server loaded 279 of the preset's 283 mods and reported the
other four as "required mod ... not found". None of it was a download problem:
all 256 workshop items fetched, and every one of these four is in the preset's
own WorkshopItems= list. Only the ID each is referenced by has drifted.

Checked against the id= field of the mod.info files on disk rather than folder
names, which are unrelated to mod IDs -- an early look at folder names was
misleading here.

  FWOBenchPress&Treadmill  ->  FWOBenchPressTreadmill  (stray ampersand)
  Ladders42131             ->  Ladders4220             (per-build variants;
                                                        42131 is the 42.16 one)
  ServingPlatesB42         ->  ServingPlates42
  NewMusic_OrchestraMix    ->  NewMusic_CMM            (renamed upstream to
                                                        "Classical Music Mix")

Kept as a rename map applied at template time rather than edited into the
vendored list, so the list stays a faithful copy of upstream and re-vendoring
a newer preset keeps the corrections.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 02:23:31 -04:00
Bastian de BylandClaude Opus 5 4d86a57d75 fix: stop the Zomboid entrypoint withholding the admin password
The flag was gated on the admin database not existing yet, treating the file's
presence as proof the admin user had been created. It is not: PZ creates the
DB file in ServerWorldDatabase.create() before it creates the admin user, so a
boot interrupted between the two leaves a database with no admin in it.

After that the guard withheld -adminpassword on every subsequent start, the
server fell back to an interactive "Enter new administrator password:" prompt,
read EOF because a detached container has no stdin, and died with
NoSuchElementException. systemd restarted it into exactly the same state, 50
times, which is how this was found.

Passing it unconditionally costs nothing when the admin already exists and
removes the deadlock entirely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 02:10:51 -04:00
Bastian de BylandClaude Opus 5 a90f4b7ed0 fix: force the Zomboid branch switch with an explicit -beta public
Dropping -beta from app_update does not leave a beta branch. SteamCMD wrote
"public" into the manifest's UserConfig and left MountedConfig on "unstable",
so the install stayed on the 42.14.1 build from February while reporting
success. The config had already moved to the bare mod-ID syntax that 42.20+
wants, and 42.14.1 still requires the backslash prefix, so all 283 mods
loaded as "required mod ... not found" and the server came up vanilla.

Naming the branch explicitly is what actually remounts the content.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:57:40 -04:00
Bastian de BylandClaude Opus 5 81451247b6 fix: drop the redundant daily from the Zomboid logrotate config
logrotate warns that size overrides daily, and it is right: with both set the
time directive does nothing. The daily timer is when the config gets checked;
100M is when it rotates. Saying so in a comment rather than in a directive
that has no effect.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:48:28 -04:00
Bastian de BylandClaude Opus 5 d34c6cd36d rebuild the Zomboid server on the Sophie 42 preset as "debbzoid"
The deployment had drifted badly. It was pinned to the Steam `-beta unstable`
branch, which stopped being right the moment B42 became the default at 42.20,
and it wrote every mod ID with the `\` prefix that B42 required only during
that unstable period and now rejects. Three server profiles were carried
around (vanilla, modded, b42revamp) whose mod lists were hand-curated blobs
from a Steam collection that has since moved on.

Moves to the Sophie 42 community preset, vendored from GerDeathstar/sophie-pz
("Sophie 42 Files.zip", 2026-07-31 == the b42 release tag): 283 mod IDs, 256
workshop items, map_distanciado over Muldraugh, plus its SandboxVars,
spawnregions and spawnpoints. Sophie's gameplay settings are kept exactly as
shipped -- PVP with its damage modifiers, the safety system, MaxPlayers=32,
PauseEmpty, no sleep, safehouses off. Only the keys this deployment actually
owns are templated over the top.

The modlist is stored as YAML lists rather than the INI's semicolon blobs, so
the next Sophie update produces a diff you can read. Exclusions live in
zomboid_mods_excluded with a reason each, seeded with IconsInventory -- already
absent upstream, listed so a re-vendor cannot quietly bring it back.

Config is now seeded before first boot instead of patched after it. The old
approach could not write the INI until the server had generated one, so every
setting went through lineinfile guarded on a stat; seeding the whole file from
the preset removes the chicken-and-egg and puts the config in git. force is
off by design: these files are a starting point, not managed state, so the
server can be stopped, hand-edited and regenerated without Ansible clobbering
the edits. Push them again deliberately with -e zomboid_config_force=true.

The B41 Discord keys were dead. B42 replaced DiscordChannel/DiscordChannelID
with DiscordChatChannel, which takes a channel name rather than a snowflake,
so the DiscordChannelID=... this role had been appending was a key the server
ignores and the chat bridge has not been working.

Renaming the server to debbzoid is what starts the new world: PZ derives the
INI, save directory and player DB from the name, so gregboid stays on disk
untouched as a rollback.

Re-enabled, because the reason it was off is now fixed. It was disabled for
saturating the SSD -- ~10 MB/s of log writes into the shared journal, which
starved Gitea CI badly enough to stretch a firmware build to 17 minutes. The
container now logs to its own rotating k8s-file instead of the journal every
other service shares, and logrotate caps the server's own server-console.txt
and Logs/ at 100M keeping 3. copytruncate is mandatory there: PZ holds those
fds for its whole life, so a rename would leave it writing into an unlinked
inode. Memory is unchanged; idle draw is nowhere near MAX_RAM.

Also stops the entrypoint recursively chowning the install tree on every boot.
With 256 workshop mods that is a large inode walk and pure churn once the
first run has set ownership.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:41:00 -04:00
Bastian de BylandClaude Opus 5 eea55def6c retire Graylog behind a flag, fix Caddy reloads, reap awsddns zombies
Graylog was the worst cost/benefit tenant on this 4-core box: two JVMs plus
MongoDB holding ~1.6 GB resident and ~3% CPU around the clock to store ~3k
messages a day -- about 28 MB across its four live indices. journald already
retains ~25 days of the same logs at its 500M cap, so this costs searchability,
not the logs.

The switch is `graylog_enabled` in inventory rather than a role default,
because three roles read it (common, podman, graylog-config). The disabled
path is an active teardown, not a skipped create: the containers already on
the host keep running and their systemd user units keep restarting them at
boot unless something stops and removes them. fluent-bit follows the same
flag -- with the GELF sink down it would spin retrying a dead 127.0.0.1:12202
and fill the journal it exists to drain -- but only the service state follows,
so re-enabling is a restart rather than a reinstall.

Caddy reloads were silently no-ops. The handler read /etc/caddy/Caddyfile,
which is a single-file bind mount, and podman binds those by inode; the
template module writes a temp file and renames it into place, so every deploy
gave the host file a new inode while the container kept seeing the one it was
created with. Config changes only ever landed when something recreated the
container. {{ caddy_path }}/config is also mounted, as a *directory*, and
directory mounts resolve names at open() time -- so /config/Caddyfile is
always the file Ansible just wrote.

awsddns and its four siblings had accumulated 12 zombies over 30 days of
uptime. The image's PID 1 is busybox crond, which only waitpid()s the job PIDs
it tracks and does no generic orphan reaping, so whenever the run-parts/sh
layer exited before the script it left a permanent <defunct>. init: true puts
catatonit at PID 1 to reap them, and the recreation clears the existing ones.

Also bumps fulfillr and greg-time-bot images.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:40:10 -04:00
Bastian de BylandClaude Opus 5 773a2bbc9c harden TrueNAS CIFS mounts so immich self-heals
TrueNAS was power-cycled, the CIFS mounts failed, and systemd never retried
-- mount units are not restarted on failure. SMB came back, nothing
remounted, and immich-server served an empty library for days while its
database still listed 14,181 assets pointing at /mnt/media/originals.

fstab carried no _netdev, no nofail and no automount, so there was no path
back without a human. Now:

  x-systemd.automount  any access re-attempts the mount; failure stops
                       being terminal
  _netdev / nofail     ordered after network-online, dead NAS cannot block boot
  soft                 I/O errors instead of blocking forever, so the
                       container can be restarted rather than wedging in
                       uninterruptible sleep
  idle-timeout         unmount when unused, clearing stale handles
  resilienthandles     SMB3 rides out brief blips

ansible.posix.mount mounts directly and never starts the generated
.automount unit, leaving the on-access trigger inactive -- enable it
explicitly, or the headline fix silently does nothing.

The containers are systemd USER units while the mounts are SYSTEM units, so
RequiresMountsFor= is unavailable. cifs-watchdog bridges the scopes: checks
health, recovers, and restarts ONLY immich-server (the sole consumer of both
paths; postgres/redis/ML use local volumes).

Two bugs the umount test caught, both worth knowing:
- `ls` cannot test mountedness. An unmounted mount point is an ordinary
  empty directory, so ls succeeds and recovery was skipped entirely.
- A drop repaired within a single run leaves prev=healthy, so keying the
  restart solely on the stored state skipped it while the container still
  held its stale view.

Also moves the SMB password out of /etc/fstab, which is 0644 and was
readable by every local user, into a 0600 credentials file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 14:49:22 -04:00
Bastian de BylandClaude Opus 5 0939204061 docs(podman): add role README with backup, restore and LibreSign runbook
No restore procedure existed anywhere in this repo. The backup pipeline
is well commented but nothing described how to get data back, and the
README the backup script already referenced was missing.

Covers the shared backup engine and its stage ordering, a restore
procedure for both MariaDB and Postgres with the rootless-podman command
form, data-tree restore including the uid 33 chown, and the
maintenance-mode/files:scan reconcile.

Two things worth stating plainly:
- Untested backups are not a control. No rehearsal has been recorded.
- Do not blindly re-run libresign:configure:openssl on a restored
  instance -- it mints a new root CA and invalidates the trust chain on
  every document already signed under the old one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:55:04 -04:00
Bastian de BylandClaude Opus 5 fec7d62acb feat(skudak-cloud): repair LibreSign, brand its mail, add Redis
LibreSign had been silently broken since it was first deployed in
January. Every step of the old before-starting hook ended in `|| echo`,
so six months of failures logged nothing.

LibreSign repair
- Root cause was a stale config_path: a valid OpenSSL root CA existed at
  generation 1, a failed CFSSL attempt left an empty generation 2, and
  config_path was left pointing at the empty one. Regenerated as
  "Skudak LLP" (was the pre-rename "Skudak Rennsport LLP").
- Deleted the hook. Java/PDFtk/jSignPdf live under data/appdata_*, a
  persisted volume, so they only ever needed installing once. Install and
  verification are now explicit tasks that actually fail.
- PHP_MEMORY_LIMIT 1024M -- the 512M image default fails opaquely
  mid-signature. LC_ALL/LANG so the JVM is not ANSI_X3.4-1968.
- signature_render_mode=GRAPHIC_ONLY. Any other mode halves the stamp
  width and overlays a name/date block that collides with the drawn mark
  and duplicates what our documents already typeset. The value must be
  exactly GRAPHIC_ONLY; a bare "GRAPHIC" is accepted by occ, matches no
  radio in the UI, and silently reverts to default.
- write_qrcode_on_footer=false, written with --type=boolean because
  FooterHandler reads it via getValueBool and the typed appconfig API
  does not coerce a string "0". The validation URL text is kept.
- identification_documents=0 -- the default gates signing behind an ID
  upload plus admin approval, so signers saw no way to sign.
- shareapi_restrict_user_enumeration_full_match=no, so an email owned by
  an existing account can be added as a signer. Root cause is in core
  (MailPlugin.php:163), not LibreSign. Do NOT set full_match_email=no --
  that disables email signer search entirely.

Mail branding (skudakmail app)
- Two supported extension points, no core patch and no LibreSign fork:
  mail_template_class for layout, subjects, button labels and the footer
  LibreSign never adds; and a BeforeMessageSent listener to embed the
  wordmark as a cid: part so it survives remote-image blocking.
- A third listener adds scoped CSS fixing the signing page being clipped
  on iOS Safari (100vh -> 100dvh). Patched upstream too.
- skudakmail-verify.php.j2 asserts all of the above through the real
  useTemplate() path and fails the play on drift. Every assertion was
  proven to fail when deliberately regressed.

Redis
- memcache.locking was unset, so Nextcloud used DBLockingProvider and
  every file lock became a MariaDB write -- the contention behind the
  intermittent multi-second stalls. Verified after: db locks static,
  redis keys growing.
- requirepass lives in a mounted 0640 conf, not --requirepass, which
  would leak it into podman inspect, the systemd unit and ps. The file is
  chowned to uid 999 because redis-server does not run as root and the
  :ro mount stops the image fixing it itself.
- No maxmemory: cache is evictable, locks are NOT, and evicting a held
  lock permits concurrent writers to one file. No persistence either --
  a restored RDB could reinstate locks whose owner is long dead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:54:52 -04:00
Bastian de BylandClaude Opus 5 c184099b2d refactor(backup): remove duplicate direct-to-S3 stage
A host-to-iDrive S3 stage was added here and is now removed. It would
have written the same data into the same `backup-all` bucket that the
TrueNAS cloud-sync task already fills -- duplicate storage, two writers
to one prefix, for no additional coverage.

Offsite to business-owned storage was already solved: the rsync feeds
/mnt/glacier/skudakcloud and TrueNAS cloud-syncs that to Skudak's own
iDrive e2 account. The earlier note in skudak/cloud.yml proposed adding
S3 *and then dropping the rsync* -- replacement, not addition -- and
building both was a misreading of it.

If offsite is ever moved onto this host it must REPLACE the rsync, not
run beside it. Settle first whether the TrueNAS -> iDrive leg is
independently verifiable; keeping this chain means trusting it.

Also removes the now-orphaned /etc/backup_s3 credential file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:54:27 -04:00
Bastian de BylandClaude Opus 5 f11391b28f bound log growth and reclaim ~52 GB of container disk
Caddy was rotating on implicit defaults (100MiB/keep 10/90d) that were not
holding -- 20 rotated files per stream and a 190-day-old .gz, 1.3 GB across
16 log streams. Made explicit at 10MiB/keep 3/7d.

Note roll_size et al are subdirectives of `output file`, NOT of `log`.
Getting that wrong does not degrade gracefully: Caddy refuses to start on a
bad config, so every site went down until it was corrected. Worth a
`caddy validate` gate before reload.

journald had no SystemMaxUse and had reached 4 GB, drifting toward its
10%-of-filesystem default (~190 GB on this root). Capped at 500M.

Both are safe to keep short because fluent-bit ships the journal and every
Caddy access log into Graylog -- though note its GELF output has been
erroring for days, which weakens that premise and wants investigating.

The larger find was unrelated to logs: 896 images totalling 59.6 GB with
75% unused (94 tags of greg-time-bot, 73 of fulfillr -- one per deploy) and
5.4 GB of dangling volumes, mostly 804 MB Nextcloud /var/www/html trees
orphaned by container recreations. Pruned to 22 images / 15.4 GB, and added
a weekly timer keeping 30 days so a rollback still needs no rebuild.

Also dropped the decommissioned 6379/tcp redis rule (nothing listening;
Immich's redis is on the shared podman network) and the orphaned nosql, s3
and searxng volume dirs.

Backup log exclusions turned out to be unnecessary: Gitea logs to console
so its log dirs are empty, Nextcloud already excludes its own, BookStack
mounts only uploads, and Caddy is not backed up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:13:35 -04:00
Bastian de BylandClaude Opus 5 bc110ce69e back up Gitea + Skudak app data; drop PartKeepr and Pi-hole
Extends the Nextcloud backup machinery rather than adding a second
mechanism. cloud-backup.sh.j2 gains three guarded options, all no-ops for
the existing callers:

  backup_podman_user  Gitea runs rootless under `git`, not `podman`
  backup_db_type      postgres (Gitea) and mysql (BookStack) alongside
                      mariadb; each engine's completion trailer differs,
                      and grepping for the wrong one fails every run
  backup_sqlite_dbs   `sqlite3 .backup` for live WAL-mode SQLite, gated on
                      `pragma integrity_check` before promotion -- rsync
                      is either stale (no -wal) or torn (with it)

New instances: gitea-debyl, skudak-gitea, bookstack, partsy-skudak. The
alert handler is rendered once and shared, so its wording is now generic
rather than per-product; TAG stays nextcloud-backup because an external
Graylog rule matches on it.

`apply:` on the includes is load-bearing -- tags on a dynamic
include_tasks do not reach the tasks inside it.

Business data (skudak-gitea, bookstack, partsy-skudak) goes to TrueNAS
and on to Skudak's own iDrive account; the personal bucket's
/skudak*/** excludes are permanent, not a stopgap.

Removals: PartKeepr is superseded by Partsy, and its teardown never
finished -- it targeted /etc/systemd/system/podman-partkeepr*.service,
wrong prefix and wrong scope, leaving enabled user units in failed state.
Pi-hole's role was already orphaned (absent from deploy_home.yml); its
port 53 rule went with it after confirming nothing listens there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 18:08:47 -04:00
Bastian de BylandClaude Opus 5 5776dbe1bf UPS monitoring, Nextcloud cron, container image bumps
ups role: NUT server on home.debyl.io for the CyberPower PR1500RT2U that backs
both it and truenas.localdomain, with staged shutdown (TrueNAS sheds at t+2min,
host at 10% charge) and best-effort IPMI power-on when mains returns. The 10%
threshold leans on ignorelb + override.battery.charge.low rather than a custom
poller, because CyberPower asserts its own low-battery flag far too early.
Credentials come from vault vars; nothing sensitive is templated in the clear.

Nextcloud background jobs: both instances have backgroundjobs_mode "cron", which
expects an external caller every ~5 minutes, and nothing was calling. The
personal instance had not run a background job since 2026-05-14 and skudak since
2024-11-20. Consequently trash and file versions never expired, stale chunked
uploads accumulated, calendar reminders never fired, and nextcloud.log was never
rotated -- which quietly made the existing log_rotate_size cap inert. Added a
systemd timer per instance, skipping cleanly when the container is down or in
maintenance so deploy windows don't show up as failed units.

Trash retention on the personal instance: the default "auto" only expires when
disk space demands it, so 66 GB of >30-day deletions sat on a host with 1.3 TB
free -- effectively unbounded. "auto, 30" makes the 30-day expiry unconditional
while still purging early under pressure.

Image bumps:
  nextcloud       33.0.0  -> 34.0.2   (both cloud and skudak-cloud)
  greg-time-bot   3.9.25  -> 3.10.0
  fulfillr        20260723.2044 -> 20260728.2155  (prod and dev)

The fulfillr bump records what is already deployed: both containers were rolled
to that image on 2026-07-29 for SCRUM-156 (digital product releases + customer
update campaign). Committing it keeps the repo from claiming an older tag than
the host is actually running, which would otherwise roll fulfillr backwards on
the next clean-checkout deploy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 13:00:51 -04:00
Bastian de BylandClaude Opus 5 16145fb6bd only email genuine backup failures
A spurious invocation carries no action for a reader, so mailing it just
trains them to skip past the subject line -- which defeats the point of the
alert. Send mail only when the unit actually reports failure; spurious
triggers still leave their journald record, so they stay greppable and can
still feed a Graylog rule.

Verified both paths: a spurious trigger leaves /var/log/msmtp.log untouched
and logs alert_mail=skipped reason=spurious, while a genuinely failed unit
still composes mail with the FAILED subject.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:34:45 -04:00
Bastian de BylandClaude Opus 5 6e99794d0f restore full dump flags; stop spurious alerts claiming FAILED
mariadb-upgrade --force has now been run against both instances and repaired
the system tables, so --routines and --events no longer abort the dump.
Restore them for completeness. Both verified: rc=0 with a clean completion
trailer, and both Nextcloud instances report installed with unchanged table
counts afterwards.

The upgrade still exits non-zero on these containers because it cannot create
the `sys` schema: /var/lib/mysql is owned by daemon rather than mysql, so
mysqld may not create top-level databases. `sys` is diagnostic only and
unused by Nextcloud, but the same permission would block creating any new
database, so it is recorded in the template comment.

The alert handler claimed FAILED in its subject line regardless of what the
unit actually reported, so starting it by hand mailed out a failure notice
for a run that succeeded. Derive the subject and log line from the real
Result, and emit status=spurious rather than status=failed so such triggers
cannot match a Graylog alert rule keyed on genuine failures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:04:32 -04:00
Bastian de BylandClaude Opus 5 7c72aee4f9 add home_smtp vault key for system mail
Referenced by roles/common/templates/msmtp/msmtprc.j2 to authenticate
outbound alert mail against the OpenSRS relay. Without this committed a
fresh checkout cannot render /etc/msmtprc.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 17:22:30 -04:00
Bastian de BylandClaude Opus 5 0ab423ca55 harden nextcloud backups: db dumps, alerting, drift fix
The data-only rsync left no way to restore a working instance: mysql/ and
config/ were never backed up, so a recovery would have files but no shares,
users or metadata. Dump the database before syncing files (a DB older than
the files is repairable with occ files:scan; a newer one references blobs
that never made it into the backup) and ship config/ alongside it.

Capture the --chmod=Du=rwx,Dgo=rx flag that had been hand-added to the
deployed skudak-cloud script. It was outside git, so every deploy silently
reverted it. It now lives in backup_rsync_extra_args.

Add OnFailure= alerting. The units failed silently before, which is how an
iDrive sync failure sat unnoticed since May. msmtp rather than the esmtp
already installed: the OpenSRS relay is port 465 (implicit TLS) and libesmtp
only speaks STARTTLS.

Exclude nextcloud.log* from the sync and cap log_rotate_size. skudak-cloud
was running at loglevel 0 and had written a 64 GB log that was being rsynced
and pushed to S3; set it to 2 to match the home instance.

Stagger the timers (04:00 / 04:30) so both finish before the 05:00 TrueNAS
snapshot task, and bound TimeoutStartSec so a wedged rsync cannot leave the
unit activating forever and skip every subsequent trigger.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 16:03:18 -04:00
Bastian de Byl 1e1d53ecd8 feat(gitea-actions): install jq in the ESP-IDF CI image
esp32-stm32-vcu's scripts/release.sh uses jq to rewrite the protocol
manifest before publishing it to S3. The image did not have it, so the
release aborted at:

    ./scripts/release.sh: line 61: jq: command not found

The failure mode is nastier than a red job. jq is only reached *after*
the firmware .bin and version.json have already been uploaded, so
clients were served the new build while the git tag, the Gitea release
and the protocol manifest were never written — the repo still showed
the previous release as latest while production served a newer one.

Only jq was missing; aws and sha256sum are already present (verified in
the rebuilt image: jq-1.7, aws-ok, sha256sum-ok).

This surfaced when the skudak firmware jobs were moved into this image
(they previously ran on the default gitea-ci image, which has jq).
2026-07-25 12:01:16 -04:00
Bastian de BylandClaude Opus 4.8 7abec05cc2 SCRUM-148/SCRUM-149: Bump fulfillr + fulfillr-dev images to 20260723.2044
Per-product sales metrics + storefront image proxy endpoints.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 16:52:24 -04:00
Bastian de BylandClaude Fable 5 8703875825 SCRUM-140: Bump fulfillr images to 20260717.1955 (import status ratchet)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 21:36:43 -04:00
Bastian de BylandClaude Fable 5 b380e33c92 SCRUM-140: Bump fulfillr images to 20260717.0202; snipcart_api_key via EXTRA_VARS
The import-only Snipcart key renders empty on normal deploys (import
endpoint 503s); it was injected via EXTRA_VARS just for the historical
refund re-import and is now blanked again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 22:43:13 -04:00
Bastian de BylandClaude Fable 5 6737b4621d SCRUM-140: Bump fulfillr + fulfillr-dev images to 20260716.2108
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 17:19:42 -04:00
Bastian de BylandClaude Fable 5 3f0a78da39 SCRUM-129: Bump fulfillr + fulfillr-dev images to 20260714.2120
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 17:30:32 -04:00
Bastian de BylandClaude Fable 5 51beb906ba gitea-actions: pre-bake esptoolpy + mkspiffs in the PlatformIO CI image
Both are pulled on demand by pio run/buildfs; without them baked in, every
ephemeral job container re-downloads them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 17:01:59 -04:00
Bastian de BylandClaude Fable 5 9dfa812a80 gitea-actions: add PlatformIO CI image for Arduino-framework ESP32 builds
New opt-in job image localhost/gitea-ci-platformio:7.0.1 (pattern matches the
ESP-IDF image): PlatformIO core 6.1.19 with the espressif32@7.0.1 platform,
xtensa toolchain and Arduino framework pre-baked via a seed project so
ephemeral job containers download nothing. First consumer is
Skudak/esp32-web-interface.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 16:59:57 -04:00
Bastian de BylandClaude Fable 5 5588e81fb0 SCRUM-126: Bump fulfillr + fulfillr-dev images to 20260714.2006
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 16:29:39 -04:00
Bastian de BylandClaude Fable 5 66d2502d31 SCRUM-114: Bump fulfillr + fulfillr-dev images to 20260712.1826
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 15:07:54 -04:00
Bastian de BylandClaude Opus 4.8 c4c0067756 SCRUM-108: Bump prod fulfillr image to 20260708.0147
Ship-email delivery visibility + resend-tracking endpoint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 21:55:12 -04:00
Bastian de BylandClaude Opus 4.8 6f158bc2fb SCRUM-108: Bump fulfillr-dev image to 20260708.0147
Ship-email delivery visibility + resend-tracking endpoint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 21:52:43 -04:00
Bastian de BylandClaude Opus 4.8 a78cf91be9 SCRUM-105: Bump prod fulfillr image to 20260706.1823
Promotes the ASCII-safe address transliteration fix (go-fulfillr#24) to prod
after dev verification.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 14:36:44 -04:00
Bastian de BylandClaude Opus 4.8 1cf264bbdf SCRUM-105: Bump fulfillr-dev image to 20260706.1823
Deploys the ASCII-safe address transliteration fix (go-fulfillr#24) to the
staging back-office. Prod image tag unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 14:29:13 -04:00
Bastian de BylandClaude Opus 4.8 a696d8c47b SCRUM-102: Bump prod fulfillr image to 20260703.1535
Deploys the go-store banner-columns build to the prod back-office so the Coupons
admin round-trips the global-sale banner fields on prod. Applied via
`make deploy TAGS=fulfillr`.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:31:25 -04:00
Bastian de BylandClaude Opus 4.8 f9e9adc507 SCRUM-102: Bump fulfillr-dev image to 20260703.1535
Deploys the go-store banner-columns bump (coupon banner_label/banner_text) to
the staging back-office so the Coupons admin round-trips the global-sale fields.
Applied via `make deploy TAGS=fulfillr-dev`. Prod tag unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 21:41:41 -04:00
Bastian de BylandClaude Opus 4.8 704c1a8311 immich: upgrade to v3.0.0 + migrate DB to VectorChord
v3.0.0 dropped pgvecto.rs support, which crash-looped immich-server
("No vector extension found") and left photos.debyl.io on a white page.
Swap db_image to Immich's bundled postgres image (VectorChord 0.4.3 +
pgvecto.rs 0.2.0) so v3 reindexes the existing embeddings into
VectorChord in place; bump server/ml from v2.7.5 to v3.0.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 17:21:14 -04:00
Bastian de BylandClaude Opus 4.8 b8f68acd0e SCRUM-98: Bump fulfillr (prod) image to 20260629.2000
Stripe relink endpoints (stripe-candidates / link-stripe) + ch_ refund support
now live on prod back-office.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 23:05:11 -04:00
Bastian de BylandClaude Opus 4.8 4f6e3f709e SCRUM-98: Bump fulfillr-dev image to 20260629.2000
Stripe relink endpoints (stripe-candidates / link-stripe) + ch_ refund support.
Prod (fulfillr) left on 20260628.1930 pending prod cutover.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 16:10:39 -04:00
Bastian de BylandClaude Opus 4.8 5d0d15f414 SCRUM-97: Healthcheck + restart-on-unhealthy for fulfillr containers
After a power cycle a transient HMAC Secrets Manager blip leaves
go-fulfillr's gated routes unregistered (404) with the process still up,
so nothing restarts it. Add a podman healthcheck probing the new
dependency-free /api/v1/health/startup (503 until those routes register)
with healthcheck_failure_action: restart, so podman restarts the
container in place and the next boot self-heals.

- fulfillr.yml + fulfillr-dev.yml: healthcheck via busybox wget (ships in
  the alpine image), interval 30s / timeout 5s / retries 3 /
  start_period 30s (covers the ~14s HMAC retry backoff), failure_action
  restart. Existing restart_policy on-failure:3 kept (process-exit case).
- main.yml: bump fulfillr + fulfillr-dev image to 20260628.1930 (the
  build carrying the /health/startup probe).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 15:35:22 -04:00
Bastian de BylandClaude Opus 4.8 285ef2ad01 gitea-actions: add python3-yaml + python3-jinja2 to the ESP-IDF CI image
The esp32-stm32-vcu firmware build generates common-yaml headers with
`python3 generate.py`, which needs pyyaml + jinja2. The runner's base Python is
PEP 668 externally-managed (pip install fails) and the IDF venv isn't on PATH in
the docker-exec step shell, so install both as distro packages. Lets firmware
jobs run a plain `python3 generate.py` with no pip and no IDF sourcing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 16:03:36 -04:00
Bastian de BylandClaude Opus 4.8 2df697f5f6 SCRUM-51: Bump fulfillr image to 20260614.1925 (dev + prod)
Adds email_sent/email_failed timeline events for ticket label emails.
Deployed to fulfillr-dev and fulfillr (prod) on home.debyl.io.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 15:31:47 -04:00
Bastian de BylandClaude Opus 4.8 b4dec16cad SCRUM-50: Bump fulfillr image to 20260614.1518 (dev + prod)
Ticket refund + replacement-shipment endpoints and guarded transitions.
Deployed to fulfillr-dev and fulfillr (prod) on home.debyl.io.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 11:25:14 -04:00
Bastian de BylandClaude Opus 4.8 16053e1cbb fulfillr: drop Snipcart key, add outreach/recovery schedule config, bump image
- remove snipcart_api_key from dev/production config (Snipcart decommissioned
  post-migration)
- add review-outreach and cart-recovery schedule_name/schedule_group blocks
  (dev + prod) for the EventBridge-driven outreach and cart-recovery jobs
- bump fulfillr image 20260607.0217 -> 20260613.0117

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:19:56 -04:00
Bastian de BylandClaude Opus 4.8 a30ff9b165 gitea-actions: add ARM/Python CI deps and SSH bind-mount for submodule clones
- Containerfile.ci: add python3-yaml + python3-jinja2 and the
  gcc-arm-none-eabi / binutils / libnewlib toolchain for embedded builds
- bind-mount the runner's SSH key + known_hosts read-only into each job
  container at /root/.ssh so submodule clones over
  ssh://git@git.skudak.com:2222 succeed; staged as a dedicated
  container_file_t-labelled ci-ssh copy (tasks/user.yml) and allowlisted
  via valid_volumes (config.yaml.j2)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:19:45 -04:00
Bastian de BylandClaude Opus 4.8 7d4a398bba Drop self-hosted AI (Ollama + SearXNG); gregtime switches to xAI Grok
The Ollama role and SearXNG container backed FISTO AI responses in the
greg-time Discord bot. greg-time 3.9.6 drops both (plus the Gemini path)
in favor of a single xAI Grok backend, so:

- remove the ollama role and its wiring in deploy_home.yml
- remove the searxng container task, template, and searxng_path default
- gregtime: swap OLLAMA_*/SEARXNG_URL/GEMINI_API_KEY env for XAI_API_KEY,
  bump image 3.6.5 -> 3.9.6
- vault: add xai_api_key, drop gemini_api_key

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:19:45 -04:00
Bastian de BylandClaude Opus 4.8 87cf953364 SCRUM-45: Revert Caddy /webhooks/easypost carve-out
The EasyPost tracker webhook moved to debyltech-api (publicly reachable Lambda);
the fulfillr host is LAN-restricted and no longer hosts it, so the carve-out is
no longer needed. Removes the handle blocks for prod and dev.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 21:13:56 -04:00
Bastian de BylandClaude Opus 4.8 c896f69ff9 SCRUM-45: Caddy carve-out for the EasyPost return webhook
The Fulfillr host is IP-restricted, so EasyPost's servers can't reach it. Add a
narrow `handle /webhooks/easypost` before the IP restriction (handle blocks are
mutually exclusive, first match wins) for prod (:9054) and dev (:9055) so the
HMAC-verified tracker webhook is reachable while the rest of the host stays locked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 20:29:44 -04:00
Bastian de Byl 3b9c46a11b fulfillr prod: bump to 20260606.2328 (immutable note write handler) 2026-06-06 19:34:10 -04:00
Bastian de Byl 7c58a2a358 fulfillr prod: bump to 20260606.2231 (immutable notes + go-store v0.2.1) 2026-06-06 19:24:14 -04:00
Bastian de Byl 2ce6c531ee fulfillr prod: bump to 20260606.1840 (go-store v0.2.1 order INSERT fix) 2026-06-06 18:29:36 -04:00
Bastian de Byl cc0cb2911f vault: correct fulfillr_prod_store_auth_token (was invalid -> 401) 2026-06-06 17:30:38 -04:00
Bastian de Byl 2335b4980d fulfillr(prod): wire prod Turso store + live Stripe (fulfillr_prod_* vars) + image 20260606.1735 2026-06-06 17:28:00 -04:00
Bastian de Byl da98a2c5dc fulfillr(prod): add download_base_url=https://api.debyltech.com to production.json.j2 (cutover prep) 2026-06-06 16:55:55 -04:00
Bastian de Byl 5d1db841f0 fulfillr-dev: bump to 20260606.1735 (no double shipped-email) 2026-06-06 14:40:08 -04:00
Bastian de Byl 1f16749935 fulfillr-dev: bump to 20260606.1727 (importer fixes + tickets/custom-shipment on Turso) 2026-06-06 13:37:52 -04:00
Bastian de Byl fcde86153c fulfillr-dev: bump to 20260606.1639 (refund + internal notes) 2026-06-06 12:42:58 -04:00
Bastian de BylandClaude Opus 4.8 e149d860d5 gitea-ci: add zip to the CI image
Lambda packaging steps in some workflows shell out to `zip`; the image
only had `unzip`. Add `zip` alongside it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 11:46:23 -04:00
Bastian de Byl bafc32226c fulfillr-dev: bump to 20260606.1523 (resend downloads + new-products-inactive seed) 2026-06-06 11:33:22 -04:00
Bastian de Byl 935de1fcfe fulfillr-dev: download_base_url for resend-download links 2026-06-06 11:32:42 -04:00
Bastian de Byl a024078a55 fulfillr-dev: bump image to 20260606.1425 (digital file upload + download-admin + tickets payment refresh) 2026-06-06 10:29:49 -04:00
Bastian de Byl 35213d81c3 fulfillr-dev: point aws.bucket at debyltech.digital.dev (digital file uploads) 2026-06-06 08:42:33 -04:00
Bastian de BylandClaude Opus 4.8 e82ace6de3 fulfillr-dev: staging back-office container + Turso store prep
Add a second go-fulfillr container (fulfillr-dev) wired to the staging
Turso store + EasyPost/Stripe test keys via dev.json, served at
fulfillr-dev.debyltech.com (Caddy -> :9055), LAN-restricted like prod.

- fulfillr-dev.yml + dev.json.j2: the staging container, volumes, config
- defaults: fulfillr_dev_* vars; prod store URL stubbed off until cutover
- Caddyfile + caddy.yml: fulfillr-dev site block and static mount
- awsddns.yml: Route53 DDNS for the fulfillr-dev hostname
- production.json.j2: add store_database_url/store_auth, rename stripe key
  var to fulfillr_stripe_api_key
- vault.yml: dev + store/stripe secrets

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 00:23:07 -04:00
Bastian de BylandClaude Opus 4.8 2640d09cb5 gitea-actions: run CI jobs in rootless-podman containers
Switch the act_runners from :host execution to docker:// images backed by
a rootless podman socket under the gitea-runner user, so each job runs in
its own ephemeral container with per-job Go caches. This eliminates the
cross-repo GOMODCACHE/go-build poisoning that forced the debyl runner to
capacity:1.

- deps.yml: enable the rootless --user podman.socket, ensure subuid/subgid,
  register gitea_runner_uid; drop the rootful system socket override,
  podman-docker and host golang
- images.yml + Containerfile.ci/.espidf: build localhost/gitea-ci and
  localhost/gitea-ci-espidf into the runner's rootless image store
- config.yaml.j2: docker:// labels (per-runner overridable), docker_host
  -> rootless socket, force_pull false
- act_runner.service.j2: XDG_RUNTIME_DIR + DOCKER_HOST -> user socket
- defaults: uniform capacity:4 (drop the debyl capacity:1 workaround);
  esp_idf_version now tags the espressif/idf-based image
- main.yml: import images.yml, drop the host esp-idf install (firmware jobs
  use the espressif/idf job container instead)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 00:16:54 -04:00
Bastian de Byl 72ecc63e17 fulfillr-dev: bump image to 20260606.0357 (inventory editor, logs page, branded shipped email, U5 trim) 2026-06-06 00:10:39 -04:00
Bastian de BylandClaude Opus 4.8 2df5b7fc03 Deploy fulfillr 20260603.0222 and wire tickets_table
Bump fulfillr image to the build with the tickets feature, and add the
tickets_table to the fulfillr production.json config (new debyltech-tickets-prod
DynamoDB table) so the /api/v1/tickets routes register.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 22:32:52 -04:00
Bastian de BylandClaude Opus 4.8 5e189289e7 fulfillr: deploy Stripe payment requests (key + image 20260530.2348)
- add stripe_api_key to fulfillr production.json template
- add restricted Stripe key to ansible vault (encrypted)
- bump fulfillr image to the CI build containing the Stripe endpoints

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 20:58:40 -04:00
Bastian de BylandClaude Opus 4.7 1bc1a7f619 chore: bump fulfillr container to 20260527.2345
Records the back-in-stock notify-route fix image now running in prod.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:08:19 -04:00
Bastian de Byl 4287f5774f noticket - updates, cleanup, housekeeping 2026-05-27 11:19:09 -04:00
Bastian de BylandClaude Opus 4.7 0249044475 chore: bump fulfillr container to 20260519.0014
Picks up /api/v1/orders/search smart-search endpoint.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 20:17:19 -04:00