Nothing was group-restricted, so customers saw Dashboard, Photos, Office
and the rest, plus Nextcloud's first-run and promo apps.
- Disable for everyone: firstrunwizard, recommendations, related_resources,
weather_status, survey_client, support, app_api, contactsinteraction,
photos. A refused disable now fails the play (occ exits 0 on "can't be
disabled").
- Restrict dashboard and office to staff. defaultapp=dashboard,files so
staff land on the dashboard and customers fall through to Files.
- libresign groups_request_sign pinned to staff and asserted in verify.
LibreSign itself is deliberately NOT group-restricted: that also blocks
anonymous requests and would break public signing links.
- profile.enabled=false; lookup_server="" (lookup_server_connector
can't be disabled).
- README: what customers can open.
Checked as a probe customer: apps=files,activity,libresign,text,viewer,
lands in Files, can't request signatures; staff land on Dashboard.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Customers are admin-created accounts in per-customer groups. They should
read what staff share with them and nothing else. Checked against a test
customer in the UI, and by running the sharee and contacts-menu search
services as that user.
- shareapi_exclude_groups=allow, list [admin]: only staff can share. Exclude
mode ("yes") only disables sharing for users whose groups are ALL
excluded, so it never catches a customer in their own group.
- User/group autocomplete off. By default a customer typing "bas" found the
owner's account. Staff share by exact group name; LibreSign signers are
found by email.
- New shares default to View only (shareapi_default_permissions=1).
- files default_quota 0 B, so customers get no personal storage. Staff in
cloud_debyltech_staff_users are exempted (skipped if not yet created).
- No "Leon Green" sample contact in new address books.
- The verify script fails the deploy if any isolation setting drifts.
- README: staff and customer onboarding checklist.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Gitea expands a job's matrix when the run is created, before plan has any
outputs, so fromJSON(needs.plan.outputs.matrix) collapsed to a single empty
"Build ${{ matrix.key }}" job and nothing was ever built. The matrix is now
the fixed list of image keys; plan emits every image's spec with a build flag
and each matrix job looks its own entry up, no-opping when it wasn't picked.
Also document that both registry tokens need write:package -- the vault token
was read-only, so the gitea_ci_build_local bootstrap failed its push.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- theming background -> backgroundColor, dropping the stock image for the
navy theming colour.
- skeletondirectory/templatedirectory set to "" so customer accounts start
empty instead of getting Nextcloud's sample Manual, intro video and
Templates folder.
- The system-config compare now tells an unset key apart from an empty
value. Both print nothing, so an intentionally empty setting was
silently skipped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Enable LibreSign's email identify method (click-to-sign, no account
creation), remove the stamp background and collect signer metadata.
On Skudak these were only ever set in the admin UI. Without the email
method a fresh instance answers "No signers." for any outside address.
The verify script now asserts it.
- Store add_footer=true explicitly. The code already defaults to it, but
the 14.2 admin page shows unset as unchecked.
- trusted_proxies = the container's own address, read per deploy. Behind
rootless port forwarding every request arrives from it, so without this
X-Forwarded-For was ignored and every client shared one IP for
brute-force throttling.
- maintenance_window_start and default_phone_region, which clears the setup
warnings.
- Mail declares color-scheme "light only" so Apple Mail's dark mode doesn't
repaint the white ground and bury the black wordmark (asserted in verify).
- Idempotency: redis image fully qualified (docker.io/library/...), the
debyltechmail copy owned by the mapped www-data uid, and theming
compared before setting.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A de Byl Technologies LLC Nextcloud cloned from the Skudak instance:
LibreSign signing for people without an account, registration off
(admin-created accounts only), no Group Folders. DNS is a terraform-managed
ALIAS to fulfillr.debyltech.com.
- containers/debyltech/cloud.yml: nextcloud/mariadb/redis on port 8091.
It installs unattended on the first deploy, sends mail through SES as
noreply@debyltech.com, and re-asserts the Skudak LibreSign settings.
- files/debyltechmail: skudakmail rebranded, with a new black-and-white
wordmark and white web-UI logos.
- LibreSign is pinned to 14.2.2 from the GitHub release (sha256-checked)
rather than `occ app:install`. The app store served a same-day 14.2.3
whose tarball has no binary-signature metadata. 14.2.x also doesn't
create its own download dirs, so they're pre-created.
- The backup runs nightly at 04:15 to TrueNAS /mnt/glacier/debyltechcloud and
reaches personal iDrive via the "iDrive E2 Backup" task; the TrueNAS side
excludes /debyltechcloud/_backup/config/**.
- Fix the libresign:configure:check gate in both instances: '\berror\b'
becomes a backspace in Jinja and never matched, so a check reporting three
errors passed clean. Now '\\berror\\b'.
- vault: cloud_debyltech_* secrets.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Bedroom Light HS200 dropped off Wi-Fi for 5 s at 03:01 and came back
reporting "on"; the Bedroom On device trigger treated unavailable -> on as
someone flipping the switch and lit the bedroom Hue lamps at 100%.
Bedroom On/Off and TV On/Off now use state triggers with not_from
unavailable/unknown, so reconnects and HA restarts no longer count as a
flip. The driveway's switch-reconnect catch-up is dropped for the same
reason (it would undo a manual off); the HA-restart catch-up stays.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sunset automation required TV mode off, so an afternoon of TV skipped
the driveway string lights entirely (Sep 22 and 23) - not an outage. The
driveway now has its own sunset -1h automation, with catch-up on restart
or switch reconnect before 23:00.
Evening brightness lives in one script (evening_lights_apply) that blends
between the old step levels; a 5-minute ramp from 20:30 eases lights that
are on and leaves alone any a person has changed by hand. The Dining Hall
no longer bumps to 50% at 21:30.
TV off after 23:30 now only turns off the living room glow instead of
bringing the whole house back to full brightness, and TV on/off leave the
lights alone in daylight. Lights-out moves to 23:30, and a 01:00 sweep
catches anything switched back on at the wall.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1.26.1 -> 1.27.3 picks up the security fixes in 1.27.0 through 1.27.3.
The image was one shared variable, so the two instances could only move
together; split it into gitea_debyl_image / gitea_skudak_image and tag
the debyl tasks gitea-debyl so each can be upgraded and verified on its
own. Skudak stays on 1.26.1 in this commit.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The daily quote keeps a ledger of what it has posted and re-rolls ZenQuotes
until it finds something the channel has not read, with the header framing and
the offline fallback pool drawn against that same ledger.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The runner's job images were built by ansible into localhost/ only, so the
nightly CI prune deleted them and every idle stretch ended with CI failing in
under a second on `docker pull localhost/gitea-ci:latest` until someone re-ran
the role and waited out a rebuild. The previous commit moved them to the Gitea
registry; this moves the *build* off the deploy path entirely.
- .gitea/workflows/ci-images.yml builds files/Containerfile.* and pushes to
git.debyl.io/gitbot/. Per-image change detection, so an ESP-IDF pin bump does
not rebuild the other two; weekly schedule for base-image updates; a
workflow_dispatch selector. PRs build under a throwaway :pr-<n> tag and drop
it -- the build lands in the live runner's store, and act_runner will not
re-pull a tag it already has, so a PR using the real tag would hand every
later job on this host an unmerged image.
- The Containerfiles stop being ansible templates: their version vars are now
--build-arg, read by the workflow out of the same defaults/main.yml the role
interpolates, so CI and ansible build the same bytes from one set of pins.
- LABEL io.debyl.ci-base moves into each Containerfile so neither builder can
forget the prune exemption; the workflow re-checks it before pushing.
- roles/gitea-actions pulls instead of building. gitea_ci_build_local=true
restores the local build+push for seeding a cold registry or when CI is
down -- the workflow that builds gitea-ci runs in gitea-ci.
- Lint .gitea/ alongside ansible/, and document the flow in the role README.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nights are checkboxes again: a household ticks every night that works, which
also says which nights do not. At least one tick (or "Any of these works")
is required.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The CI job images only existed under localhost/ in the gitea-runner store,
and the nightly CI prune deletes any image older than 48h that no container
holds. After every idle stretch CI failed in 0-1s pulling
localhost/gitea-ci:latest until the role was re-run and the images rebuilt.
- Build under git.debyl.io/gitbot/..., push after every run, and pull from the
registry instead of rebuilding when the Containerfile is unchanged.
- Log gitea-runner in via ~/.docker/config.json, which both act_runner (job
image pulls) and podman read.
- Label the base images io.debyl.ci-base and skip that label in the CI prune;
its `until` counts from build time, so a re-pulled image would otherwise be
deleted again the next night.
Workflows pinning `container: image: localhost/gitea-ci-*` must move to the
registry paths.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Render an analytics block (property 353859448 + service-account key) into the
fulfillr dev and prod configs once fulfillr_ga4_credentials is in the vault.
Without the vault var the block is omitted and the portal reports GA as not
connected. Remember to restart the container after deploy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Compacts the admin invite table so it fits without clipping: short token link
with Copy/Msg, status as an emoji, Qty, allergens/notes behind popovers, and
relative update times.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fixes the admin invite table clipping its Edit/Revoke column, and adds
default-headcount people estimates to the invited/awaiting summary.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds the anonymized "Who's coming" card for guests who have answered, and a
message-template copy button on the admin invite table.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A single Go binary with SQLite, built and loaded as localhost/rsvpd:<VERSION>
by make deploy-remote in ~/src/rsvp-debylio. Both of its listeners are
published on 127.0.0.1 only: Caddy proxies the public one to everyone and the
admin one (/admin, no login) only to caddy_local_networks. A Caddyfile mistake
alone cannot expose admin, and neither can a port mistake alone.
Guests' invite links are the credential and they sit in the URL path, which
shapes the vhost:
- It does not import common_headers. That snippet sets Referrer-Policy
same-origin, which would replace the app's no-referrer and let a token leak
in a Referer header.
- Its access log rewrites request>uri to /i/REDACTED and drops the Location
response header, since every POST 303s back to /i/<token>.
- Caddy's error logger is separate from the site's and wrote the raw URI to
caddy.log when the upstream was down. The global log now excludes
http.log.error.rsvp and a filtered rsvp-errors logger takes it instead.
Verified with zero token occurrences in both logs, locally and live.
The data directory is owned directly by the host uid of the container's uid
10001 (subuid + 10000). Setting it to the podman user and chowning back each run
flipped ownership on every deploy and briefly locked the app out of its
database.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
podman-check deleted the old container as soon as the pinned image differed,
and the create task pulled afterwards. A tag that did not exist, or a registry
that was down, therefore left the service with no container at all. The pull
now happens first, so that failure stops the play with the old container still
running. localhost/ images are built and loaded by hand and are never pulled.
The pull is skipped when the container does not exist yet: there is nothing to
protect, and containers[0] is not there to compare against. Without that guard
the first deploy of any new service failed on the conditional -- rsvp was the
first to hit it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A Pi 3B+ (stickah.local) shares a Phomemo PM246 to the LAN as a plain CUPS
queue, so any machine can print 4x6 labels -- fulfillr-site's shipping labels
in particular -- without installing the vendor driver, which is x86-64 only.
The role builds the TSPL CUPS driver from source instead.
It is Debian, not Fedora, so it lives in its own inventory and playbook
(make deploy-labelprint / check-labelprint) and the home.debyl.io roles can
never run against it. make bootfs renders its cloud-init first-boot files onto
a freshly imaged SD card from the same templates the role uses. The Wi-Fi
credentials for the home and rescue networks are in the vault.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rolling a world back meant hand-work over ssh: stop the service, move the live
save aside, unzip the right archive, chown into the container's subuid range,
relabel, start. That is the wrong shape of task to do by hand, and it is always
done under time pressure -- by construction, because the archive you want is
being deleted while you work.
PZ keeps BackupsCount=10 per set and writes one every BackupsPeriod=30 minutes,
so a periodic backup is reachable for about five hours and then gone. On
2026-09-05 the snapshot the admins asked for (05:16, four minutes before the
incident) had about 90 minutes of life left when the request came in.
Same shape as the wipe: the Discord bot writes a trigger file into its own rw
volume, zomboid-restore.path notices it, and zomboid-restore.service runs the
script as the podman user. The bot gets no ssh, no systemd, and keeps only its
existing read-only mount of the Zomboid volume.
Two details carry most of the correctness.
Resolution is by mtime, not by index. The rotation renames the files -- today's
backup_7.zip is backup_8.zip half an hour from now, and a new backup_7.zip holds
a different world -- so an index is valid only while the listing is fresh, which
is not long enough to survive a human reading a confirmation prompt. The trigger
names a set and an mtime; the script resolves the path itself, whitelists the
filename, and refuses if nothing matches. It never accepts a path.
Everything that can fail is checked before the server is touched. A rotated-out
target, an archive with no debbzoid world in it, a bad action, a traversal
attempt in the set name: each aborts with the server still running and writes a
result file the bot reports back. The live world is moved aside rather than
deleted, so a restore is undoable and the last three are kept.
One thing PZ does not advertise: its backups do not cover the whole save
directory. blam/, a mod's own state, is in none of them -- not the 05:16 archive
and not the newest one. Restoring only what the archive holds therefore lands
the world slightly *behind* the target rather than on it, so anything present in
the displaced world and absent from the archive is carried across.
The gregtime tag moves to 3.17.0 for the bot half of this -- `backups`,
`restore <n>`, `restore confirm`, `restore undo`, gated to the same two admins
as the wipe. That image is built and running on the host; its source is not
committed yet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
The vendored Sophie preset and the running debbzoid world had drifted, and the
repo only held the preset. That is not a restore point: force-pushing it would
have reverted the admins' in-game tuning rather than recovering it, which is
exactly how a Sophie world quietly became an Apocalypse one once already -- 139
values reverted, loot from 0.35 back to 0.9, CharacterFreePoints 0 to 60.
files/zomboid/sophie/SandboxVars.lua is now a snapshot of the live world taken
2026-08-31, not the preset as shipped. The modlist is untouched and still
upstream, which is why zomboid_preset_version now names the two halves and their
separate dates. server.ini.j2 carries the eight keys that had drifted:
PlayerSafehouse false -> true
SafehouseAllowNonResidential false -> true (the diner/gas-station case)
SafehouseAllowRespawn false -> true
SafehouseAllowLoot true -> false
SafehouseAllowFire true -> false
TrashDeleteAll false -> true
MapRemotePlayerVisibility 1 -> 4
ResetID 6953472 -> 826046
ResetID is in that list on purpose, and matters most. It is the world's
soft-reset token: a file value that differs from the one the live world was
created with tells every connected client to roll a new character. Carrying the
live value makes a deliberate force-push a no-op instead of a server-wide wipe
prompt.
Spawn config gets its own switch, zomboid_spawn_force. spawnregions.lua and
spawnpoints.lua are the only config a running world re-reads -- at every server
start, where SandboxVars is read once, when the world is created -- so a spawn
edit is deployable on the live world without a wipe. Sharing zomboid_config_force
between them would have meant force-pushing the whole preset to land a one-line
spawn edit, rewriting the INI (hence the ResetID hazard above) and the world's
SandboxVars along with it.
config-template/ now tracks the repo unconditionally. Nothing on the host writes
that directory and the server cannot see it, so it has no hand edits to protect;
if it does not track the repo it is not a restore point, just an older world's
settings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
After the host rebooted, truenas began mailing NOCOMM continuously. The UPS
driver was fine throughout -- nut-driver@cyberpower stayed running and the
CyberPower is still on USB here. What died was upsd, which publishes its state
to truenas over the network:
upsd: not listening on 192.168.1.10 port 3493
upsd: Fatal error: some listening interfaces were not available
nut-server.service: Start request repeated too quickly
upsd binds an explicit address, but the packaged unit only orders itself
After=network.target, which is satisfied when networking STARTS rather than
when an address exists. It tried to bind 11 seconds into boot, before
NetworkManager had assigned the address, then burned all five default restart
attempts inside one second, tripped the start limit and stayed dead. The
shipped unit carries these two lines commented out, because upstream knows the
case.
network-online.target is the correct ordering; NetworkManager-wait-online is
enabled here so it genuinely waits for addresses. StartLimitIntervalSec=0 and
RestartSec are belt and braces: a slow address can no longer exhaust the
attempts, and retries are spaced instead of hammered.
truenas needs no change -- it is a correctly configured SLAVE that reconnected
on its own, and its NUT config is regenerated from the middleware database
anyway. Verified: upsd listening on both addresses, UPS OL at 100%, and truenas
querying it again with zero failures since.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
restorecon -Frv over the podman volumes tree descends into two SMB shares
mounted inside it -- volumes/photos/immich and volumes/photos/storage, both
from truenas -- so it was relabelling the entire remote photo library over the
network, on a filesystem that cannot store SELinux xattrs at all.
On 2026-08-28 that pinned CPU#0 at 100% system time and the kernel logged six
escalating soft lockups:
watchdog: BUG: soft lockup - CPU#0 stuck for 1423s! [restorecon:132780]
New process creation starved, so sshd accepted connections and then hung during
session setup, and the host needed a hard reboot. An earlier run the same day
had already been SIGKILLed at 49s, which was the same bug surfacing quietly.
-x keeps it on the local filesystem. Measured: 1,279,231 local files walked and
relabelled in 29.2s, against never finishing before. The mounts are also
x-systemd.automount, so merely walking into them triggers a mount -- there was
never anything there to relabel.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Smart search, face detection and OCR were all failing with
Machine learning request to "http://immich-machine-learning:3003" failed
while the container still reported Up. Podman only sees PID 1: gunicorn's
master was alive and holding the listening socket, but its worker had died at a
WORKER TIMEOUT and was never respawned -- hence a connect timeout rather than a
refusal. The dead worker was a zombie whose remaining thread was stuck in
uninterruptible sleep in exit_mmap, so it survived SIGKILL, podman rm -f and
rm -f -t 0, and kept the container name and network alias until the host was
rebooted.
MACHINE_LEARNING_MODEL_TTL=0 addresses the cause rather than the symptom. The
default unloads models after 300s idle, so every search following a gap
reloaded four of them (CLIP, buffalo_l detection + recognition, PP-OCRv5) on a
CPU-only 4-core box and then tore those mappings back down -- and that teardown
is what wedged. Keeping them resident costs ~1-2 GB and removes the path.
IMMICH_MACHINE_LEARNING_URL is now set explicitly instead of relying on
immich's implicit default, so the ML container can be renamed without silently
losing search, and the name is a variable so a replacement can be stood up
beside a broken one without editing tasks.
Verified after deploy: zero ML failures, "in-memory cache with unloading
disabled", ping 200 from immich-server, 26/26 containers healthy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
podman_prune_users listed only the podman and git users, so the two stores that
turn over fastest were never touched. gitea-runner had reached 1205 images /
113.1 GB with 100% of it reclaimable, and actions-runner had 137 exited job
containers. That layer count is what makes overlayfs lookups -- and so CI
itself -- slow; the disk was the lesser problem.
Split into two policies, because the stores are not the same kind of thing:
service users keep the 30-day rollback window, and their containers are
deliberately NOT pruned. They are the live services, and reaping one that
merely happens to be stopped would turn a transient crash into a unit that
cannot start again until the next deploy.
CI users get 48h and their exited job containers reaped too. Build layers
carry no rollback value. Containers are reaped BEFORE images on purpose: an
exited container pins the image it ran from, so pruning images first would
leave those layers behind for another day.
Timer moved weekly -> daily; a week of CI turnover is what let the store reach
113 GB between runs. Persistent=true is kept so a missed run catches up.
First run reclaimed 134 GB: gitea-runner 113.1 -> 4.2 GB, actions-runner
7.6 GB -> 0, podman 25.7 -> 15.1 GB. Disk 449G -> 315G, all 26 containers up.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three findings from investigating sustained I/O pressure on the root SSD.
Grouped because the logging and storage edits land in the same task file.
rsyslog was writing a second complete copy of the journal to
/var/log/messages: 6.7 GB of rotated copies, two weekly files of which were
2.8 GB and 2.5 GB. It loads imjournal, so it reads the journal directly and
ForwardToSyslog=no alone does not stop it -- the unit itself has to go.
Verified nothing consumes those files first: fail2ban runs backend=systemd and
matches on the journal ("No file is currently monitored"), and lsof showed only
rsyslogd holding them. Measured afterwards: writes 46 -> 23 GB/day, /var/log
6.7 GB -> 655 MB, journald still capturing container stdout.
The SSD was on bfq, which fedora's stock 60-block-scheduler.rules picks for any
rotational=0 disk. bfq is built for spinning disks and desktop interactivity:
it costs CPU per request, lets reads queue behind write bursts, and hard-caps
nr_requests at 64. With 26 containers and two CI runners writing at once that
is the wrong trade. mq-deadline rather than none because this is SATA with a
32-deep NCQ queue, not NVMe -- the merging and the read-expiry deadline both
earn their place.
Writeback was at the stock percent-of-RAM ratios, so on 31 GB the kernel would
sit on 3.1 GB before starting writeback and 6.2 GB before blocking writers.
Flushing that to a QLC drive that falls to ~80-160 MB/s once its SLC cache is
spent takes tens of seconds with everything stalled behind it. Capped in
absolute bytes instead: one long stall traded for frequent short ones.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
assistant.debyl.io returned 400 on every request. With use_x_forwarded_for
set, home assistant hard-fails any request carrying X-Forwarded-For from an
address outside trusted_proxies:
Received X-Forwarded-For header from an untrusted proxy 169.254.2.1
caddy runs --network host and proxies to localhost:8123, so the source address
is whatever pasta presents inside the container's netns. Rootless podman
switched from slirp4netns to pasta in v5 (this host runs 5.7.1), which moved
that address from 10.0.2.x -- covered by the existing 10.0.0.0/8 entry -- to a
link-local tap0 address that nothing in the list matched. The container has no
10.x address at all any more.
Latent since the pasta migration; it only surfaced when the site was next
opened. Reproduced directly:
curl -H 'X-Forwarded-For: 1.2.3.4' http://localhost:8123/ -> 400
Bumped to 2026.8.3 in the same pass (identical digest to the stable tag).
The config directory was snapshotted first: the 2026.5.1 -> 2026.8.3 database
migration is not reversible.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reverts the restore-from-template behaviour. A wipe now resets the world and the
player database only; Server/<name>_{SandboxVars,spawnregions,spawnpoints}.lua
carry over untouched, so hand tuning survives it. Verified by checksum: all three
files byte-identical either side of a wipe.
That gives up the guarantee the restore bought. PZ writes the running world's
settings back over those files, so if a world is ever created with defaults the
file inherits them and later wipes regenerate from them -- which is how a Sophie
world became an Apocalypse one. Nothing corrects that automatically now, so the
wipe takes a snapshot of the settings before it starts, keeping the last ten
under config-backup/. Greg's edits were lost once because the only record of them
was a file the server had since overwritten; that is the hole this fills.
config-template/ still holds the pristine Sophie preset to copy back from.
The snapshot is best-effort throughout. The first version of it created the
directory with plain mkdir, the parent belongs to the container's subuid rather
than the podman user, and set -e turned that into a failed wipe -- the same shape
as the tee that broke this script before. Ansible owns the directory now and
every step of the snapshot tolerates failure, because a backup that cannot be
written is not a reason to refuse to wipe.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A wipe was quietly turning a Sophie world into an Apocalypse one. 139 sandbox
values had reverted -- loot rates from 0.35 back to 0.9, ranged weapons and ammo
to 2.0, CharacterFreePoints from 0 to 60 -- and ten of Sophie's mod-added
options had vanished entirely.
The reset script was not deleting the settings. PZ reads
Server/<name>_SandboxVars.lua only when it creates a world, and then writes the
running world's settings back over that same file. The file is an output, not an
input. So once any world came up with defaults, the server stamped those defaults
into the file, and every wipe afterwards regenerated from them -- inheriting the
corruption rather than causing it, and with no way back out on its own. The same
shape as the admin-password deadlock.
The intended settings now live in config-template/, which the server has no
reason to touch, and a wipe restores Server/<name>_{SandboxVars,spawnregions,
spawnpoints}.lua from there before restarting. That is also the file to hand-edit
when changing the world: edit the template, wipe, and the new world has the edits.
Verified by wiping: 139 differences before, 0 after. The server still rewrites
the live file on boot -- 996 keys become 1061 as it adds newer options -- but
every Sophie value survives that now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A Build 41 mod (versionMin=41.60) carried into Sophie's B42 list. It indexes
the ModOptions framework at VehicleDoorsHotkey_Options.lua:4, that framework is
not in the list, and so it dies there every boot:
attempted index: ModOptions of non-table: null
Twenty-one exceptions a boot for a hotkey that has never once worked. Excluded
rather than fixed by adding ModOptions, since nothing else in the list wants it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reset script died on its very first line. log() piped through tee into
zomboid/logs/world-reset.log, that directory was owned by the container's
subuid, and the script runs as the podman user -- so tee returned EACCES and
set -e killed the run before it stopped the server or touched a save. Every
`@bot` reset since had been a no-op that reported nothing.
Nothing mounts zomboid/logs into a container; it only holds output from
host-side helpers running as the podman user, so it is now owned by that user
rather than by the subuid the container volumes need.
log() no longer treats the file as load-bearing either. stdout is already
captured by the journal, so an unwritable log is worth continuing past rather
than aborting a wipe over.
Verified end to end by writing the trigger exactly as the bot does: server
stopped, saves and player database deleted, service restarted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The dedicated server (app 380870) and the client are packaged separately and
have drifted. Base.Log_Stack_01 was removed from the client in 42.16, but the
server package still ships entity_logstack.txt.
The server registers every script it can see into a new world's
WorldDictionary, so the entry lands in the save, and every client then fails at
world load with:
WorldDictionaryException: [SpriteConfigs] Missing dictionary script on
client: Base.Log_Stack_01
Nothing in multiplayer can legitimately reference a script that no client has,
so dropping it server-side costs nothing.
The strip runs after the SteamCMD update rather than once at install time on
purpose: validate restores the file on every boot, so it has to be removed on
every boot too.
Note this does not repair a world whose dictionary already records the entry --
that save still needs regenerating.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`@bot restart` saves, sends RCON quit, and relies on the unit's Restart=always
to bring the server back. It never came back. The container sat at exited(0)
while the unit reported active/running with NRestarts=0, so the bot waited for
a startup that was never going to happen and the server stayed down until
someone noticed.
podman generate systemd emits Type=forking with ExecStart=podman start and a
PIDFile pointing at conmon. podman start returns immediately, so the process
systemd was told to supervise was never its child -- it warns about exactly
this in the journal, once per poll, and then does not notice the exit:
zomboid.service: Supervising process 2431312 which is not our child.
We'll most likely not notice when it exits.
That PIDFile is also stale by design: it embeds the container ID, so every
deploy that recreates the container leaves it pointing at nothing.
Type=simple with `podman start -a` keeps podman in the foreground as systemd's
own child, so the exit is seen and Restart=always does what it always claimed
to. stdout and stderr are discarded deliberately -- attaching re-emits the
container's output for systemd to capture straight back into the journal, which
is the flood the k8s-file driver exists to prevent. podman logs still has it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The anti-scanner rules dropped 53-byte Steam query packets above 1/hour with a
burst of 2, per source IP, on the theory that only scanners send queries and
that real players would already have been marked verified by the priority-2
rule when they sent something larger.
That premise is inverted. A client's first contact with the server IS a 53-byte
query, so nobody can be verified before querying, and nobody can query more
than twice an hour without being dropped. The counters said so plainly: five
packets had ever matched the verified-accept rule, against 30,563 drops.
fail2ban then banned each dropped player for a week -- 32 live bans, 550 total,
firing every ten minutes, every one of them a residential address.
Keeps the shape of the protection and moves the threshold somewhere no real
client reaches: opening the server browser or retrying a connection is a
handful of queries, a flood is thousands. Removes the old rules first so hosts
carrying them converge instead of stacking a second copy.
The bans themselves were also hooked into INPUT for tcp only, so they never
blocked the UDP game traffic they were meant to -- the drop rule was doing all
the damage on its own. Cleared the outstanding 32.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nextcloud follows the personal instance, which is now verified on 34.0.3:
`occ upgrade` completed, needsDbUpgrade is false, and cloud.debyl.io reports
34.0.3.2 out of maintenance mode.
BookStack skips 26.4 -- 26.3.4 is simply where this instance sat. Note the tag
is 26.5.4, not 26.05.4: solidnerd drops the month's leading zero.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Immich moves the server and machine-learning images only; the Postgres and
Redis images stay on their pinned digests, which are tied to the vectorchord
extension version rather than the app release. v3.1.0 is a quality-of-life
release -- its one breaking change is dropping iOS 14 in the mobile client,
which does not touch the server.
Nextcloud 34.0.2 -> 34.0.3 is a patch release; the official image's entrypoint
runs `occ upgrade` itself on first start with a newer version.
Skudak's Nextcloud is deliberately left on 34.0.2 until this one is verified.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The logrotate config added with the Sophie overhaul claimed to bound PZ's own
logs. It did not. data/Logs had reached 18 GB across 112 entries and 146k
files, and rotation had never once fired there.
Two reasons, both wrong assumptions on my part. PZ rolls its logs into dated
directories -- logs_2025-12-14/ through logs_2026-08-26/, up to 279 MB each --
so the Logs/*.txt glob only ever matched a handful of loose files at the top.
And it rotates by size: not one file in that tree exceeds 100M, because the
growth is in the number of files, not the size of any of them.
Age-based pruning is the right tool, so this adds a daily zomboid-log-prune
timer keeping 14 days and deleting the emptied directories behind it. It runs
under podman unshare, since the files belong to the container's UID.
logrotate keeps server-console.txt, which is a single ever-growing file and
genuinely is what it is good at. Narrowed its scope to say so rather than
implying coverage it never had.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On 42.20.3 the server loaded 279 of the preset's 283 mods and reported the
other four as "required mod ... not found". None of it was a download problem:
all 256 workshop items fetched, and every one of these four is in the preset's
own WorkshopItems= list. Only the ID each is referenced by has drifted.
Checked against the id= field of the mod.info files on disk rather than folder
names, which are unrelated to mod IDs -- an early look at folder names was
misleading here.
FWOBenchPress&Treadmill -> FWOBenchPressTreadmill (stray ampersand)
Ladders42131 -> Ladders4220 (per-build variants;
42131 is the 42.16 one)
ServingPlatesB42 -> ServingPlates42
NewMusic_OrchestraMix -> NewMusic_CMM (renamed upstream to
"Classical Music Mix")
Kept as a rename map applied at template time rather than edited into the
vendored list, so the list stays a faithful copy of upstream and re-vendoring
a newer preset keeps the corrections.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>