Rolling a world back meant hand-work over ssh: stop the service, move the live
save aside, unzip the right archive, chown into the container's subuid range,
relabel, start. That is the wrong shape of task to do by hand, and it is always
done under time pressure -- by construction, because the archive you want is
being deleted while you work.
PZ keeps BackupsCount=10 per set and writes one every BackupsPeriod=30 minutes,
so a periodic backup is reachable for about five hours and then gone. On
2026-09-05 the snapshot the admins asked for (05:16, four minutes before the
incident) had about 90 minutes of life left when the request came in.
Same shape as the wipe: the Discord bot writes a trigger file into its own rw
volume, zomboid-restore.path notices it, and zomboid-restore.service runs the
script as the podman user. The bot gets no ssh, no systemd, and keeps only its
existing read-only mount of the Zomboid volume.
Two details carry most of the correctness.
Resolution is by mtime, not by index. The rotation renames the files -- today's
backup_7.zip is backup_8.zip half an hour from now, and a new backup_7.zip holds
a different world -- so an index is valid only while the listing is fresh, which
is not long enough to survive a human reading a confirmation prompt. The trigger
names a set and an mtime; the script resolves the path itself, whitelists the
filename, and refuses if nothing matches. It never accepts a path.
Everything that can fail is checked before the server is touched. A rotated-out
target, an archive with no debbzoid world in it, a bad action, a traversal
attempt in the set name: each aborts with the server still running and writes a
result file the bot reports back. The live world is moved aside rather than
deleted, so a restore is undoable and the last three are kept.
One thing PZ does not advertise: its backups do not cover the whole save
directory. blam/, a mod's own state, is in none of them -- not the 05:16 archive
and not the newest one. Restoring only what the archive holds therefore lands
the world slightly *behind* the target rather than on it, so anything present in
the displaced world and absent from the archive is carried across.
The gregtime tag moves to 3.17.0 for the bot half of this -- `backups`,
`restore <n>`, `restore confirm`, `restore undo`, gated to the same two admins
as the wipe. That image is built and running on the host; its source is not
committed yet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
The vendored Sophie preset and the running debbzoid world had drifted, and the
repo only held the preset. That is not a restore point: force-pushing it would
have reverted the admins' in-game tuning rather than recovering it, which is
exactly how a Sophie world quietly became an Apocalypse one once already -- 139
values reverted, loot from 0.35 back to 0.9, CharacterFreePoints 0 to 60.
files/zomboid/sophie/SandboxVars.lua is now a snapshot of the live world taken
2026-08-31, not the preset as shipped. The modlist is untouched and still
upstream, which is why zomboid_preset_version now names the two halves and their
separate dates. server.ini.j2 carries the eight keys that had drifted:
PlayerSafehouse false -> true
SafehouseAllowNonResidential false -> true (the diner/gas-station case)
SafehouseAllowRespawn false -> true
SafehouseAllowLoot true -> false
SafehouseAllowFire true -> false
TrashDeleteAll false -> true
MapRemotePlayerVisibility 1 -> 4
ResetID 6953472 -> 826046
ResetID is in that list on purpose, and matters most. It is the world's
soft-reset token: a file value that differs from the one the live world was
created with tells every connected client to roll a new character. Carrying the
live value makes a deliberate force-push a no-op instead of a server-wide wipe
prompt.
Spawn config gets its own switch, zomboid_spawn_force. spawnregions.lua and
spawnpoints.lua are the only config a running world re-reads -- at every server
start, where SandboxVars is read once, when the world is created -- so a spawn
edit is deployable on the live world without a wipe. Sharing zomboid_config_force
between them would have meant force-pushing the whole preset to land a one-line
spawn edit, rewriting the INI (hence the ResetID hazard above) and the world's
SandboxVars along with it.
config-template/ now tracks the repo unconditionally. Nothing on the host writes
that directory and the server cannot see it, so it has no hand edits to protect;
if it does not track the repo it is not a restore point, just an older world's
settings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
Reverts the restore-from-template behaviour. A wipe now resets the world and the
player database only; Server/<name>_{SandboxVars,spawnregions,spawnpoints}.lua
carry over untouched, so hand tuning survives it. Verified by checksum: all three
files byte-identical either side of a wipe.
That gives up the guarantee the restore bought. PZ writes the running world's
settings back over those files, so if a world is ever created with defaults the
file inherits them and later wipes regenerate from them -- which is how a Sophie
world became an Apocalypse one. Nothing corrects that automatically now, so the
wipe takes a snapshot of the settings before it starts, keeping the last ten
under config-backup/. Greg's edits were lost once because the only record of them
was a file the server had since overwritten; that is the hole this fills.
config-template/ still holds the pristine Sophie preset to copy back from.
The snapshot is best-effort throughout. The first version of it created the
directory with plain mkdir, the parent belongs to the container's subuid rather
than the podman user, and set -e turned that into a failed wipe -- the same shape
as the tee that broke this script before. Ansible owns the directory now and
every step of the snapshot tolerates failure, because a backup that cannot be
written is not a reason to refuse to wipe.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A wipe was quietly turning a Sophie world into an Apocalypse one. 139 sandbox
values had reverted -- loot rates from 0.35 back to 0.9, ranged weapons and ammo
to 2.0, CharacterFreePoints from 0 to 60 -- and ten of Sophie's mod-added
options had vanished entirely.
The reset script was not deleting the settings. PZ reads
Server/<name>_SandboxVars.lua only when it creates a world, and then writes the
running world's settings back over that same file. The file is an output, not an
input. So once any world came up with defaults, the server stamped those defaults
into the file, and every wipe afterwards regenerated from them -- inheriting the
corruption rather than causing it, and with no way back out on its own. The same
shape as the admin-password deadlock.
The intended settings now live in config-template/, which the server has no
reason to touch, and a wipe restores Server/<name>_{SandboxVars,spawnregions,
spawnpoints}.lua from there before restarting. That is also the file to hand-edit
when changing the world: edit the template, wipe, and the new world has the edits.
Verified by wiping: 139 differences before, 0 after. The server still rewrites
the live file on boot -- 996 keys become 1061 as it adds newer options -- but
every Sophie value survives that now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reset script died on its very first line. log() piped through tee into
zomboid/logs/world-reset.log, that directory was owned by the container's
subuid, and the script runs as the podman user -- so tee returned EACCES and
set -e killed the run before it stopped the server or touched a save. Every
`@bot` reset since had been a no-op that reported nothing.
Nothing mounts zomboid/logs into a container; it only holds output from
host-side helpers running as the podman user, so it is now owned by that user
rather than by the subuid the container volumes need.
log() no longer treats the file as load-bearing either. stdout is already
captured by the journal, so an unwritable log is worth continuing past rather
than aborting a wipe over.
Verified end to end by writing the trigger exactly as the bot does: server
stopped, saves and player database deleted, service restarted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The dedicated server (app 380870) and the client are packaged separately and
have drifted. Base.Log_Stack_01 was removed from the client in 42.16, but the
server package still ships entity_logstack.txt.
The server registers every script it can see into a new world's
WorldDictionary, so the entry lands in the save, and every client then fails at
world load with:
WorldDictionaryException: [SpriteConfigs] Missing dictionary script on
client: Base.Log_Stack_01
Nothing in multiplayer can legitimately reference a script that no client has,
so dropping it server-side costs nothing.
The strip runs after the SteamCMD update rather than once at install time on
purpose: validate restores the file on every boot, so it has to be removed on
every boot too.
Note this does not repair a world whose dictionary already records the entry --
that save still needs regenerating.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The logrotate config added with the Sophie overhaul claimed to bound PZ's own
logs. It did not. data/Logs had reached 18 GB across 112 entries and 146k
files, and rotation had never once fired there.
Two reasons, both wrong assumptions on my part. PZ rolls its logs into dated
directories -- logs_2025-12-14/ through logs_2026-08-26/, up to 279 MB each --
so the Logs/*.txt glob only ever matched a handful of loose files at the top.
And it rotates by size: not one file in that tree exceeds 100M, because the
growth is in the number of files, not the size of any of them.
Age-based pruning is the right tool, so this adds a daily zomboid-log-prune
timer keeping 14 days and deleting the emptied directories behind it. It runs
under podman unshare, since the files belong to the container's UID.
logrotate keeps server-console.txt, which is a single ever-growing file and
genuinely is what it is good at. Narrowed its scope to say so rather than
implying coverage it never had.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On 42.20.3 the server loaded 279 of the preset's 283 mods and reported the
other four as "required mod ... not found". None of it was a download problem:
all 256 workshop items fetched, and every one of these four is in the preset's
own WorkshopItems= list. Only the ID each is referenced by has drifted.
Checked against the id= field of the mod.info files on disk rather than folder
names, which are unrelated to mod IDs -- an early look at folder names was
misleading here.
FWOBenchPress&Treadmill -> FWOBenchPressTreadmill (stray ampersand)
Ladders42131 -> Ladders4220 (per-build variants;
42131 is the 42.16 one)
ServingPlatesB42 -> ServingPlates42
NewMusic_OrchestraMix -> NewMusic_CMM (renamed upstream to
"Classical Music Mix")
Kept as a rename map applied at template time rather than edited into the
vendored list, so the list stays a faithful copy of upstream and re-vendoring
a newer preset keeps the corrections.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The flag was gated on the admin database not existing yet, treating the file's
presence as proof the admin user had been created. It is not: PZ creates the
DB file in ServerWorldDatabase.create() before it creates the admin user, so a
boot interrupted between the two leaves a database with no admin in it.
After that the guard withheld -adminpassword on every subsequent start, the
server fell back to an interactive "Enter new administrator password:" prompt,
read EOF because a detached container has no stdin, and died with
NoSuchElementException. systemd restarted it into exactly the same state, 50
times, which is how this was found.
Passing it unconditionally costs nothing when the admin already exists and
removes the deadlock entirely.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dropping -beta from app_update does not leave a beta branch. SteamCMD wrote
"public" into the manifest's UserConfig and left MountedConfig on "unstable",
so the install stayed on the 42.14.1 build from February while reporting
success. The config had already moved to the bare mod-ID syntax that 42.20+
wants, and 42.14.1 still requires the backslash prefix, so all 283 mods
loaded as "required mod ... not found" and the server came up vanilla.
Naming the branch explicitly is what actually remounts the content.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
logrotate warns that size overrides daily, and it is right: with both set the
time directive does nothing. The daily timer is when the config gets checked;
100M is when it rotates. Saying so in a comment rather than in a directive
that has no effect.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The deployment had drifted badly. It was pinned to the Steam `-beta unstable`
branch, which stopped being right the moment B42 became the default at 42.20,
and it wrote every mod ID with the `\` prefix that B42 required only during
that unstable period and now rejects. Three server profiles were carried
around (vanilla, modded, b42revamp) whose mod lists were hand-curated blobs
from a Steam collection that has since moved on.
Moves to the Sophie 42 community preset, vendored from GerDeathstar/sophie-pz
("Sophie 42 Files.zip", 2026-07-31 == the b42 release tag): 283 mod IDs, 256
workshop items, map_distanciado over Muldraugh, plus its SandboxVars,
spawnregions and spawnpoints. Sophie's gameplay settings are kept exactly as
shipped -- PVP with its damage modifiers, the safety system, MaxPlayers=32,
PauseEmpty, no sleep, safehouses off. Only the keys this deployment actually
owns are templated over the top.
The modlist is stored as YAML lists rather than the INI's semicolon blobs, so
the next Sophie update produces a diff you can read. Exclusions live in
zomboid_mods_excluded with a reason each, seeded with IconsInventory -- already
absent upstream, listed so a re-vendor cannot quietly bring it back.
Config is now seeded before first boot instead of patched after it. The old
approach could not write the INI until the server had generated one, so every
setting went through lineinfile guarded on a stat; seeding the whole file from
the preset removes the chicken-and-egg and puts the config in git. force is
off by design: these files are a starting point, not managed state, so the
server can be stopped, hand-edited and regenerated without Ansible clobbering
the edits. Push them again deliberately with -e zomboid_config_force=true.
The B41 Discord keys were dead. B42 replaced DiscordChannel/DiscordChannelID
with DiscordChatChannel, which takes a channel name rather than a snowflake,
so the DiscordChannelID=... this role had been appending was a key the server
ignores and the chat bridge has not been working.
Renaming the server to debbzoid is what starts the new world: PZ derives the
INI, save directory and player DB from the name, so gregboid stays on disk
untouched as a rollback.
Re-enabled, because the reason it was off is now fixed. It was disabled for
saturating the SSD -- ~10 MB/s of log writes into the shared journal, which
starved Gitea CI badly enough to stretch a firmware build to 17 minutes. The
container now logs to its own rotating k8s-file instead of the journal every
other service shares, and logrotate caps the server's own server-console.txt
and Logs/ at 100M keeping 3. copytruncate is mandatory there: PZ holds those
fds for its whole life, so a rename would leave it writing into an unlinked
inode. Memory is unchanged; idle draw is nowhere near MAX_RAM.
Also stops the entrypoint recursively chowning the install tree on every boot.
With 256 workshop mods that is a large inode walk and pure churn once the
first run has set ownership.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Add ollama role for local LLM inference (install, service, models)
- Add searxng container for private search
- Migrate hostname from home.bdebyl.net to home.debyl.io
(inventory, awsddns, zomboid entrypoint, home_server_name)
- Update vault with new secrets
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Deploy systemd path unit that watches for trigger file from Discord
bot and executes world reset script to delete saves and restart server.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Deploy systemd timer that writes zomboid container stats to
zomboid-stats.json every 30 seconds for gregtime to read.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>