Commit Graph
16 Commits
Author SHA1 Message Date
Bastian de BylandClaude Opus 5 60d5ec4ae3 fix: leave world settings alone on a wipe, and snapshot them first
Reverts the restore-from-template behaviour. A wipe now resets the world and the
player database only; Server/<name>_{SandboxVars,spawnregions,spawnpoints}.lua
carry over untouched, so hand tuning survives it. Verified by checksum: all three
files byte-identical either side of a wipe.

That gives up the guarantee the restore bought. PZ writes the running world's
settings back over those files, so if a world is ever created with defaults the
file inherits them and later wipes regenerate from them -- which is how a Sophie
world became an Apocalypse one. Nothing corrects that automatically now, so the
wipe takes a snapshot of the settings before it starts, keeping the last ten
under config-backup/. Greg's edits were lost once because the only record of them
was a file the server had since overwritten; that is the hole this fills.
config-template/ still holds the pristine Sophie preset to copy back from.

The snapshot is best-effort throughout. The first version of it created the
directory with plain mkdir, the parent belongs to the container's subuid rather
than the podman user, and set -e turned that into a failed wipe -- the same shape
as the tee that broke this script before. Ansible owns the directory now and
every step of the snapshot tolerates failure, because a backup that cannot be
written is not a reason to refuse to wipe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:10:37 -04:00
Bastian de BylandClaude Opus 5 1c467b1a76 fix: stop the world wipe taking the world settings with it
A wipe was quietly turning a Sophie world into an Apocalypse one. 139 sandbox
values had reverted -- loot rates from 0.35 back to 0.9, ranged weapons and ammo
to 2.0, CharacterFreePoints from 0 to 60 -- and ten of Sophie's mod-added
options had vanished entirely.

The reset script was not deleting the settings. PZ reads
Server/<name>_SandboxVars.lua only when it creates a world, and then writes the
running world's settings back over that same file. The file is an output, not an
input. So once any world came up with defaults, the server stamped those defaults
into the file, and every wipe afterwards regenerated from them -- inheriting the
corruption rather than causing it, and with no way back out on its own. The same
shape as the admin-password deadlock.

The intended settings now live in config-template/, which the server has no
reason to touch, and a wipe restores Server/<name>_{SandboxVars,spawnregions,
spawnpoints}.lua from there before restarting. That is also the file to hand-edit
when changing the world: edit the template, wipe, and the new world has the edits.

Verified by wiping: 139 differences before, 0 after. The server still rewrites
the live file on boot -- 996 keys become 1061 as it adds newer options -- but
every Sophie value survives that now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:55:06 -04:00
Bastian de BylandClaude Opus 5 a47551724b fix: make the Discord world reset actually delete the world
The reset script died on its very first line. log() piped through tee into
zomboid/logs/world-reset.log, that directory was owned by the container's
subuid, and the script runs as the podman user -- so tee returned EACCES and
set -e killed the run before it stopped the server or touched a save. Every
`@bot` reset since had been a no-op that reported nothing.

Nothing mounts zomboid/logs into a container; it only holds output from
host-side helpers running as the podman user, so it is now owned by that user
rather than by the subuid the container volumes need.

log() no longer treats the file as load-bearing either. stdout is already
captured by the journal, so an unwritable log is worth continuing past rather
than aborting a wipe over.

Verified end to end by writing the trigger exactly as the bot does: server
stopped, saves and player database deleted, service restarted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:33:57 -04:00
Bastian de BylandClaude Opus 5 62c4015410 fix: let systemd actually notice when the Zomboid container exits
`@bot restart` saves, sends RCON quit, and relies on the unit's Restart=always
to bring the server back. It never came back. The container sat at exited(0)
while the unit reported active/running with NRestarts=0, so the bot waited for
a startup that was never going to happen and the server stayed down until
someone noticed.

podman generate systemd emits Type=forking with ExecStart=podman start and a
PIDFile pointing at conmon. podman start returns immediately, so the process
systemd was told to supervise was never its child -- it warns about exactly
this in the journal, once per poll, and then does not notice the exit:

  zomboid.service: Supervising process 2431312 which is not our child.
  We'll most likely not notice when it exits.

That PIDFile is also stale by design: it embeds the container ID, so every
deploy that recreates the container leaves it pointing at nothing.

Type=simple with `podman start -a` keeps podman in the foreground as systemd's
own child, so the exit is seen and Restart=always does what it always claimed
to. stdout and stderr are discarded deliberately -- attaching re-emits the
container's output for systemd to capture straight back into the journal, which
is the flood the k8s-file driver exists to prevent. podman logs still has it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:43:53 -04:00
Bastian de BylandClaude Opus 5 32748d1786 fix: stop the query filter banning every player who tries to connect
The anti-scanner rules dropped 53-byte Steam query packets above 1/hour with a
burst of 2, per source IP, on the theory that only scanners send queries and
that real players would already have been marked verified by the priority-2
rule when they sent something larger.

That premise is inverted. A client's first contact with the server IS a 53-byte
query, so nobody can be verified before querying, and nobody can query more
than twice an hour without being dropped. The counters said so plainly: five
packets had ever matched the verified-accept rule, against 30,563 drops.
fail2ban then banned each dropped player for a week -- 32 live bans, 550 total,
firing every ten minutes, every one of them a residential address.

Keeps the shape of the protection and moves the threshold somewhere no real
client reaches: opening the server browser or retrying a connection is a
handful of queries, a flood is thousands. Removes the old rules first so hosts
carrying them converge instead of stacking a second copy.

The bans themselves were also hooked into INPUT for tcp only, so they never
blocked the UDP game traffic they were meant to -- the drop rule was doing all
the damage on its own. Cleared the outstanding 32.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:26:45 -04:00
Bastian de BylandClaude Opus 5 91d9f6b419 fix: actually bound the Zomboid log archive, which logrotate could not
The logrotate config added with the Sophie overhaul claimed to bound PZ's own
logs. It did not. data/Logs had reached 18 GB across 112 entries and 146k
files, and rotation had never once fired there.

Two reasons, both wrong assumptions on my part. PZ rolls its logs into dated
directories -- logs_2025-12-14/ through logs_2026-08-26/, up to 279 MB each --
so the Logs/*.txt glob only ever matched a handful of loose files at the top.
And it rotates by size: not one file in that tree exceeds 100M, because the
growth is in the number of files, not the size of any of them.

Age-based pruning is the right tool, so this adds a daily zomboid-log-prune
timer keeping 14 days and deleting the emptied directories behind it. It runs
under podman unshare, since the files belong to the container's UID.

logrotate keeps server-console.txt, which is a single ever-growing file and
genuinely is what it is good at. Narrowed its scope to say so rather than
implying coverage it never had.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 02:34:30 -04:00
Bastian de BylandClaude Opus 5 d34c6cd36d rebuild the Zomboid server on the Sophie 42 preset as "debbzoid"
The deployment had drifted badly. It was pinned to the Steam `-beta unstable`
branch, which stopped being right the moment B42 became the default at 42.20,
and it wrote every mod ID with the `\` prefix that B42 required only during
that unstable period and now rejects. Three server profiles were carried
around (vanilla, modded, b42revamp) whose mod lists were hand-curated blobs
from a Steam collection that has since moved on.

Moves to the Sophie 42 community preset, vendored from GerDeathstar/sophie-pz
("Sophie 42 Files.zip", 2026-07-31 == the b42 release tag): 283 mod IDs, 256
workshop items, map_distanciado over Muldraugh, plus its SandboxVars,
spawnregions and spawnpoints. Sophie's gameplay settings are kept exactly as
shipped -- PVP with its damage modifiers, the safety system, MaxPlayers=32,
PauseEmpty, no sleep, safehouses off. Only the keys this deployment actually
owns are templated over the top.

The modlist is stored as YAML lists rather than the INI's semicolon blobs, so
the next Sophie update produces a diff you can read. Exclusions live in
zomboid_mods_excluded with a reason each, seeded with IconsInventory -- already
absent upstream, listed so a re-vendor cannot quietly bring it back.

Config is now seeded before first boot instead of patched after it. The old
approach could not write the INI until the server had generated one, so every
setting went through lineinfile guarded on a stat; seeding the whole file from
the preset removes the chicken-and-egg and puts the config in git. force is
off by design: these files are a starting point, not managed state, so the
server can be stopped, hand-edited and regenerated without Ansible clobbering
the edits. Push them again deliberately with -e zomboid_config_force=true.

The B41 Discord keys were dead. B42 replaced DiscordChannel/DiscordChannelID
with DiscordChatChannel, which takes a channel name rather than a snowflake,
so the DiscordChannelID=... this role had been appending was a key the server
ignores and the chat bridge has not been working.

Renaming the server to debbzoid is what starts the new world: PZ derives the
INI, save directory and player DB from the name, so gregboid stays on disk
untouched as a rollback.

Re-enabled, because the reason it was off is now fixed. It was disabled for
saturating the SSD -- ~10 MB/s of log writes into the shared journal, which
starved Gitea CI badly enough to stretch a firmware build to 17 minutes. The
container now logs to its own rotating k8s-file instead of the journal every
other service shares, and logrotate caps the server's own server-console.txt
and Logs/ at 100M keeping 3. copytruncate is mandatory there: PZ holds those
fds for its whole life, so a rename would leave it writing into an unlinked
inode. Memory is unchanged; idle draw is nowhere near MAX_RAM.

Also stops the entrypoint recursively chowning the install tree on every boot.
With 256 workshop mods that is a large inode walk and pure churn once the
first run has set ownership.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:41:00 -04:00
Bastian de BylandClaude Opus 4.5 9d562c7188 feat: smart zomboid traffic filtering with packet-size detection
Replace per-IP hashlimit with smarter filtering that distinguishes
legitimate players from scanner bots based on packet behavior:
- Players send varied packet sizes (53, 37, 1472 bytes)
- Scanners only send 53-byte query packets

New firewall rule chain:
- Priority 2: Mark + ACCEPT non-query packets (verifies player)
- Priority 3: ACCEPT queries from verified IPs (1 hour TTL)
- Priority 4: LOG rate-limited queries from unverified IPs
- Priority 5: DROP rate-limited queries (2 burst, then 1/hour)

Also includes:
- Fail2ban zomboid jail with tighter thresholds (5 retries/4h, 1w ban)
- Graylog streams for zomboid-connections, zomboid-ratelimit, fail2ban
- GeoIP pipeline enrichment for zomboid traffic
- Fluent-bit inputs for ratelimit logs and fail2ban events
- Remove Legendary Katana mod (Workshop 3418366499) - removed from Steam
- Bump Immich to v2.5.0
- Fix fulfillr config (nil → null)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-27 15:09:26 -05:00
Bastian de Byl bc26fcd1f9 chore: fluent-bit zomboid, zomboid stats, home assistant, gregbot 2026-01-24 17:08:05 -05:00
Bastian de Byl 9a95eecfd5 chore: zomboid stats for gregtime, updates 2026-01-23 12:02:57 -05:00
Bastian de BylandClaude Opus 4.5 c2d117bd95 feat: add systemd timer for zomboid container stats
Deploy systemd timer that writes zomboid container stats to
zomboid-stats.json every 30 seconds for gregtime to read.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-22 23:10:05 -05:00
Bastian de Byl 9e665a841d chore: non-cifs nextcloud, partsy, zomboid updates 2026-01-15 16:48:07 -05:00
Bastian de Byl 4d835e86a0 chore: zomboid improvements, gregtime improvements with rcon 2025-12-22 12:31:43 -05:00
Bastian de Byl 38561cb968 gitea, zomboid updates, ssh key fixes 2025-12-19 10:39:56 -05:00
Bastian de Byl adce3e2dd4 chore: zomboid improvements, immich and other updates 2025-12-14 22:07:49 -05:00
Bastian de Byl 8c21923358 zomboid added, caddyfile updates, debylio migration, ddns migration 2025-12-13 21:18:33 -05:00