The sunset automation required TV mode off, so an afternoon of TV skipped
the driveway string lights entirely (Sep 22 and 23) - not an outage. The
driveway now has its own sunset -1h automation, with catch-up on restart
or switch reconnect before 23:00.
Evening brightness lives in one script (evening_lights_apply) that blends
between the old step levels; a 5-minute ramp from 20:30 eases lights that
are on and leaves alone any a person has changed by hand. The Dining Hall
no longer bumps to 50% at 21:30.
TV off after 23:30 now only turns off the living room glow instead of
bringing the whole house back to full brightness, and TV on/off leave the
lights alone in daylight. Lights-out moves to 23:30, and a 01:00 sweep
catches anything switched back on at the wall.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A single Go binary with SQLite, built and loaded as localhost/rsvpd:<VERSION>
by make deploy-remote in ~/src/rsvp-debylio. Both of its listeners are
published on 127.0.0.1 only: Caddy proxies the public one to everyone and the
admin one (/admin, no login) only to caddy_local_networks. A Caddyfile mistake
alone cannot expose admin, and neither can a port mistake alone.
Guests' invite links are the credential and they sit in the URL path, which
shapes the vhost:
- It does not import common_headers. That snippet sets Referrer-Policy
same-origin, which would replace the app's no-referrer and let a token leak
in a Referer header.
- Its access log rewrites request>uri to /i/REDACTED and drops the Location
response header, since every POST 303s back to /i/<token>.
- Caddy's error logger is separate from the site's and wrote the raw URI to
caddy.log when the upstream was down. The global log now excludes
http.log.error.rsvp and a filtered rsvp-errors logger takes it instead.
Verified with zero token occurrences in both logs, locally and live.
The data directory is owned directly by the host uid of the container's uid
10001 (subuid + 10000). Setting it to the podman user and chowning back each run
flipped ownership on every deploy and briefly locked the app out of its
database.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rolling a world back meant hand-work over ssh: stop the service, move the live
save aside, unzip the right archive, chown into the container's subuid range,
relabel, start. That is the wrong shape of task to do by hand, and it is always
done under time pressure -- by construction, because the archive you want is
being deleted while you work.
PZ keeps BackupsCount=10 per set and writes one every BackupsPeriod=30 minutes,
so a periodic backup is reachable for about five hours and then gone. On
2026-09-05 the snapshot the admins asked for (05:16, four minutes before the
incident) had about 90 minutes of life left when the request came in.
Same shape as the wipe: the Discord bot writes a trigger file into its own rw
volume, zomboid-restore.path notices it, and zomboid-restore.service runs the
script as the podman user. The bot gets no ssh, no systemd, and keeps only its
existing read-only mount of the Zomboid volume.
Two details carry most of the correctness.
Resolution is by mtime, not by index. The rotation renames the files -- today's
backup_7.zip is backup_8.zip half an hour from now, and a new backup_7.zip holds
a different world -- so an index is valid only while the listing is fresh, which
is not long enough to survive a human reading a confirmation prompt. The trigger
names a set and an mtime; the script resolves the path itself, whitelists the
filename, and refuses if nothing matches. It never accepts a path.
Everything that can fail is checked before the server is touched. A rotated-out
target, an archive with no debbzoid world in it, a bad action, a traversal
attempt in the set name: each aborts with the server still running and writes a
result file the bot reports back. The live world is moved aside rather than
deleted, so a restore is undoable and the last three are kept.
One thing PZ does not advertise: its backups do not cover the whole save
directory. blam/, a mod's own state, is in none of them -- not the 05:16 archive
and not the newest one. Restoring only what the archive holds therefore lands
the world slightly *behind* the target rather than on it, so anything present in
the displaced world and absent from the archive is carried across.
The gregtime tag moves to 3.17.0 for the bot half of this -- `backups`,
`restore <n>`, `restore confirm`, `restore undo`, gated to the same two admins
as the wipe. That image is built and running on the host; its source is not
committed yet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
The vendored Sophie preset and the running debbzoid world had drifted, and the
repo only held the preset. That is not a restore point: force-pushing it would
have reverted the admins' in-game tuning rather than recovering it, which is
exactly how a Sophie world quietly became an Apocalypse one once already -- 139
values reverted, loot from 0.35 back to 0.9, CharacterFreePoints 0 to 60.
files/zomboid/sophie/SandboxVars.lua is now a snapshot of the live world taken
2026-08-31, not the preset as shipped. The modlist is untouched and still
upstream, which is why zomboid_preset_version now names the two halves and their
separate dates. server.ini.j2 carries the eight keys that had drifted:
PlayerSafehouse false -> true
SafehouseAllowNonResidential false -> true (the diner/gas-station case)
SafehouseAllowRespawn false -> true
SafehouseAllowLoot true -> false
SafehouseAllowFire true -> false
TrashDeleteAll false -> true
MapRemotePlayerVisibility 1 -> 4
ResetID 6953472 -> 826046
ResetID is in that list on purpose, and matters most. It is the world's
soft-reset token: a file value that differs from the one the live world was
created with tells every connected client to roll a new character. Carrying the
live value makes a deliberate force-push a no-op instead of a server-wide wipe
prompt.
Spawn config gets its own switch, zomboid_spawn_force. spawnregions.lua and
spawnpoints.lua are the only config a running world re-reads -- at every server
start, where SandboxVars is read once, when the world is created -- so a spawn
edit is deployable on the live world without a wipe. Sharing zomboid_config_force
between them would have meant force-pushing the whole preset to land a one-line
spawn edit, rewriting the INI (hence the ResetID hazard above) and the world's
SandboxVars along with it.
config-template/ now tracks the repo unconditionally. Nothing on the host writes
that directory and the server cannot see it, so it has no hand edits to protect;
if it does not track the repo it is not a restore point, just an older world's
settings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
Smart search, face detection and OCR were all failing with
Machine learning request to "http://immich-machine-learning:3003" failed
while the container still reported Up. Podman only sees PID 1: gunicorn's
master was alive and holding the listening socket, but its worker had died at a
WORKER TIMEOUT and was never respawned -- hence a connect timeout rather than a
refusal. The dead worker was a zombie whose remaining thread was stuck in
uninterruptible sleep in exit_mmap, so it survived SIGKILL, podman rm -f and
rm -f -t 0, and kept the container name and network alias until the host was
rebooted.
MACHINE_LEARNING_MODEL_TTL=0 addresses the cause rather than the symptom. The
default unloads models after 300s idle, so every search following a gap
reloaded four of them (CLIP, buffalo_l detection + recognition, PP-OCRv5) on a
CPU-only 4-core box and then tore those mappings back down -- and that teardown
is what wedged. Keeping them resident costs ~1-2 GB and removes the path.
IMMICH_MACHINE_LEARNING_URL is now set explicitly instead of relying on
immich's implicit default, so the ML container can be renamed without silently
losing search, and the name is a variable so a replacement can be stood up
beside a broken one without editing tasks.
Verified after deploy: zero ML failures, "in-memory cache with unloading
disabled", ping 200 from immich-server, 26/26 containers healthy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reverts the restore-from-template behaviour. A wipe now resets the world and the
player database only; Server/<name>_{SandboxVars,spawnregions,spawnpoints}.lua
carry over untouched, so hand tuning survives it. Verified by checksum: all three
files byte-identical either side of a wipe.
That gives up the guarantee the restore bought. PZ writes the running world's
settings back over those files, so if a world is ever created with defaults the
file inherits them and later wipes regenerate from them -- which is how a Sophie
world became an Apocalypse one. Nothing corrects that automatically now, so the
wipe takes a snapshot of the settings before it starts, keeping the last ten
under config-backup/. Greg's edits were lost once because the only record of them
was a file the server had since overwritten; that is the hole this fills.
config-template/ still holds the pristine Sophie preset to copy back from.
The snapshot is best-effort throughout. The first version of it created the
directory with plain mkdir, the parent belongs to the container's subuid rather
than the podman user, and set -e turned that into a failed wipe -- the same shape
as the tee that broke this script before. Ansible owns the directory now and
every step of the snapshot tolerates failure, because a backup that cannot be
written is not a reason to refuse to wipe.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A wipe was quietly turning a Sophie world into an Apocalypse one. 139 sandbox
values had reverted -- loot rates from 0.35 back to 0.9, ranged weapons and ammo
to 2.0, CharacterFreePoints from 0 to 60 -- and ten of Sophie's mod-added
options had vanished entirely.
The reset script was not deleting the settings. PZ reads
Server/<name>_SandboxVars.lua only when it creates a world, and then writes the
running world's settings back over that same file. The file is an output, not an
input. So once any world came up with defaults, the server stamped those defaults
into the file, and every wipe afterwards regenerated from them -- inheriting the
corruption rather than causing it, and with no way back out on its own. The same
shape as the admin-password deadlock.
The intended settings now live in config-template/, which the server has no
reason to touch, and a wipe restores Server/<name>_{SandboxVars,spawnregions,
spawnpoints}.lua from there before restarting. That is also the file to hand-edit
when changing the world: edit the template, wipe, and the new world has the edits.
Verified by wiping: 139 differences before, 0 after. The server still rewrites
the live file on boot -- 996 keys become 1061 as it adds newer options -- but
every Sophie value survives that now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reset script died on its very first line. log() piped through tee into
zomboid/logs/world-reset.log, that directory was owned by the container's
subuid, and the script runs as the podman user -- so tee returned EACCES and
set -e killed the run before it stopped the server or touched a save. Every
`@bot` reset since had been a no-op that reported nothing.
Nothing mounts zomboid/logs into a container; it only holds output from
host-side helpers running as the podman user, so it is now owned by that user
rather than by the subuid the container volumes need.
log() no longer treats the file as load-bearing either. stdout is already
captured by the journal, so an unwritable log is worth continuing past rather
than aborting a wipe over.
Verified end to end by writing the trigger exactly as the bot does: server
stopped, saves and player database deleted, service restarted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`@bot restart` saves, sends RCON quit, and relies on the unit's Restart=always
to bring the server back. It never came back. The container sat at exited(0)
while the unit reported active/running with NRestarts=0, so the bot waited for
a startup that was never going to happen and the server stayed down until
someone noticed.
podman generate systemd emits Type=forking with ExecStart=podman start and a
PIDFile pointing at conmon. podman start returns immediately, so the process
systemd was told to supervise was never its child -- it warns about exactly
this in the journal, once per poll, and then does not notice the exit:
zomboid.service: Supervising process 2431312 which is not our child.
We'll most likely not notice when it exits.
That PIDFile is also stale by design: it embeds the container ID, so every
deploy that recreates the container leaves it pointing at nothing.
Type=simple with `podman start -a` keeps podman in the foreground as systemd's
own child, so the exit is seen and Restart=always does what it always claimed
to. stdout and stderr are discarded deliberately -- attaching re-emits the
container's output for systemd to capture straight back into the journal, which
is the flood the k8s-file driver exists to prevent. podman logs still has it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The anti-scanner rules dropped 53-byte Steam query packets above 1/hour with a
burst of 2, per source IP, on the theory that only scanners send queries and
that real players would already have been marked verified by the priority-2
rule when they sent something larger.
That premise is inverted. A client's first contact with the server IS a 53-byte
query, so nobody can be verified before querying, and nobody can query more
than twice an hour without being dropped. The counters said so plainly: five
packets had ever matched the verified-accept rule, against 30,563 drops.
fail2ban then banned each dropped player for a week -- 32 live bans, 550 total,
firing every ten minutes, every one of them a residential address.
Keeps the shape of the protection and moves the threshold somewhere no real
client reaches: opening the server browser or retrying a connection is a
handful of queries, a flood is thousands. Removes the old rules first so hosts
carrying them converge instead of stacking a second copy.
The bans themselves were also hooked into INPUT for tcp only, so they never
blocked the UDP game traffic they were meant to -- the drop rule was doing all
the damage on its own. Cleared the outstanding 32.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The logrotate config added with the Sophie overhaul claimed to bound PZ's own
logs. It did not. data/Logs had reached 18 GB across 112 entries and 146k
files, and rotation had never once fired there.
Two reasons, both wrong assumptions on my part. PZ rolls its logs into dated
directories -- logs_2025-12-14/ through logs_2026-08-26/, up to 279 MB each --
so the Logs/*.txt glob only ever matched a handful of loose files at the top.
And it rotates by size: not one file in that tree exceeds 100M, because the
growth is in the number of files, not the size of any of them.
Age-based pruning is the right tool, so this adds a daily zomboid-log-prune
timer keeping 14 days and deleting the emptied directories behind it. It runs
under podman unshare, since the files belong to the container's UID.
logrotate keeps server-console.txt, which is a single ever-growing file and
genuinely is what it is good at. Narrowed its scope to say so rather than
implying coverage it never had.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The deployment had drifted badly. It was pinned to the Steam `-beta unstable`
branch, which stopped being right the moment B42 became the default at 42.20,
and it wrote every mod ID with the `\` prefix that B42 required only during
that unstable period and now rejects. Three server profiles were carried
around (vanilla, modded, b42revamp) whose mod lists were hand-curated blobs
from a Steam collection that has since moved on.
Moves to the Sophie 42 community preset, vendored from GerDeathstar/sophie-pz
("Sophie 42 Files.zip", 2026-07-31 == the b42 release tag): 283 mod IDs, 256
workshop items, map_distanciado over Muldraugh, plus its SandboxVars,
spawnregions and spawnpoints. Sophie's gameplay settings are kept exactly as
shipped -- PVP with its damage modifiers, the safety system, MaxPlayers=32,
PauseEmpty, no sleep, safehouses off. Only the keys this deployment actually
owns are templated over the top.
The modlist is stored as YAML lists rather than the INI's semicolon blobs, so
the next Sophie update produces a diff you can read. Exclusions live in
zomboid_mods_excluded with a reason each, seeded with IconsInventory -- already
absent upstream, listed so a re-vendor cannot quietly bring it back.
Config is now seeded before first boot instead of patched after it. The old
approach could not write the INI until the server had generated one, so every
setting went through lineinfile guarded on a stat; seeding the whole file from
the preset removes the chicken-and-egg and puts the config in git. force is
off by design: these files are a starting point, not managed state, so the
server can be stopped, hand-edited and regenerated without Ansible clobbering
the edits. Push them again deliberately with -e zomboid_config_force=true.
The B41 Discord keys were dead. B42 replaced DiscordChannel/DiscordChannelID
with DiscordChatChannel, which takes a channel name rather than a snowflake,
so the DiscordChannelID=... this role had been appending was a key the server
ignores and the chat bridge has not been working.
Renaming the server to debbzoid is what starts the new world: PZ derives the
INI, save directory and player DB from the name, so gregboid stays on disk
untouched as a rollback.
Re-enabled, because the reason it was off is now fixed. It was disabled for
saturating the SSD -- ~10 MB/s of log writes into the shared journal, which
starved Gitea CI badly enough to stretch a firmware build to 17 minutes. The
container now logs to its own rotating k8s-file instead of the journal every
other service shares, and logrotate caps the server's own server-console.txt
and Logs/ at 100M keeping 3. copytruncate is mandatory there: PZ holds those
fds for its whole life, so a rename would leave it writing into an unlinked
inode. Memory is unchanged; idle draw is nowhere near MAX_RAM.
Also stops the entrypoint recursively chowning the install tree on every boot.
With 256 workshop mods that is a large inode walk and pure churn once the
first run has set ownership.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TrueNAS was power-cycled, the CIFS mounts failed, and systemd never retried
-- mount units are not restarted on failure. SMB came back, nothing
remounted, and immich-server served an empty library for days while its
database still listed 14,181 assets pointing at /mnt/media/originals.
fstab carried no _netdev, no nofail and no automount, so there was no path
back without a human. Now:
x-systemd.automount any access re-attempts the mount; failure stops
being terminal
_netdev / nofail ordered after network-online, dead NAS cannot block boot
soft I/O errors instead of blocking forever, so the
container can be restarted rather than wedging in
uninterruptible sleep
idle-timeout unmount when unused, clearing stale handles
resilienthandles SMB3 rides out brief blips
ansible.posix.mount mounts directly and never starts the generated
.automount unit, leaving the on-access trigger inactive -- enable it
explicitly, or the headline fix silently does nothing.
The containers are systemd USER units while the mounts are SYSTEM units, so
RequiresMountsFor= is unavailable. cifs-watchdog bridges the scopes: checks
health, recovers, and restarts ONLY immich-server (the sole consumer of both
paths; postgres/redis/ML use local volumes).
Two bugs the umount test caught, both worth knowing:
- `ls` cannot test mountedness. An unmounted mount point is an ordinary
empty directory, so ls succeeds and recovery was skipped entirely.
- A drop repaired within a single run leaves prev=healthy, so keying the
restart solely on the stored state skipped it while the container still
held its stale view.
Also moves the SMB password out of /etc/fstab, which is 0644 and was
readable by every local user, into a 0600 credentials file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Extends the Nextcloud backup machinery rather than adding a second
mechanism. cloud-backup.sh.j2 gains three guarded options, all no-ops for
the existing callers:
backup_podman_user Gitea runs rootless under `git`, not `podman`
backup_db_type postgres (Gitea) and mysql (BookStack) alongside
mariadb; each engine's completion trailer differs,
and grepping for the wrong one fails every run
backup_sqlite_dbs `sqlite3 .backup` for live WAL-mode SQLite, gated on
`pragma integrity_check` before promotion -- rsync
is either stale (no -wal) or torn (with it)
New instances: gitea-debyl, skudak-gitea, bookstack, partsy-skudak. The
alert handler is rendered once and shared, so its wording is now generic
rather than per-product; TAG stays nextcloud-backup because an external
Graylog rule matches on it.
`apply:` on the includes is load-bearing -- tags on a dynamic
include_tasks do not reach the tasks inside it.
Business data (skudak-gitea, bookstack, partsy-skudak) goes to TrueNAS
and on to Skudak's own iDrive account; the personal bucket's
/skudak*/** excludes are permanent, not a stopgap.
Removals: PartKeepr is superseded by Partsy, and its teardown never
finished -- it targeted /etc/systemd/system/podman-partkeepr*.service,
wrong prefix and wrong scope, leaving enabled user units in failed state.
Pi-hole's role was already orphaned (absent from deploy_home.yml); its
port 53 rule went with it after confirming nothing listens there.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ups role: NUT server on home.debyl.io for the CyberPower PR1500RT2U that backs
both it and truenas.localdomain, with staged shutdown (TrueNAS sheds at t+2min,
host at 10% charge) and best-effort IPMI power-on when mains returns. The 10%
threshold leans on ignorelb + override.battery.charge.low rather than a custom
poller, because CyberPower asserts its own low-battery flag far too early.
Credentials come from vault vars; nothing sensitive is templated in the clear.
Nextcloud background jobs: both instances have backgroundjobs_mode "cron", which
expects an external caller every ~5 minutes, and nothing was calling. The
personal instance had not run a background job since 2026-05-14 and skudak since
2024-11-20. Consequently trash and file versions never expired, stale chunked
uploads accumulated, calendar reminders never fired, and nextcloud.log was never
rotated -- which quietly made the existing log_rotate_size cap inert. Added a
systemd timer per instance, skipping cleanly when the container is down or in
maintenance so deploy windows don't show up as failed units.
Trash retention on the personal instance: the default "auto" only expires when
disk space demands it, so 66 GB of >30-day deletions sat on a host with 1.3 TB
free -- effectively unbounded. "auto, 30" makes the 30-day expiry unconditional
while still purging early under pressure.
Image bumps:
nextcloud 33.0.0 -> 34.0.2 (both cloud and skudak-cloud)
greg-time-bot 3.9.25 -> 3.10.0
fulfillr 20260723.2044 -> 20260728.2155 (prod and dev)
The fulfillr bump records what is already deployed: both containers were rolled
to that image on 2026-07-29 for SCRUM-156 (digital product releases + customer
update campaign). Committing it keeps the repo from claiming an older tag than
the host is actually running, which would otherwise roll fulfillr backwards on
the next clean-checkout deploy.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The data-only rsync left no way to restore a working instance: mysql/ and
config/ were never backed up, so a recovery would have files but no shares,
users or metadata. Dump the database before syncing files (a DB older than
the files is repairable with occ files:scan; a newer one references blobs
that never made it into the backup) and ship config/ alongside it.
Capture the --chmod=Du=rwx,Dgo=rx flag that had been hand-added to the
deployed skudak-cloud script. It was outside git, so every deploy silently
reverted it. It now lives in backup_rsync_extra_args.
Add OnFailure= alerting. The units failed silently before, which is how an
iDrive sync failure sat unnoticed since May. msmtp rather than the esmtp
already installed: the OpenSRS relay is port 465 (implicit TLS) and libesmtp
only speaks STARTTLS.
Exclude nextcloud.log* from the sync and cap log_rotate_size. skudak-cloud
was running at loglevel 0 and had written a 64 GB log that was being rsynced
and pushed to S3; set it to 2 to match the home instance.
Stagger the timers (04:00 / 04:30) so both finish before the 05:00 TrueNAS
snapshot task, and bound TimeoutStartSec so a wedged rsync cannot leave the
unit activating forever and skip every subsequent trigger.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Ollama role and SearXNG container backed FISTO AI responses in the
greg-time Discord bot. greg-time 3.9.6 drops both (plus the Gemini path)
in favor of a single xAI Grok backend, so:
- remove the ollama role and its wiring in deploy_home.yml
- remove the searxng container task, template, and searxng_path default
- gregtime: swap OLLAMA_*/SEARXNG_URL/GEMINI_API_KEY env for XAI_API_KEY,
bump image 3.6.5 -> 3.9.6
- vault: add xai_api_key, drop gemini_api_key
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
SSH keys moved to /etc/ssh/backup_keys/ (ssh_home_t) and backup scripts
to /usr/local/bin/ (bin_t) to fix SELinux denials - container_file_t
context blocked rsync from exec'ing ssh. Also fixes skudak key path
mismatch (was truenas_skudak, key deployed as truenas_skudak-cloud).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replaces old 168-mod collection (3636931465) with new 385-mod collection.
Cleaned BBCode artifacts from mod IDs, updated map folders for 32 maps.
LogCabin retained for player connect/disconnect logging.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add ollama role for local LLM inference (install, service, models)
- Add searxng container for private search
- Migrate hostname from home.bdebyl.net to home.debyl.io
(inventory, awsddns, zomboid entrypoint, home_server_name)
- Update vault with new secrets
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Benchmarked uncensored models for the gregtime FISTO bot. dolphin-mistral
produces the best uncensored creative content, dolphin-phi is faster fallback.
Added OLLAMA_NUM_PREDICT env var (300) and bumped image to 3.3.0.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add uptime-kuma-personal container on port 3002
- Add Caddy config for uptime.debyl.io with IP restriction
- Update both uptime-kuma instances to 2.0.2
- Rename debyltech tag from uptime-kuma to uptime-debyltech
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Deploy systemd timer that writes zomboid container stats to
zomboid-stats.json every 30 seconds for gregtime to read.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit finalizes the comprehensive migration from nginx + ModSecurity + manual LetsEncrypt
to Caddy v2 with automatic HTTPS. The migration eliminates over 2000 lines of complex
configuration in favor of a single, simplified Caddyfile.
## Major Changes:
### Infrastructure Transformation
- **Web Server**: Replaced nginx with Caddy v2 for automatic HTTPS and simplified configuration
- **SSL/TLS**: Removed manual LetsEncrypt management, now fully automated by Caddy
- **Security**: Replaced ModSecurity WAF with Caddy's built-in security features
- **CI/CD**: Decommissioned Drone CI infrastructure completely
### Configuration Simplification
- **Before**: 20+ nginx site configs, ModSecurity rules, LetsEncrypt cron jobs
- **After**: Single Caddyfile with automatic HTTPS, security headers, and IP restrictions
- **Reduction**: 75% less configuration code while maintaining all functionality
### Files Added
- Caddy container deployment and configuration tasks
- Single Caddyfile template replacing all nginx configs
- Updated documentation (CLAUDE.md, TODO.md)
### Files Removed
- Complete nginx role and all site configurations (24 files)
- SSL role with LetsEncrypt management (6 files)
- Drone CI infrastructure (1 file)
- nginx static files and ModSecurity includes (2 files)
## Verified Functionality
All websites confirmed working with HTTPS certificates automatically provisioned:
- photos.bdebyl.net, parts.bdebyl.net, cloud.bdebyl.net
- wiki.skudakrennsport.com, cloud.skudakrennsport.com
- fulfillr.debyltech.com (with IP restrictions)
- Proper security headers and WebSocket support
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>