Commit Graph
65 Commits
Author SHA1 Message Date
Bastian de Byl f7f4d903a0 Merge remote-tracking branch 'origin/master' into feat/ci-images-registry 2026-09-14 19:57:33 -04:00
Bastian de BylandClaude Opus 5 42e6f5271d fix(gitea-actions): serve CI images from the Gitea registry
The CI job images only existed under localhost/ in the gitea-runner store,
and the nightly CI prune deletes any image older than 48h that no container
holds. After every idle stretch CI failed in 0-1s pulling
localhost/gitea-ci:latest until the role was re-run and the images rebuilt.

- Build under git.debyl.io/gitbot/..., push after every run, and pull from the
  registry instead of rebuilding when the Containerfile is unchanged.
- Log gitea-runner in via ~/.docker/config.json, which both act_runner (job
  image pulls) and podman read.
- Label the base images io.debyl.ci-base and skip that label in the CI prune;
  its `until` counts from build time, so a re-pulled image would otherwise be
  deleted again the next night.

Workflows pinning `container: image: localhost/gitea-ci-*` must move to the
registry paths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 16:40:57 -04:00
Bastian de BylandClaude Opus 5 e174e259eb SCRUM-196: Read GA4 key from fulfillr_ga4_credentials_json
Match the vault variable name holding the service-account key file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 16:21:04 -04:00
Bastian de BylandClaude Opus 5 e681a46b78 SCRUM-196: GA4 analytics config for fulfillr Traffic & funnel tab
Render an analytics block (property 353859448 + service-account key) into the
fulfillr dev and prod configs once fulfillr_ga4_credentials is in the vault.
Without the vault var the block is omitted and the portal reports GA as not
connected. Remember to restart the container after deploy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 14:02:38 -04:00
Bastian de BylandClaude Opus 5 f59ade748b feat(rsvp): deploy rsvp.debyl.io, the invite-only party RSVP app
A single Go binary with SQLite, built and loaded as localhost/rsvpd:<VERSION>
by make deploy-remote in ~/src/rsvp-debylio. Both of its listeners are
published on 127.0.0.1 only: Caddy proxies the public one to everyone and the
admin one (/admin, no login) only to caddy_local_networks. A Caddyfile mistake
alone cannot expose admin, and neither can a port mistake alone.

Guests' invite links are the credential and they sit in the URL path, which
shapes the vhost:

- It does not import common_headers. That snippet sets Referrer-Policy
  same-origin, which would replace the app's no-referrer and let a token leak
  in a Referer header.
- Its access log rewrites request>uri to /i/REDACTED and drops the Location
  response header, since every POST 303s back to /i/<token>.
- Caddy's error logger is separate from the site's and wrote the raw URI to
  caddy.log when the upstream was down. The global log now excludes
  http.log.error.rsvp and a filtered rsvp-errors logger takes it instead.
  Verified with zero token occurrences in both logs, locally and live.

The data directory is owned directly by the host uid of the container's uid
10001 (subuid + 10000). Setting it to the podman user and chowning back each run
flipped ownership on every deploy and briefly locked the app out of its
database.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:46 -04:00
Bastian de BylandClaude Opus 5 ba9c4f2bfe fix(zomboid): re-sync the world settings from the live server
The vendored Sophie preset and the running debbzoid world had drifted, and the
repo only held the preset. That is not a restore point: force-pushing it would
have reverted the admins' in-game tuning rather than recovering it, which is
exactly how a Sophie world quietly became an Apocalypse one once already -- 139
values reverted, loot from 0.35 back to 0.9, CharacterFreePoints 0 to 60.

files/zomboid/sophie/SandboxVars.lua is now a snapshot of the live world taken
2026-08-31, not the preset as shipped. The modlist is untouched and still
upstream, which is why zomboid_preset_version now names the two halves and their
separate dates. server.ini.j2 carries the eight keys that had drifted:

  PlayerSafehouse              false -> true
  SafehouseAllowNonResidential false -> true   (the diner/gas-station case)
  SafehouseAllowRespawn        false -> true
  SafehouseAllowLoot           true  -> false
  SafehouseAllowFire           true  -> false
  TrashDeleteAll               false -> true
  MapRemotePlayerVisibility    1     -> 4
  ResetID                      6953472 -> 826046

ResetID is in that list on purpose, and matters most. It is the world's
soft-reset token: a file value that differs from the one the live world was
created with tells every connected client to roll a new character. Carrying the
live value makes a deliberate force-push a no-op instead of a server-wide wipe
prompt.

Spawn config gets its own switch, zomboid_spawn_force. spawnregions.lua and
spawnpoints.lua are the only config a running world re-reads -- at every server
start, where SandboxVars is read once, when the world is created -- so a spawn
edit is deployable on the live world without a wipe. Sharing zomboid_config_force
between them would have meant force-pushing the whole preset to land a one-line
spawn edit, rewriting the INI (hence the ResetID hazard above) and the world's
SandboxVars along with it.

config-template/ now tracks the repo unconditionally. Nothing on the host writes
that directory and the server cannot see it, so it has no hand edits to protect;
if it does not track the repo it is not a restore point, just an older world's
settings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
2026-09-05 09:11:31 -04:00
Bastian de BylandClaude Opus 5 ae99ba415e fix(immich): keep the ML models resident, and pin the ML URL explicitly
Smart search, face detection and OCR were all failing with

  Machine learning request to "http://immich-machine-learning:3003" failed

while the container still reported Up. Podman only sees PID 1: gunicorn's
master was alive and holding the listening socket, but its worker had died at a
WORKER TIMEOUT and was never respawned -- hence a connect timeout rather than a
refusal. The dead worker was a zombie whose remaining thread was stuck in
uninterruptible sleep in exit_mmap, so it survived SIGKILL, podman rm -f and
rm -f -t 0, and kept the container name and network alias until the host was
rebooted.

MACHINE_LEARNING_MODEL_TTL=0 addresses the cause rather than the symptom. The
default unloads models after 300s idle, so every search following a gap
reloaded four of them (CLIP, buffalo_l detection + recognition, PP-OCRv5) on a
CPU-only 4-core box and then tore those mappings back down -- and that teardown
is what wedged. Keeping them resident costs ~1-2 GB and removes the path.

IMMICH_MACHINE_LEARNING_URL is now set explicitly instead of relying on
immich's implicit default, so the ML container can be renamed without silently
losing search, and the name is a variable so a replacement can be stood up
beside a broken one without editing tasks.

Verified after deploy: zero ML failures, "in-memory cache with unloading
disabled", ping 200 from immich-server, 26/26 containers healthy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:39:57 -04:00
Bastian de BylandClaude Opus 5 2a6390c0ad fix(podman-prune): reap the CI stores, which had reached 113 GB
podman_prune_users listed only the podman and git users, so the two stores that
turn over fastest were never touched. gitea-runner had reached 1205 images /
113.1 GB with 100% of it reclaimable, and actions-runner had 137 exited job
containers. That layer count is what makes overlayfs lookups -- and so CI
itself -- slow; the disk was the lesser problem.

Split into two policies, because the stores are not the same kind of thing:

  service users keep the 30-day rollback window, and their containers are
  deliberately NOT pruned. They are the live services, and reaping one that
  merely happens to be stopped would turn a transient crash into a unit that
  cannot start again until the next deploy.

  CI users get 48h and their exited job containers reaped too. Build layers
  carry no rollback value. Containers are reaped BEFORE images on purpose: an
  exited container pins the image it ran from, so pruning images first would
  leave those layers behind for another day.

Timer moved weekly -> daily; a week of CI turnover is what let the store reach
113 GB between runs. Persistent=true is kept so a missed run catches up.

First run reclaimed 134 GB: gitea-runner 113.1 -> 4.2 GB, actions-runner
7.6 GB -> 0, podman 25.7 -> 15.1 GB. Disk 449G -> 315G, all 26 containers up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 11:39:45 -04:00
Bastian de BylandClaude Opus 5 05fc3b5c65 fix: drop VehicleDoorsHotkey, which never loaded
A Build 41 mod (versionMin=41.60) carried into Sophie's B42 list. It indexes
the ModOptions framework at VehicleDoorsHotkey_Options.lua:4, that framework is
not in the list, and so it dies there every boot:

  attempted index: ModOptions of non-table: null

Twenty-one exceptions a boot for a hotkey that has never once worked. Excluded
rather than fixed by adding ModOptions, since nothing else in the list wants it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 19:33:48 -04:00
Bastian de BylandClaude Opus 5 32748d1786 fix: stop the query filter banning every player who tries to connect
The anti-scanner rules dropped 53-byte Steam query packets above 1/hour with a
burst of 2, per source IP, on the theory that only scanners send queries and
that real players would already have been marked verified by the priority-2
rule when they sent something larger.

That premise is inverted. A client's first contact with the server IS a 53-byte
query, so nobody can be verified before querying, and nobody can query more
than twice an hour without being dropped. The counters said so plainly: five
packets had ever matched the verified-accept rule, against 30,563 drops.
fail2ban then banned each dropped player for a week -- 32 live bans, 550 total,
firing every ten minutes, every one of them a residential address.

Keeps the shape of the protection and moves the threshold somewhere no real
client reaches: opening the server browser or retrying a connection is a
handful of queries, a flood is thousands. Removes the old rules first so hosts
carrying them converge instead of stacking a second copy.

The bans themselves were also hooked into INPUT for tcp only, so they never
blocked the UDP game traffic they were meant to -- the drop rule was doing all
the damage on its own. Cleared the outstanding 32.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:26:45 -04:00
Bastian de BylandClaude Opus 5 91d9f6b419 fix: actually bound the Zomboid log archive, which logrotate could not
The logrotate config added with the Sophie overhaul claimed to bound PZ's own
logs. It did not. data/Logs had reached 18 GB across 112 entries and 146k
files, and rotation had never once fired there.

Two reasons, both wrong assumptions on my part. PZ rolls its logs into dated
directories -- logs_2025-12-14/ through logs_2026-08-26/, up to 279 MB each --
so the Logs/*.txt glob only ever matched a handful of loose files at the top.
And it rotates by size: not one file in that tree exceeds 100M, because the
growth is in the number of files, not the size of any of them.

Age-based pruning is the right tool, so this adds a daily zomboid-log-prune
timer keeping 14 days and deleting the emptied directories behind it. It runs
under podman unshare, since the files belong to the container's UID.

logrotate keeps server-console.txt, which is a single ever-growing file and
genuinely is what it is good at. Narrowed its scope to say so rather than
implying coverage it never had.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 02:34:30 -04:00
Bastian de BylandClaude Opus 5 a0421226a2 fix: correct four stale mod IDs in the Sophie preset
On 42.20.3 the server loaded 279 of the preset's 283 mods and reported the
other four as "required mod ... not found". None of it was a download problem:
all 256 workshop items fetched, and every one of these four is in the preset's
own WorkshopItems= list. Only the ID each is referenced by has drifted.

Checked against the id= field of the mod.info files on disk rather than folder
names, which are unrelated to mod IDs -- an early look at folder names was
misleading here.

  FWOBenchPress&Treadmill  ->  FWOBenchPressTreadmill  (stray ampersand)
  Ladders42131             ->  Ladders4220             (per-build variants;
                                                        42131 is the 42.16 one)
  ServingPlatesB42         ->  ServingPlates42
  NewMusic_OrchestraMix    ->  NewMusic_CMM            (renamed upstream to
                                                        "Classical Music Mix")

Kept as a rename map applied at template time rather than edited into the
vendored list, so the list stays a faithful copy of upstream and re-vendoring
a newer preset keeps the corrections.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 02:23:31 -04:00
Bastian de BylandClaude Opus 5 d34c6cd36d rebuild the Zomboid server on the Sophie 42 preset as "debbzoid"
The deployment had drifted badly. It was pinned to the Steam `-beta unstable`
branch, which stopped being right the moment B42 became the default at 42.20,
and it wrote every mod ID with the `\` prefix that B42 required only during
that unstable period and now rejects. Three server profiles were carried
around (vanilla, modded, b42revamp) whose mod lists were hand-curated blobs
from a Steam collection that has since moved on.

Moves to the Sophie 42 community preset, vendored from GerDeathstar/sophie-pz
("Sophie 42 Files.zip", 2026-07-31 == the b42 release tag): 283 mod IDs, 256
workshop items, map_distanciado over Muldraugh, plus its SandboxVars,
spawnregions and spawnpoints. Sophie's gameplay settings are kept exactly as
shipped -- PVP with its damage modifiers, the safety system, MaxPlayers=32,
PauseEmpty, no sleep, safehouses off. Only the keys this deployment actually
owns are templated over the top.

The modlist is stored as YAML lists rather than the INI's semicolon blobs, so
the next Sophie update produces a diff you can read. Exclusions live in
zomboid_mods_excluded with a reason each, seeded with IconsInventory -- already
absent upstream, listed so a re-vendor cannot quietly bring it back.

Config is now seeded before first boot instead of patched after it. The old
approach could not write the INI until the server had generated one, so every
setting went through lineinfile guarded on a stat; seeding the whole file from
the preset removes the chicken-and-egg and puts the config in git. force is
off by design: these files are a starting point, not managed state, so the
server can be stopped, hand-edited and regenerated without Ansible clobbering
the edits. Push them again deliberately with -e zomboid_config_force=true.

The B41 Discord keys were dead. B42 replaced DiscordChannel/DiscordChannelID
with DiscordChatChannel, which takes a channel name rather than a snowflake,
so the DiscordChannelID=... this role had been appending was a key the server
ignores and the chat bridge has not been working.

Renaming the server to debbzoid is what starts the new world: PZ derives the
INI, save directory and player DB from the name, so gregboid stays on disk
untouched as a rollback.

Re-enabled, because the reason it was off is now fixed. It was disabled for
saturating the SSD -- ~10 MB/s of log writes into the shared journal, which
starved Gitea CI badly enough to stretch a firmware build to 17 minutes. The
container now logs to its own rotating k8s-file instead of the journal every
other service shares, and logrotate caps the server's own server-console.txt
and Logs/ at 100M keeping 3. copytruncate is mandatory there: PZ holds those
fds for its whole life, so a rename would leave it writing into an unlinked
inode. Memory is unchanged; idle draw is nowhere near MAX_RAM.

Also stops the entrypoint recursively chowning the install tree on every boot.
With 256 workshop mods that is a large inode walk and pure churn once the
first run has set ownership.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:41:00 -04:00
Bastian de BylandClaude Opus 5 773a2bbc9c harden TrueNAS CIFS mounts so immich self-heals
TrueNAS was power-cycled, the CIFS mounts failed, and systemd never retried
-- mount units are not restarted on failure. SMB came back, nothing
remounted, and immich-server served an empty library for days while its
database still listed 14,181 assets pointing at /mnt/media/originals.

fstab carried no _netdev, no nofail and no automount, so there was no path
back without a human. Now:

  x-systemd.automount  any access re-attempts the mount; failure stops
                       being terminal
  _netdev / nofail     ordered after network-online, dead NAS cannot block boot
  soft                 I/O errors instead of blocking forever, so the
                       container can be restarted rather than wedging in
                       uninterruptible sleep
  idle-timeout         unmount when unused, clearing stale handles
  resilienthandles     SMB3 rides out brief blips

ansible.posix.mount mounts directly and never starts the generated
.automount unit, leaving the on-access trigger inactive -- enable it
explicitly, or the headline fix silently does nothing.

The containers are systemd USER units while the mounts are SYSTEM units, so
RequiresMountsFor= is unavailable. cifs-watchdog bridges the scopes: checks
health, recovers, and restarts ONLY immich-server (the sole consumer of both
paths; postgres/redis/ML use local volumes).

Two bugs the umount test caught, both worth knowing:
- `ls` cannot test mountedness. An unmounted mount point is an ordinary
  empty directory, so ls succeeds and recovery was skipped entirely.
- A drop repaired within a single run leaves prev=healthy, so keying the
  restart solely on the stored state skipped it while the container still
  held its stale view.

Also moves the SMB password out of /etc/fstab, which is 0644 and was
readable by every local user, into a 0600 credentials file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 14:49:22 -04:00
Bastian de BylandClaude Opus 5 fec7d62acb feat(skudak-cloud): repair LibreSign, brand its mail, add Redis
LibreSign had been silently broken since it was first deployed in
January. Every step of the old before-starting hook ended in `|| echo`,
so six months of failures logged nothing.

LibreSign repair
- Root cause was a stale config_path: a valid OpenSSL root CA existed at
  generation 1, a failed CFSSL attempt left an empty generation 2, and
  config_path was left pointing at the empty one. Regenerated as
  "Skudak LLP" (was the pre-rename "Skudak Rennsport LLP").
- Deleted the hook. Java/PDFtk/jSignPdf live under data/appdata_*, a
  persisted volume, so they only ever needed installing once. Install and
  verification are now explicit tasks that actually fail.
- PHP_MEMORY_LIMIT 1024M -- the 512M image default fails opaquely
  mid-signature. LC_ALL/LANG so the JVM is not ANSI_X3.4-1968.
- signature_render_mode=GRAPHIC_ONLY. Any other mode halves the stamp
  width and overlays a name/date block that collides with the drawn mark
  and duplicates what our documents already typeset. The value must be
  exactly GRAPHIC_ONLY; a bare "GRAPHIC" is accepted by occ, matches no
  radio in the UI, and silently reverts to default.
- write_qrcode_on_footer=false, written with --type=boolean because
  FooterHandler reads it via getValueBool and the typed appconfig API
  does not coerce a string "0". The validation URL text is kept.
- identification_documents=0 -- the default gates signing behind an ID
  upload plus admin approval, so signers saw no way to sign.
- shareapi_restrict_user_enumeration_full_match=no, so an email owned by
  an existing account can be added as a signer. Root cause is in core
  (MailPlugin.php:163), not LibreSign. Do NOT set full_match_email=no --
  that disables email signer search entirely.

Mail branding (skudakmail app)
- Two supported extension points, no core patch and no LibreSign fork:
  mail_template_class for layout, subjects, button labels and the footer
  LibreSign never adds; and a BeforeMessageSent listener to embed the
  wordmark as a cid: part so it survives remote-image blocking.
- A third listener adds scoped CSS fixing the signing page being clipped
  on iOS Safari (100vh -> 100dvh). Patched upstream too.
- skudakmail-verify.php.j2 asserts all of the above through the real
  useTemplate() path and fails the play on drift. Every assertion was
  proven to fail when deliberately regressed.

Redis
- memcache.locking was unset, so Nextcloud used DBLockingProvider and
  every file lock became a MariaDB write -- the contention behind the
  intermittent multi-second stalls. Verified after: db locks static,
  redis keys growing.
- requirepass lives in a mounted 0640 conf, not --requirepass, which
  would leak it into podman inspect, the systemd unit and ps. The file is
  chowned to uid 999 because redis-server does not run as root and the
  :ro mount stops the image fixing it itself.
- No maxmemory: cache is evictable, locks are NOT, and evicting a held
  lock permits concurrent writers to one file. No persistence either --
  a restored RDB could reinstate locks whose owner is long dead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 15:54:52 -04:00
Bastian de BylandClaude Opus 5 f11391b28f bound log growth and reclaim ~52 GB of container disk
Caddy was rotating on implicit defaults (100MiB/keep 10/90d) that were not
holding -- 20 rotated files per stream and a 190-day-old .gz, 1.3 GB across
16 log streams. Made explicit at 10MiB/keep 3/7d.

Note roll_size et al are subdirectives of `output file`, NOT of `log`.
Getting that wrong does not degrade gracefully: Caddy refuses to start on a
bad config, so every site went down until it was corrected. Worth a
`caddy validate` gate before reload.

journald had no SystemMaxUse and had reached 4 GB, drifting toward its
10%-of-filesystem default (~190 GB on this root). Capped at 500M.

Both are safe to keep short because fluent-bit ships the journal and every
Caddy access log into Graylog -- though note its GELF output has been
erroring for days, which weakens that premise and wants investigating.

The larger find was unrelated to logs: 896 images totalling 59.6 GB with
75% unused (94 tags of greg-time-bot, 73 of fulfillr -- one per deploy) and
5.4 GB of dangling volumes, mostly 804 MB Nextcloud /var/www/html trees
orphaned by container recreations. Pruned to 22 images / 15.4 GB, and added
a weekly timer keeping 30 days so a rollback still needs no rebuild.

Also dropped the decommissioned 6379/tcp redis rule (nothing listening;
Immich's redis is on the shared podman network) and the orphaned nosql, s3
and searxng volume dirs.

Backup log exclusions turned out to be unnecessary: Gitea logs to console
so its log dirs are empty, Nextcloud already excludes its own, BookStack
mounts only uploads, and Caddy is not backed up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:13:35 -04:00
Bastian de BylandClaude Opus 5 bc110ce69e back up Gitea + Skudak app data; drop PartKeepr and Pi-hole
Extends the Nextcloud backup machinery rather than adding a second
mechanism. cloud-backup.sh.j2 gains three guarded options, all no-ops for
the existing callers:

  backup_podman_user  Gitea runs rootless under `git`, not `podman`
  backup_db_type      postgres (Gitea) and mysql (BookStack) alongside
                      mariadb; each engine's completion trailer differs,
                      and grepping for the wrong one fails every run
  backup_sqlite_dbs   `sqlite3 .backup` for live WAL-mode SQLite, gated on
                      `pragma integrity_check` before promotion -- rsync
                      is either stale (no -wal) or torn (with it)

New instances: gitea-debyl, skudak-gitea, bookstack, partsy-skudak. The
alert handler is rendered once and shared, so its wording is now generic
rather than per-product; TAG stays nextcloud-backup because an external
Graylog rule matches on it.

`apply:` on the includes is load-bearing -- tags on a dynamic
include_tasks do not reach the tasks inside it.

Business data (skudak-gitea, bookstack, partsy-skudak) goes to TrueNAS
and on to Skudak's own iDrive account; the personal bucket's
/skudak*/** excludes are permanent, not a stopgap.

Removals: PartKeepr is superseded by Partsy, and its teardown never
finished -- it targeted /etc/systemd/system/podman-partkeepr*.service,
wrong prefix and wrong scope, leaving enabled user units in failed state.
Pi-hole's role was already orphaned (absent from deploy_home.yml); its
port 53 rule went with it after confirming nothing listens there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 18:08:47 -04:00
Bastian de BylandClaude Opus 5 0ab423ca55 harden nextcloud backups: db dumps, alerting, drift fix
The data-only rsync left no way to restore a working instance: mysql/ and
config/ were never backed up, so a recovery would have files but no shares,
users or metadata. Dump the database before syncing files (a DB older than
the files is repairable with occ files:scan; a newer one references blobs
that never made it into the backup) and ship config/ alongside it.

Capture the --chmod=Du=rwx,Dgo=rx flag that had been hand-added to the
deployed skudak-cloud script. It was outside git, so every deploy silently
reverted it. It now lives in backup_rsync_extra_args.

Add OnFailure= alerting. The units failed silently before, which is how an
iDrive sync failure sat unnoticed since May. msmtp rather than the esmtp
already installed: the OpenSRS relay is port 465 (implicit TLS) and libesmtp
only speaks STARTTLS.

Exclude nextcloud.log* from the sync and cap log_rotate_size. skudak-cloud
was running at loglevel 0 and had written a 64 GB log that was being rsynced
and pushed to S3; set it to 2 to match the home instance.

Stagger the timers (04:00 / 04:30) so both finish before the 05:00 TrueNAS
snapshot task, and bound TimeoutStartSec so a wedged rsync cannot leave the
unit activating forever and skip every subsequent trigger.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 16:03:18 -04:00
Bastian de BylandClaude Opus 4.8 7d4a398bba Drop self-hosted AI (Ollama + SearXNG); gregtime switches to xAI Grok
The Ollama role and SearXNG container backed FISTO AI responses in the
greg-time Discord bot. greg-time 3.9.6 drops both (plus the Gemini path)
in favor of a single xAI Grok backend, so:

- remove the ollama role and its wiring in deploy_home.yml
- remove the searxng container task, template, and searxng_path default
- gregtime: swap OLLAMA_*/SEARXNG_URL/GEMINI_API_KEY env for XAI_API_KEY,
  bump image 3.6.5 -> 3.9.6
- vault: add xai_api_key, drop gemini_api_key

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:19:45 -04:00
Bastian de Byl 2335b4980d fulfillr(prod): wire prod Turso store + live Stripe (fulfillr_prod_* vars) + image 20260606.1735 2026-06-06 17:28:00 -04:00
Bastian de BylandClaude Opus 4.8 e82ace6de3 fulfillr-dev: staging back-office container + Turso store prep
Add a second go-fulfillr container (fulfillr-dev) wired to the staging
Turso store + EasyPost/Stripe test keys via dev.json, served at
fulfillr-dev.debyltech.com (Caddy -> :9055), LAN-restricted like prod.

- fulfillr-dev.yml + dev.json.j2: the staging container, volumes, config
- defaults: fulfillr_dev_* vars; prod store URL stubbed off until cutover
- Caddyfile + caddy.yml: fulfillr-dev site block and static mount
- awsddns.yml: Route53 DDNS for the fulfillr-dev hostname
- production.json.j2: add store_database_url/store_auth, rename stripe key
  var to fulfillr_stripe_api_key
- vault.yml: dev + store/stripe secrets

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 00:23:07 -04:00
Bastian de BylandClaude Opus 4.8 2df5b7fc03 Deploy fulfillr 20260603.0222 and wire tickets_table
Bump fulfillr image to the build with the tickets feature, and add the
tickets_table to the fulfillr production.json config (new debyltech-tickets-prod
DynamoDB table) so the /api/v1/tickets routes register.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 22:32:52 -04:00
Bastian de BylandClaude Opus 4.7 829befeb1c chore: bump container versions and remove n8n
- gitea: 1.25.2 -> 1.26.1 (debyl + skudak)
- caddy: 2.10.2 -> 2.11.2
- uptime-kuma: 2.0.2 -> 2.3.2 (debyl + skudak)
- bookstack: 25.7 -> 26.3.4
- home-assistant: 2026.1 -> 2026.5.1
- immich (server + ML): v2.5.0 -> v2.7.5
- remove n8n service (unused)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 15:44:09 -04:00
Bastian de BylandClaude Opus 4.7 4cc65f2a99 feat: deploy go-fulfillr cases dashboard to home.debyl.io
- Bump fulfillr container image from 20260124.0411 to 20260509.1940
  (built from go-fulfillr commit 48b9f60 which adds /api/v1/cases
  endpoints for the contact-form CRM dashboard).
- Add fulfillr_cases_table default ("debyltech-cases-prod") so the
  HasCasesConfig() guard flips on at startup and the cases routes
  register.
- Add cases_table to production.json.j2 so it lands in /config inside
  the container.

Verified after deploy: GET /api/v1/cases returns the existing test
cases, PATCH succeeds, GSI1PK rewrite works.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 16:04:52 -04:00
Bastian de BylandClaude Opus 4.6 43fbcf59a5 add n8n workflow automation and fix cloud backup rsync
- Add n8n container (n8nio/n8n:2.11.3) with Caddy reverse proxy at n8n.debyl.io
- Add --exclude .ssh to cloud backup rsync to prevent overwriting
  authorized_keys on TrueNAS backup targets

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 12:12:19 -04:00
Bastian de BylandClaude Opus 4.6 8fd220a16e noticket - update zomboid b42revamp modpack to collection 3672556207
Replaces old 168-mod collection (3636931465) with new 385-mod collection.
Cleaned BBCode artifacts from mod IDs, updated map folders for 32 maps.
LogCabin retained for player connect/disconnect logging.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 13:59:33 -05:00
Bastian de BylandClaude Opus 4.6 3637b3ba23 noticket - remove karrio, update gregtime, fix caddy duplicate redirect
Remove Karrio shipping platform (containers, config, vault secrets,
Caddy site block). Bump gregtime 3.4.1 -> 3.4.3. Remove duplicate
home.debyl.io redirect in Caddyfile. Update zomboid b42revamp mod list.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-24 17:40:00 -05:00
Bastian de BylandClaude Opus 4.6 495943b837 feat: add ollama and searxng, migrate to debyl.io hostname
- Add ollama role for local LLM inference (install, service, models)
- Add searxng container for private search
- Migrate hostname from home.bdebyl.net to home.debyl.io
  (inventory, awsddns, zomboid entrypoint, home_server_name)
- Update vault with new secrets

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-12 15:13:25 -05:00
Bastian de BylandClaude Opus 4.5 9d562c7188 feat: smart zomboid traffic filtering with packet-size detection
Replace per-IP hashlimit with smarter filtering that distinguishes
legitimate players from scanner bots based on packet behavior:
- Players send varied packet sizes (53, 37, 1472 bytes)
- Scanners only send 53-byte query packets

New firewall rule chain:
- Priority 2: Mark + ACCEPT non-query packets (verifies player)
- Priority 3: ACCEPT queries from verified IPs (1 hour TTL)
- Priority 4: LOG rate-limited queries from unverified IPs
- Priority 5: DROP rate-limited queries (2 burst, then 1/hour)

Also includes:
- Fail2ban zomboid jail with tighter thresholds (5 retries/4h, 1w ban)
- Graylog streams for zomboid-connections, zomboid-ratelimit, fail2ban
- GeoIP pipeline enrichment for zomboid traffic
- Fluent-bit inputs for ratelimit logs and fail2ban events
- Remove Legendary Katana mod (Workshop 3418366499) - removed from Steam
- Bump Immich to v2.5.0
- Fix fulfillr config (nil → null)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-27 15:09:26 -05:00
Bastian de BylandClaude Opus 4.5 33eceff1fe feat: add personal uptime kuma instance at uptime.debyl.io
- Add uptime-kuma-personal container on port 3002
- Add Caddy config for uptime.debyl.io with IP restriction
- Update both uptime-kuma instances to 2.0.2
- Rename debyltech tag from uptime-kuma to uptime-debyltech

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-27 08:04:33 -05:00
Bastian de Byl bc26fcd1f9 chore: fluent-bit zomboid, zomboid stats, home assistant, gregbot 2026-01-24 17:08:05 -05:00
Bastian de BylandClaude Opus 4.5 9e04727b0e feat: update zomboid b42revamp server name and mods
- Rename b42revamp server from "zomboidb42revamp" to "gregboid"
- Remove mod 3238830225 from workshop items
- Replace Real Firearms with B42RainsFirearmsAndGunPartsExpanded4213
- Remove 2788256295/ammomaker mod

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-22 23:11:56 -05:00
Bastian de BylandClaude Opus 4.5 c96aeafb3f feat: add git.skudak.com Gitea instance and skudak domain migrations
Gitea Skudak (git.skudak.com):
- New Gitea instance with PostgreSQL in podman pod under git user
- SSH access via Gitea's built-in SSH server on port 2222
- Registration restricted to @skudak.com emails with email confirmation
- SMTP configured for email delivery

Domain migrations:
- wiki.skudakrennsport.com → wiki.skudak.com (302 redirect)
- cloud.skudakrennsport.com + cloud.skudak.com (dual-domain serving)
- BookStack APP_URL updated to wiki.skudak.com
- Nextcloud trusted_domains updated for cloud.skudak.com

Infrastructure:
- SELinux context for git user container storage (container_file_t)
- Firewall rule for port 2222/tcp (Gitea Skudak SSH)
- Caddy reverse proxy for git.skudak.com

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-15 22:27:02 -05:00
Bastian de Byl 9e665a841d chore: non-cifs nextcloud, partsy, zomboid updates 2026-01-15 16:48:07 -05:00
Bastian de BylandClaude Opus 4.5 6af3c5dc69 feat: add comprehensive access logging to Graylog with GeoIP
- Add fluent-bit inputs for Caddy access logs (JSON) and SSH logs
- Create GeoIP task to download MaxMind GeoLite2-City database
- Mount GeoIP database in Graylog container
- Enable Gitea access logging via environment variables
- Add parsers.conf for Caddy JSON log parsing
- Remove unused nosql/redis container and configuration

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-09 15:16:21 -05:00
Bastian de BylandClaude Opus 4.5 cf200d82d6 chore: gitea-actions improvements, graylog/fluent-bit logging, zomboid mod
- Gitea actions: add handlers, improve deps and service template
- Graylog: simplify container config, add Caddy reverse proxy
- Add fluent-bit container for log forwarding
- Add ClimbDownRope mod (Workshop ID: 3000725405) to zomboid

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-03 17:20:18 -05:00
Bastian de BylandClaude Opus 4.5 2fd44fd450 feat: deploy gelf-proxy as container via Gitea registry
- Add Gitea container registry login task
- Add graylog.yml with full stack (MongoDB, OpenSearch, Graylog, gelf-proxy)
- Use container image instead of binary for gelf-proxy
- Image tagged from git.debyl.io/debyltech/gelf-proxy

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-31 18:53:36 -05:00
Bastian de Byl 4d835e86a0 chore: zomboid improvements, gregtime improvements with rcon 2025-12-22 12:31:43 -05:00
Bastian de Byl f9507f4685 chore: zomboid mod updates 2025-12-19 19:45:38 -05:00
Bastian de Byl 38561cb968 gitea, zomboid updates, ssh key fixes 2025-12-19 10:39:56 -05:00
Bastian de Byl 8c21923358 zomboid added, caddyfile updates, debylio migration, ddns migration 2025-12-13 21:18:33 -05:00
Bastian de Byl 28fe5937fe updates for gregtime, caddyfile, added uptime-kuma 2025-11-02 14:18:45 -05:00
Bastian de Byl 37c7259cf7 replace partkeepr with partsy, make private 2025-10-21 16:40:56 -04:00
Bastian de BylandClaude 9c9da4f47c Complete infrastructure migration from nginx + ModSecurity to Caddy
This commit finalizes the comprehensive migration from nginx + ModSecurity + manual LetsEncrypt
to Caddy v2 with automatic HTTPS. The migration eliminates over 2000 lines of complex
configuration in favor of a single, simplified Caddyfile.

## Major Changes:

### Infrastructure Transformation
- **Web Server**: Replaced nginx with Caddy v2 for automatic HTTPS and simplified configuration
- **SSL/TLS**: Removed manual LetsEncrypt management, now fully automated by Caddy
- **Security**: Replaced ModSecurity WAF with Caddy's built-in security features
- **CI/CD**: Decommissioned Drone CI infrastructure completely

### Configuration Simplification
- **Before**: 20+ nginx site configs, ModSecurity rules, LetsEncrypt cron jobs
- **After**: Single Caddyfile with automatic HTTPS, security headers, and IP restrictions
- **Reduction**: 75% less configuration code while maintaining all functionality

### Files Added
- Caddy container deployment and configuration tasks
- Single Caddyfile template replacing all nginx configs
- Updated documentation (CLAUDE.md, TODO.md)

### Files Removed
- Complete nginx role and all site configurations (24 files)
- SSL role with LetsEncrypt management (6 files)
- Drone CI infrastructure (1 file)
- nginx static files and ModSecurity includes (2 files)

## Verified Functionality
All websites confirmed working with HTTPS certificates automatically provisioned:
- photos.bdebyl.net, parts.bdebyl.net, cloud.bdebyl.net
- wiki.skudakrennsport.com, cloud.skudakrennsport.com
- fulfillr.debyltech.com (with IP restrictions)
- Proper security headers and WebSocket support

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-11 20:38:45 -04:00
Bastian de Byl a6df909de8 noticket - removed logs references 2025-03-06 12:42:14 -05:00
Bastian de Byl 6b813362ca noticket - cleanup of unused sites, containers 2025-03-01 20:47:53 -05:00
Bastian de Byl 761bb67b5c noticket - add self-hosted bitwarden for skudak 2025-02-07 19:39:32 -05:00
Bastian de Byl fced2a0038 noticket - add base site, update secrets 2025-02-03 12:34:41 -05:00
Bastian de Byl 184cd2574d noticket - reorganized podman 2024-02-01 15:35:11 -05:00
Bastian de Byl 8bd4ee9dd2 noticket - added skudak cloud (nextcloud) 2023-10-05 12:08:22 -04:00