Compare commits

...
6 Commits
Author SHA1 Message Date
Bastian de BylandClaude Opus 5 f59ade748b feat(rsvp): deploy rsvp.debyl.io, the invite-only party RSVP app
A single Go binary with SQLite, built and loaded as localhost/rsvpd:<VERSION>
by make deploy-remote in ~/src/rsvp-debylio. Both of its listeners are
published on 127.0.0.1 only: Caddy proxies the public one to everyone and the
admin one (/admin, no login) only to caddy_local_networks. A Caddyfile mistake
alone cannot expose admin, and neither can a port mistake alone.

Guests' invite links are the credential and they sit in the URL path, which
shapes the vhost:

- It does not import common_headers. That snippet sets Referrer-Policy
  same-origin, which would replace the app's no-referrer and let a token leak
  in a Referer header.
- Its access log rewrites request>uri to /i/REDACTED and drops the Location
  response header, since every POST 303s back to /i/<token>.
- Caddy's error logger is separate from the site's and wrote the raw URI to
  caddy.log when the upstream was down. The global log now excludes
  http.log.error.rsvp and a filtered rsvp-errors logger takes it instead.
  Verified with zero token occurrences in both logs, locally and live.

The data directory is owned directly by the host uid of the container's uid
10001 (subuid + 10000). Setting it to the podman user and chowning back each run
flipped ownership on every deploy and briefly locked the app out of its
database.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:46 -04:00
Bastian de BylandClaude Opus 5 a9f51b77e1 fix(podman): pull a new image before removing the running container
podman-check deleted the old container as soon as the pinned image differed,
and the create task pulled afterwards. A tag that did not exist, or a registry
that was down, therefore left the service with no container at all. The pull
now happens first, so that failure stops the play with the old container still
running. localhost/ images are built and loaded by hand and are never pulled.

The pull is skipped when the container does not exist yet: there is nothing to
protect, and containers[0] is not there to compare against. Without that guard
the first deploy of any new service failed on the conditional -- rsvp was the
first to hit it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:46 -04:00
Bastian de BylandClaude Opus 5 73e50bb19d chore(gregtime): bump to 3.17.3
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:45 -04:00
Bastian de BylandClaude Opus 5 6cd4d56de1 feat(labelprint): 4x6 label print proxy on a Raspberry Pi
A Pi 3B+ (stickah.local) shares a Phomemo PM246 to the LAN as a plain CUPS
queue, so any machine can print 4x6 labels -- fulfillr-site's shipping labels
in particular -- without installing the vendor driver, which is x86-64 only.
The role builds the TSPL CUPS driver from source instead.

It is Debian, not Fedora, so it lives in its own inventory and playbook
(make deploy-labelprint / check-labelprint) and the home.debyl.io roles can
never run against it. make bootfs renders its cloud-init first-boot files onto
a freshly imaged SD card from the same templates the role uses. The Wi-Fi
credentials for the home and rescue networks are in the vault.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:14:45 -04:00
Bastian de BylandClaude Opus 5 be0d02b938 feat(zomboid): restore the world from one of PZ's own backups
Rolling a world back meant hand-work over ssh: stop the service, move the live
save aside, unzip the right archive, chown into the container's subuid range,
relabel, start. That is the wrong shape of task to do by hand, and it is always
done under time pressure -- by construction, because the archive you want is
being deleted while you work.

PZ keeps BackupsCount=10 per set and writes one every BackupsPeriod=30 minutes,
so a periodic backup is reachable for about five hours and then gone. On
2026-09-05 the snapshot the admins asked for (05:16, four minutes before the
incident) had about 90 minutes of life left when the request came in.

Same shape as the wipe: the Discord bot writes a trigger file into its own rw
volume, zomboid-restore.path notices it, and zomboid-restore.service runs the
script as the podman user. The bot gets no ssh, no systemd, and keeps only its
existing read-only mount of the Zomboid volume.

Two details carry most of the correctness.

Resolution is by mtime, not by index. The rotation renames the files -- today's
backup_7.zip is backup_8.zip half an hour from now, and a new backup_7.zip holds
a different world -- so an index is valid only while the listing is fresh, which
is not long enough to survive a human reading a confirmation prompt. The trigger
names a set and an mtime; the script resolves the path itself, whitelists the
filename, and refuses if nothing matches. It never accepts a path.

Everything that can fail is checked before the server is touched. A rotated-out
target, an archive with no debbzoid world in it, a bad action, a traversal
attempt in the set name: each aborts with the server still running and writes a
result file the bot reports back. The live world is moved aside rather than
deleted, so a restore is undoable and the last three are kept.

One thing PZ does not advertise: its backups do not cover the whole save
directory. blam/, a mod's own state, is in none of them -- not the 05:16 archive
and not the newest one. Restoring only what the archive holds therefore lands
the world slightly *behind* the target rather than on it, so anything present in
the displaced world and absent from the archive is carried across.

The gregtime tag moves to 3.17.0 for the bot half of this -- `backups`,
`restore <n>`, `restore confirm`, `restore undo`, gated to the same two admins
as the wipe. That image is built and running on the host; its source is not
committed yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
2026-09-05 09:12:09 -04:00
Bastian de BylandClaude Opus 5 ba9c4f2bfe fix(zomboid): re-sync the world settings from the live server
The vendored Sophie preset and the running debbzoid world had drifted, and the
repo only held the preset. That is not a restore point: force-pushing it would
have reverted the admins' in-game tuning rather than recovering it, which is
exactly how a Sophie world quietly became an Apocalypse one once already -- 139
values reverted, loot from 0.35 back to 0.9, CharacterFreePoints 0 to 60.

files/zomboid/sophie/SandboxVars.lua is now a snapshot of the live world taken
2026-08-31, not the preset as shipped. The modlist is untouched and still
upstream, which is why zomboid_preset_version now names the two halves and their
separate dates. server.ini.j2 carries the eight keys that had drifted:

  PlayerSafehouse              false -> true
  SafehouseAllowNonResidential false -> true   (the diner/gas-station case)
  SafehouseAllowRespawn        false -> true
  SafehouseAllowLoot           true  -> false
  SafehouseAllowFire           true  -> false
  TrashDeleteAll               false -> true
  MapRemotePlayerVisibility    1     -> 4
  ResetID                      6953472 -> 826046

ResetID is in that list on purpose, and matters most. It is the world's
soft-reset token: a file value that differs from the one the live world was
created with tells every connected client to roll a new character. Carrying the
live value makes a deliberate force-push a no-op instead of a server-wide wipe
prompt.

Spawn config gets its own switch, zomboid_spawn_force. spawnregions.lua and
spawnpoints.lua are the only config a running world re-reads -- at every server
start, where SandboxVars is read once, when the world is created -- so a spawn
edit is deployable on the live world without a wipe. Sharing zomboid_config_force
between them would have meant force-pushing the whole preset to land a one-line
spawn edit, rewriting the INI (hence the ResetID hazard above) and the world's
SandboxVars along with it.

config-template/ now tracks the repo unconditionally. Nothing on the host writes
that directory and the server cannot see it, so it has no hand edits to protect;
if it does not track the repo it is not a restore point, just an older world's
settings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QdWYhwUtwM2NQGukiRh12
2026-09-05 09:11:31 -04:00
40 changed files with 2487 additions and 229 deletions
+20 -1
View File
@@ -14,6 +14,8 @@ This is a home infrastructure deployment repository using Ansible for automated
- `make deploy TAGS=sometag` - Deploy only specific tagged tasks
- `make deploy TARGET=specific-host` - Deploy to specific host instead of all
- `make check` - Run deployment in dry-run mode showing potential changes
- `make deploy-labelprint` / `make check-labelprint` - Deploy (or dry-run) the label print proxy Pi. Separate inventory and playbook from the home server - see `ansible/roles/labelprint/README.md`
- `make bootfs BOOTFS=/Volumes/bootfs` - Render the label print proxy's cloud-init first-boot files onto a freshly imaged SD card
- `make vault` - Edit encrypted Ansible vault file
- `make list-tags` - List all available Ansible tags
- `make list-tasks` - List all Ansible tasks
@@ -33,13 +35,16 @@ The project uses Python virtualenv for dependency management:
ansible/
├── deploy.yml # Main playbook entry point (imports deploy_home.yml)
├── deploy_home.yml # Core playbook with role definitions
├── deploy_labelprint.yml # Label print proxy Pi (separate host, Debian)
├── inventories/home/ # Inventory configuration
├── inventories/labelprint/ # Label print proxy inventory
├── roles/ # Ansible roles organized by function
│ ├── common/ # Base system configuration
│ ├── git/ # Git repository management
│ ├── podman/ # Container orchestration
│ ├── ssl/ # Legacy SSL management (deprecated - Caddy handles certificates automatically)
│ ├── github-actions/# CI/CD runner setup
│ ├── labelprint/ # 4x6 label print proxy (Raspberry Pi, CUPS/TSPL)
│ └── pihole/ # DNS filtering
└── vars/
└── vault.yml # Encrypted secrets
@@ -96,13 +101,27 @@ Tasks are tagged by service/component for selective deployment:
## Target Environment
- Single target host: `home.debyl.io`
- Primary target host: `home.debyl.io`
- OS: Fedora (ansible_user: fedora)
- Container runtime: Podman
- Web server: Caddy with automatic HTTPS and built-in security (replaced nginx + ModSecurity)
- All services accessible via HTTPS with automatic certificate renewal
- ~~CI/CD: Drone CI infrastructure completely decommissioned~~
### Label print proxy
- Second target host: `stickah.local` (Raspberry Pi 3B+, Raspberry Pi OS 64-bit,
apt not dnf) with a Phomemo PM246 4x6 label printer on USB
- Deployed only via `make deploy-labelprint` - it is in its own inventory so the
Fedora roles can never run against it
- The SD card image is written by hand with rpi-imager, but its cloud-init files
come from `make bootfs`, rendered from the same templates the role uses
- Login is `stickah` / `stickah` with SSH key auth preferred; password auth is a
deliberate LAN-only fallback so the box is never unreachable
- LAN-only (192.168.1.0/24), discovered over mDNS, no DNS record
- Falls back to its own rescue Wi-Fi AP at 192.168.4.1 when the home SSID is
unreachable
### Remote SSH Commands for Service Users
The `podman` user (and other service users) have `/bin/nologin` as their shell. To run commands as these users via SSH:
+19
View File
@@ -21,6 +21,9 @@ VAULT_FILE=ansible/vars/vault.yml
# Variables
ANSIBLE_INVENTORY=ansible/inventories/home/hosts.yml
# The label print proxy is a separate inventory and playbook: it is Debian, not
# Fedora, and shares none of the roles home.debyl.io runs.
ANSIBLE_INVENTORY_LABELPRINT=ansible/inventories/labelprint/hosts.yml
#SSH_KEY=${HOME}/.ssh/id_rsa_home_ansible
# Default to all ansible tags to run (passed via 'make deploy TAGS=sometag')
@@ -29,6 +32,9 @@ SKIP_TAGS?=none
TARGET?=all
EXTRA_VARS?=
# Mounted boot partition of the label print proxy's SD card (see `make bootfs`)
BOOTFS?=/Volumes/bootfs
${VENV}:
python3 -m venv ${VENV}
${VENV_BIN}/python3 -m pip install --upgrade pip
@@ -55,6 +61,19 @@ SKIP_FILE=./.lint-vars.sh
deploy: ${ANSIBLE} ${VAULT_FILE}
${ANSIBLE} --diff -t ${TAGS} --skip-tags ${SKIP_TAGS} -i ${ANSIBLE_INVENTORY} -l ${TARGET} --vault-password-file ${VAULT_PASS_FILE} $(if ${EXTRA_VARS},-e "${EXTRA_VARS}") ansible/deploy.yml
# Writes the Pi's cloud-init first-boot files onto a freshly imaged SD card.
bootfs: ${ANSIBLE} ${VAULT_FILE}
${ANSIBLE} --diff -i ${ANSIBLE_INVENTORY_LABELPRINT} --vault-password-file ${VAULT_PASS_FILE} -e "bootfs=${BOOTFS}" ansible/bootfs.yml
# Label print proxy (stickah.local). Override the address when the Pi has
# fallen back to its rescue AP:
# make deploy-labelprint EXTRA_VARS="ansible_host=192.168.4.1"
deploy-labelprint: ${ANSIBLE} ${VAULT_FILE}
${ANSIBLE} --diff -t ${TAGS} --skip-tags ${SKIP_TAGS} -i ${ANSIBLE_INVENTORY_LABELPRINT} -l ${TARGET} --vault-password-file ${VAULT_PASS_FILE} $(if ${EXTRA_VARS},-e "${EXTRA_VARS}") ansible/deploy_labelprint.yml
check-labelprint: ${ANSIBLE} ${VAULT_FILE}
${ANSIBLE} --check --diff -t ${TAGS} --skip-tags ${SKIP_TAGS} -i ${ANSIBLE_INVENTORY_LABELPRINT} -l ${TARGET} --vault-password-file ${VAULT_PASS_FILE} $(if ${EXTRA_VARS},-e "${EXTRA_VARS}") ansible/deploy_labelprint.yml
list-tags: ${ANSIBLE} ${VAULT_FILE}
${ANSIBLE} --list-tags -i ${ANSIBLE_INVENTORY} -l ${TARGET} --vault-password-file ${VAULT_PASS_FILE} ansible/deploy.yml
+106
View File
@@ -0,0 +1,106 @@
---
# Renders the Raspberry Pi's cloud-init first-boot files onto a freshly written
# SD card's boot partition:
#
# make bootfs BOOTFS=/Volumes/bootfs
#
# The image itself is still built by hand with rpi-imager -- this only writes
# the two cloud-init files onto it. Everything it writes comes from the same
# templates roles/labelprint uses, so the Pi boots with its Wi-Fi profiles and
# rescue access point already in place, and the first deploy has nothing to
# correct.
#
# The rendered files contain the Wi-Fi PSKs in the clear, as any Pi Wi-Fi setup
# does. They land on the SD card, never in this repo.
- hosts: localhost
gather_facts: false
connection: local
vars_files:
- vars/vault.yml
- roles/labelprint/defaults/main.yml
tasks:
- name: check that BOOTFS points at a Raspberry Pi boot partition
ansible.builtin.stat:
path: "{{ bootfs }}/config.txt"
register: labelprint_bootfs_check
tags: bootfs
- name: refuse to write to anything else
ansible.builtin.assert:
that: labelprint_bootfs_check.stat.exists
fail_msg: >-
{{ bootfs }} has no config.txt, so it is not a Raspberry Pi boot
partition. Write the image with rpi-imager first, then re-run with
BOOTFS pointing at the mounted boot volume.
tags: bootfs
- name: look for files rpi-imager already wrote
ansible.builtin.stat:
path: "{{ bootfs }}/{{ item }}"
register: labelprint_bootfs_existing
loop:
- user-data
- network-config
- cmdline.txt
tags: bootfs
# Keeps whatever rpi-imager put there, so a bad render can be undone by hand
# without reflashing. force:false means the first run's backup is the one
# that survives -- a second run must not overwrite it with our own output.
- name: back up the files rpi-imager wrote
ansible.builtin.copy:
src: "{{ bootfs }}/{{ item.item }}"
dest: "{{ bootfs }}/{{ item.item }}.rpi-imager.bak"
mode: "0644"
force: false
loop: "{{ labelprint_bootfs_existing.results }}"
loop_control:
label: "{{ item.item }}"
when: item.stat.exists
tags: bootfs
- name: render the cloud-init files
ansible.builtin.template:
src: "roles/labelprint/templates/bootfs/{{ item }}.j2"
dest: "{{ bootfs }}/{{ item }}"
mode: "0644"
loop:
- user-data
- network-config
tags: bootfs
- name: hash the rendered user-data
ansible.builtin.stat:
path: "{{ bootfs }}/user-data"
checksum_algorithm: sha1
register: labelprint_user_data_stat
tags: bootfs
- name: stamp the instance id with that hash
ansible.builtin.template:
src: roles/labelprint/templates/bootfs/meta-data.j2
dest: "{{ bootfs }}/meta-data"
mode: "0644"
vars:
labelprint_user_data_id: "{{ labelprint_user_data_stat.stat.checksum[:12] }}"
tags: bootfs
# rpi-imager writes `ds=nocloud;i=<id>` onto the kernel command line, and
# that id outranks the one in meta-data. Leave it alone and a card that has
# booted even once is seen by cloud-init as the same instance forever: it
# skips users, write_files and runcmd, silently, and the only symptom is a
# Pi that came up with none of this applied. Both places have to agree.
- name: pin the same instance id on the kernel command line
ansible.builtin.replace:
path: "{{ bootfs }}/cmdline.txt"
regexp: '(ds=nocloud[^\s]*?);i=[^\s]+'
replace: '\1;i={{ labelprint_hostname }}-{{ labelprint_user_data_stat.stat.checksum[:12] }}'
tags: bootfs
- name: what to do next
ansible.builtin.debug:
msg:
- "Wrote user-data, network-config, meta-data and cmdline.txt to {{ bootfs }}."
- "Instance id is {{ labelprint_hostname }}-{{ labelprint_user_data_stat.stat.checksum[:12] }}; the Pi re-applies this config whenever it changes."
- "Eject the volume, boot the Pi, then: make deploy-labelprint"
tags: bootfs
+14
View File
@@ -0,0 +1,14 @@
---
# Label print proxy (Raspberry Pi 3B+, Raspberry Pi OS trixie).
#
# The Pi image itself is built by hand with rpi-imager and written to the SD
# card outside of Ansible -- see roles/labelprint/README.md for what that image
# has to contain. Everything from first boot onwards lives in the labelprint
# role: the print driver, the CUPS queue, security updates, and the Wi-Fi
# rescue access point.
- hosts: labelprint
vars_files:
- vars/vault.yml
roles:
- role: labelprint
tags: labelprint
+16
View File
@@ -0,0 +1,16 @@
---
# The 4x6 label print proxy: a Raspberry Pi 3B+ with the Phomemo PM246 on USB,
# shared to the LAN over IPP/AirPrint. Deliberately a separate inventory from
# inventories/home: this host is Debian/apt and shares none of the Fedora roles
# that home.debyl.io runs, so `make deploy` can never reach it by accident.
#
# Reached by mDNS (avahi on the Pi, Bonjour/UDM Pro on the LAN) rather than by a
# DNS record. If the Pi has fallen back to its rescue access point, it is not on
# the LAN at all -- join the rescue SSID and deploy against 192.168.4.1 instead:
#
# make deploy-labelprint EXTRA_VARS="ansible_host=192.168.4.1"
labelprint:
hosts:
stickah.local:
ansible_user: stickah
ansible_python_interpreter: /usr/bin/python3
+162
View File
@@ -0,0 +1,162 @@
# labelprint — 4x6 label print proxy
A Raspberry Pi 3B+ (`stickah.local`) with a **Phomemo PM246** on USB, shared to
the LAN so any Mac, Windows or Linux machine can print 4x6 labels without
installing a printer driver. Built for
[`fulfillr-site`](../../../../debyltech/fulfillr-site), which prints shipping
labels straight out of the browser, but the queue is a plain 4x6 label printer
and anything can use it.
Deployed on its own:
```
make deploy-labelprint
make check-labelprint # dry run
make deploy-labelprint TAGS=cups # just the queue
```
## Why not the Phomemo driver
The vendor driver that runs on `yoga` installs
`/usr/lib/cups/filter/rastertolabeltspl`, and that file is an **x86-64 ELF**.
There is no ARM build, so none of it can be reused on the Pi.
The PM246 speaks TSPL over USB, so this role builds
[RunTheWall/tspl-cups-driver](https://github.com/RunTheWall/tspl-cups-driver)
(MIT) from a pinned commit instead: a CUPS raster filter, a backend and a PPD,
about a minute of compiling on the Pi. Pinned to a commit rather than installed
from the project's apt repo so that a version bump is a reviewed change in this
repo, and so the Pi carries no third-party signing key.
The PM246's USB id is not in the driver's auto-detect list. `tspl://auto` is
tried first; if the queue cannot find the printer, read the id off the Pi and
pin it — see `labelprint_device_uri` in [defaults/main.yml](defaults/main.yml).
## The image
The SD card image is built by hand, but the cloud-init files on it are not:
`make bootfs` renders them from the same templates the role uses, so the Pi
boots with its Wi-Fi profiles, its rescue access point and its login already in
place. Nothing here needs the printer driver — that is Ansible's job.
1. **Write the image** with rpi-imager: **Raspberry Pi OS (64-bit)**, Bookworm or
newer, so NetworkManager is the network stack. Verified against the
`2026-06-18` pi-gen build. Skip the customisation screen entirely — anything
set there is about to be overwritten.
2. **Render the boot files** onto the mounted boot partition:
```
make bootfs # defaults to /Volumes/bootfs
make bootfs BOOTFS=/path/to/boot
```
It refuses to write anywhere without a `config.txt`, and keeps whatever
rpi-imager wrote as `*.rpi-imager.bak`.
3. **Eject**, boot the Pi, and give cloud-init a couple of minutes.
4. `make deploy-labelprint`.
`make bootfs` is safe to re-run. The instance id is a hash of the rendered
`user-data`, so an unchanged render leaves the card alone, and a changed one
makes the Pi re-apply it on the next boot.
That id is written in **two** places, and they have to agree: `meta-data`, and
the `ds=nocloud;i=<id>` token rpi-imager puts on the kernel command line in
`cmdline.txt`. The command line wins. Leave rpi-imager's id there and a card
that has booted even once looks like the same instance to cloud-init forever —
it skips `users`, `write_files` and `runcmd` without a word, and the only
symptom is a Pi that came up with none of this applied. `make bootfs` rewrites
both.
To force a re-apply on a Pi that is already running, without pulling the card:
```
sudo sed -i 's/;i=[^ ]*/;i=<new-id>/' /boot/firmware/cmdline.txt
sudo cloud-init clean --logs
sudo reboot
```
### Getting in
- `ssh stickah@stickah.local` — your `~/.ssh/id_ed25519.pub` is installed, and
that is what Ansible uses.
- Password auth is **on**, with the password `stickah`. It is deliberately
trivial and deliberately not hashed: this is the fallback for a Pi that will
not take the key — reflashed card, someone else's laptop, standing at the
bench — and port 22 is reachable only from the LAN or from the rescue AP,
which has a real WPA2 password of its own. If that trade stops being the right
one, `labelprint_user_password` in [defaults/main.yml](defaults/main.yml) is
the only thing to change.
## Wi-Fi and the rescue AP
Normally the Pi is a station on the home SSID. When that network is
unreachable — the password changed, the router died, the Pi moved — the
`wifi-rescue` watchdog brings up an access point so there is still a way in:
- Join the rescue SSID (both it and its password are in the vault).
- The Pi is at **192.168.4.1**: `ssh pi@192.168.4.1`, or print directly to
`ipp://192.168.4.1:631/printers/labels`.
- Deploy to it there with
`make deploy-labelprint EXTRA_VARS="ansible_host=192.168.4.1"`.
The Pi also keeps the Wi-Fi profile netplan rendered from the image's original
`network-config`, at a lower autoconnect priority than `home-wifi`. It is a
deliberate fallback: if `home-wifi` is ever rendered wrong, the Pi still comes
back on the LAN rather than stranding itself on the rescue AP.
The 3B+ has one radio and cannot hold an AP and a station link at the same
time, so the watchdog cannot listen for the home SSID while the AP is up.
Instead the AP drops every five minutes to scan -- about seven seconds if the
home SSID is plainly gone, up to half a minute if it is worth a join attempt --
and comes back immediately if the home network is still missing. When the home SSID
does return, the Pi rejoins it and shuts the AP down by itself.
If you are working over the rescue AP and do not want your session cut:
```
touch /run/wifi-rescue.hold
```
The hold expires after 30 minutes, so a forgotten hold file cannot strand the
Pi. Deploys take the hold automatically for as long as they run.
Watch it work with `journalctl -t wifi-rescue -f`.
## Adding the printer from a client
Nothing to install anywhere — the Pi renders.
- **macOS** — System Settings → Printers → `+`. It appears under its own name
as an **AirPrint** printer. (`BrowseDNSSDSubTypes _print,_universal` in
cupsd.conf is what makes macOS offer the driverless add instead of guessing at
Generic PostScript.)
- **Windows 10/11** — Add printer; it is discovered as IPP Everywhere / Mopria.
- **Linux** — discovered by `cups-browsed`, or add
`ipp://stickah.local:631/printers/labels` by hand.
## Access
LAN only, `192.168.1.0/24`, plus the rescue AP subnet. Enforced twice on
purpose: at nftables ([templates/nftables.conf.j2](templates/nftables.conf.j2))
and again in cupsd's own `Location` blocks
([templates/cupsd.conf.j2](templates/cupsd.conf.j2)). CUPS is explicitly set
`--no-remote-any`; the upstream driver's `install.sh` turns that on, which would
offer this printer to anything that can route to the Pi.
The Pi patches itself: `unattended-upgrades` is on, with the Raspberry Pi
archives added to the origins allowlist (Debian's default covers only the Debian
security origin, which would leave the kernel and firmware — the packages most
specific to this hardware — unpatched). Kernel updates reboot at 04:00. Nothing
here holds state across a reboot; a queued job is spooled to disk and resumes.
## Secrets
In `ansible/vars/vault.yml` (`make vault`), no `vault_` prefix, per repo
convention:
| Key | What |
| --- | --- |
| `stickah_ssid` | Home SSID |
| `stickah_psk` | Home passphrase, or the 64-hex precomputed PSK |
| `stickah_ssid_rescue` | Rescue AP SSID |
| `stickah_psk_rescue` | Rescue AP passphrase (WPA2, 8+ characters) |
+118
View File
@@ -0,0 +1,118 @@
---
# ---------------------------------------------------------------------------
# Identity
# ---------------------------------------------------------------------------
# Reached as stickah.local. There is no DNS record for it: the UDM Pro passes
# mDNS across the LAN, so avahi on the Pi is the whole of the name service.
labelprint_hostname: stickah
# The login the image creates and Ansible connects as. Password auth stays on
# with a trivial, documented password: this box has to be reachable when the SSH
# key is not an option -- a reflashed card, a different laptop, someone standing
# at the workshop bench -- and it is only reachable from the LAN or from its own
# rescue AP in the first place. The key is what Ansible actually uses.
labelprint_user: stickah
labelprint_user_password: stickah
labelprint_authorized_key_file: ~/.ssh/id_ed25519.pub
# ---------------------------------------------------------------------------
# Printer
# ---------------------------------------------------------------------------
# Phomemo PM246, 4x6 direct thermal, 203 dpi, speaks TSPL over USB.
labelprint_queue: labels
labelprint_queue_info: 4x6 Label Printer (Phomemo PM246)
labelprint_queue_location: stickah
# The tspl backend finds the printer's usblp node by USB id. "auto" matches only
# the ids the driver already knows, and the PM246 is not yet one of them -- so if
# a deploy leaves the queue unable to find the printer, read the id off the Pi:
#
# for n in /dev/usb/lp*; do udevadm info -q property -n "$n" | grep -E 'ID_(VENDOR|MODEL)_ID|ID_SERIAL_SHORT'; done
#
# and pin it here as tspl://<vid>-<pid> (a DASH, not a colon: CUPS parses ":pid"
# as a port number and rejects the URI), or as tspl:///dev/usb/lp0.
labelprint_device_uri: "tspl://auto"
# 203dpi matches the PM246 head. The PPD defaults to 300dpi, which would render
# every label at ~2/3 scale on this printer.
labelprint_resolution: 203dpi
labelprint_media: na_index-4x6_4x6in
# Cut every other media size out of the PPD, so 4x6 is the only paper a client
# can pick. Driverless clients build their own PPD from the IPP media-supported
# list cupsd derives from ours, and there is no lpadmin option that restricts
# that list -- trimming the PPD is the only lever.
#
# The trade is real: the driver's PPD also offers 100x150mm, 4x4, 2.25x1.5, 2x1
# and a custom range, and this takes all of them away. Set false if you ever
# want this queue to run stock other than 4x6; a second queue off the untrimmed
# PPD is the better answer if you want both.
labelprint_media_only_4x6: true
# The PPD page-size keyword kept when the above is on. Pairs with
# labelprint_media, which is the same size under its IPP name.
labelprint_ppd_pagesize: w288h432
# 0-15. 8 is the driver's default and a sane starting point for the cheap
# thermal stock; raise it if barcodes scan poorly, lower it if edges bleed.
labelprint_darkness: 8
# in/sec x10. 40 = 4 in/sec.
labelprint_print_speed: 40
# ---------------------------------------------------------------------------
# Driver: RunTheWall/tspl-cups-driver (MIT)
# ---------------------------------------------------------------------------
# Built from source at a pinned commit rather than installed from the project's
# apt repo: this keeps a third-party signing key and package feed off the Pi,
# and makes the version we run a reviewed, deliberate bump in git history.
#
# The vendor Phomemo driver is not an option here -- its rastertolabeltspl
# filter ships as an x86-64 ELF only, and this host is aarch64.
labelprint_driver_repo: https://github.com/RunTheWall/tspl-cups-driver.git
labelprint_driver_version: f433b7774d80a4f6a901b6b998cb710fd79918a4
labelprint_driver_src: /usr/local/src/tspl-cups-driver
labelprint_ppd_dir: /usr/share/ppd/tspl
# The PPD the queue is actually built from.
labelprint_ppd_active: >-
{{ labelprint_ppd_dir }}/{{
'tspl-label-4x6.ppd' if labelprint_media_only_4x6 else 'tspl-label.ppd'
}}
# ---------------------------------------------------------------------------
# Network
# ---------------------------------------------------------------------------
# The only subnet allowed to reach CUPS. Everything else is dropped at nftables
# and refused again by cupsd's own access rules.
labelprint_lan_cidr: 192.168.1.0/24
# Rescue access point, brought up when the home SSID is unreachable. The Pi 3B+
# has a single radio and cannot hold an AP and a station link at once, so this
# is strictly a fallback -- see templates/wifi-rescue.sh.j2.
labelprint_ap_addr: 192.168.4.1
labelprint_ap_cidr: 192.168.4.0/24
# How often the watchdog checks, and how long the AP stays up before it drops
# for a few seconds to scan for the home SSID again.
labelprint_watchdog_interval_secs: 60
labelprint_ap_rescan_secs: 300
# How long a hand-placed /run/wifi-rescue.hold pins the radio before the
# watchdog ignores it. Bounded so a forgotten hold file cannot strand the Pi.
labelprint_hold_max_age_secs: 1800
# ---------------------------------------------------------------------------
# Packages
# ---------------------------------------------------------------------------
labelprint_deps:
[
avahi-daemon,
build-essential,
cups,
cups-filters,
dnsmasq-base,
git,
libcups2-dev,
network-manager,
nftables,
unattended-upgrades,
]
# Secrets live in ansible/vars/vault.yml (no vault_ prefix, per repo
# convention): stickah_ssid, stickah_psk,
# stickah_ssid_rescue, stickah_psk_rescue
@@ -0,0 +1,25 @@
# Cuts every media size except one out of a CUPS PPD.
#
# macOS and Windows add this printer driverless: they never see this PPD, they
# see the IPP media-supported list cupsd derives from it. Trimming here is
# therefore the only way to make 4x6 the single choice a client can offer --
# there is no lpadmin option that restricts the media list.
#
# Regenerated from the driver's own PPD on every deploy, so a driver version
# bump carries its new PPD through this filter rather than being pinned to a
# fork.
#
# awk -v keep=w288h432 -f trim-ppd-media.awk tspl-label.ppd
#
# The four families below are keyed by the PPD page-size keyword in field 2,
# which reads as "w288h432/4 x 6 in:" -- hence the split on "/".
/^\*(PageSize|PageRegion|ImageableArea|PaperDimension) / {
split($2, f, "/")
if (f[1] != keep) next
}
# Custom sizes would put "Manage Custom Sizes" back in the client's paper menu
# and let a job arrive at any dimension, which is the thing being prevented.
/^\*(CustomPageSize|ParamCustomPageSize|MaxMediaWidth|MaxMediaHeight)/ { next }
{ print }
@@ -0,0 +1,41 @@
---
- name: restart cups
become: true
ansible.builtin.systemd:
name: cups.service
state: restarted
- name: reload udev rules
become: true
ansible.builtin.command:
argv: [udevadm, control, --reload-rules]
changed_when: true
notify: trigger udev
# --action=add, not the default "change": udev only creates SYMLINK+= entries
# when a device is added, so a change event reloads the rule and leaves
# /dev/usb/tspl-label missing until the printer is next replugged or rebooted.
- name: trigger udev
become: true
ansible.builtin.command:
argv: [udevadm, trigger, --subsystem-match=usbmisc, --action=add]
changed_when: true
- name: reload systemd
become: true
ansible.builtin.systemd:
daemon_reload: true
# NetworkManager only reads new keyfiles from disk on request. This does not
# disturb the live connection.
- name: reload networkmanager connections
become: true
ansible.builtin.command:
argv: [nmcli, connection, reload]
changed_when: true
- name: reload nftables
become: true
ansible.builtin.systemd:
name: nftables.service
state: reloaded
+92
View File
@@ -0,0 +1,92 @@
---
- name: set the hostname
become: true
ansible.builtin.hostname:
name: "{{ labelprint_hostname }}"
tags: [labelprint, base]
# The hostname is also the mDNS name, and avahi publishes whatever is in
# /etc/hosts for 127.0.1.1. cloud-init writes this line on first boot from the
# image's own hostname, so it has to be corrected here too or the Pi answers to
# the wrong .local name.
- name: point 127.0.1.1 at the hostname
become: true
ansible.builtin.lineinfile:
path: /etc/hosts
regexp: '^127\.0\.1\.1\s'
line: "127.0.1.1\t{{ labelprint_hostname }}"
owner: root
group: root
mode: "0644"
tags: [labelprint, base]
- name: install the print proxy packages
become: true
ansible.builtin.apt:
name: "{{ labelprint_deps }}"
state: present
update_cache: true
cache_valid_time: 3600
tags: [labelprint, base]
- name: publish the host over mDNS
become: true
ansible.builtin.systemd:
name: avahi-daemon.service
enabled: true
state: started
tags: [labelprint, base]
# ---------------------------------------------------------------------------
# Unattended security updates
# ---------------------------------------------------------------------------
# This box sits on the LAN with an open IPP port and is not something anyone
# logs into for months at a time, so it patches itself.
- name: enable unattended upgrades
become: true
ansible.builtin.copy:
dest: /etc/apt/apt.conf.d/20auto-upgrades
content: |
APT::Periodic::Update-Package-Lists "1";
APT::Periodic::Unattended-Upgrade "1";
APT::Periodic::AutocleanInterval "7";
owner: root
group: root
mode: "0644"
tags: [labelprint, base, updates]
# Debian's stock 50unattended-upgrades allowlists the Debian security origin
# only. Raspberry Pi OS serves its own kernel, firmware and userland from the
# Raspberry Pi archives, so without these two extra origins the packages most
# specific to this hardware are exactly the ones that never get patched.
- name: allow the Raspberry Pi origins and reboot for kernel updates
become: true
ansible.builtin.copy:
dest: /etc/apt/apt.conf.d/52unattended-upgrades-labelprint
content: |
Unattended-Upgrade::Origins-Pattern {
"origin=Raspbian,codename=${distro_codename},label=Raspbian";
"origin=Raspberry Pi Foundation,codename=${distro_codename},label=Raspberry Pi Foundation";
};
Unattended-Upgrade::Remove-Unused-Kernel-Packages "true";
Unattended-Upgrade::Remove-Unused-Dependencies "true";
// Nothing here holds state across a reboot -- a queued job is spooled on
// disk and resumes -- so take the kernel update at 04:00 rather than
// leaving the Pi running an unpatched kernel until someone notices.
Unattended-Upgrade::Automatic-Reboot "true";
Unattended-Upgrade::Automatic-Reboot-Time "04:00";
owner: root
group: root
mode: "0644"
tags: [labelprint, base, updates]
- name: enable the unattended-upgrades timers
become: true
ansible.builtin.systemd:
name: "{{ item }}"
enabled: true
state: started
loop:
- apt-daily.timer
- apt-daily-upgrade.timer
tags: [labelprint, base, updates]
+136
View File
@@ -0,0 +1,136 @@
---
- name: configure cupsd
become: true
ansible.builtin.template:
src: cupsd.conf.j2
dest: /etc/cups/cupsd.conf
owner: root
group: lp
mode: "0640"
validate: /usr/sbin/cupsd -t -c %s
notify: restart cups
tags: [labelprint, cups]
- name: enable cups
become: true
ansible.builtin.systemd:
name: cups.service
enabled: true
state: started
tags: [labelprint, cups]
# cupsd.conf has to be in place and cupsd running before lpadmin can talk to it.
- name: apply pending cups changes before touching the queue
ansible.builtin.meta: flush_handlers
tags: [labelprint, cups]
# ---------------------------------------------------------------------------
# The queue
# ---------------------------------------------------------------------------
# lpadmin is not idempotent and has no "show me everything you would set" mode,
# so the desired definition is fingerprinted and the fingerprint compared with
# what was last applied. The marker is written only after lpadmin succeeds.
# Checksummed here rather than taken from the install task's return value, so
# that the fingerprint is the same whether or not this run included the driver
# tasks -- `make deploy-labelprint TAGS=cups` must not look like a change.
- name: checksum the installed PPD
become: true
ansible.builtin.stat:
path: "{{ labelprint_ppd_active }}"
checksum_algorithm: sha1
register: labelprint_ppd_stat
tags: [labelprint, cups]
- name: build the desired queue fingerprint
ansible.builtin.set_fact:
labelprint_queue_want: >-
{{
[
labelprint_device_uri,
labelprint_queue_info,
labelprint_queue_location,
labelprint_resolution,
labelprint_media,
labelprint_darkness | string,
labelprint_print_speed | string,
labelprint_ppd_stat.stat.checksum | default('none'),
] | join('|')
}}
tags: [labelprint, cups]
- name: read the queue fingerprint that was last applied
become: true
ansible.builtin.slurp:
src: "/etc/cups/.{{ labelprint_queue }}.fingerprint"
register: labelprint_queue_have
failed_when: false
tags: [labelprint, cups]
- name: create or update the label queue
become: true
ansible.builtin.command:
argv:
- lpadmin
- -p
- "{{ labelprint_queue }}"
- -E
- -v
- "{{ labelprint_device_uri }}"
- -P
- "{{ labelprint_ppd_active }}"
- -D
- "{{ labelprint_queue_info }}"
- -L
- "{{ labelprint_queue_location }}"
- -o
- printer-is-shared=true
- -o
- "Resolution={{ labelprint_resolution }}"
- -o
- "media={{ labelprint_media }}"
- -o
- "Darkness={{ labelprint_darkness }}"
- -o
- "PrintSpeed={{ labelprint_print_speed }}"
when: >-
(labelprint_queue_have.content | default('') | b64decode | trim)
!= labelprint_queue_want | trim
register: labelprint_lpadmin
changed_when: true
tags: [labelprint, cups]
- name: record the applied queue fingerprint
become: true
ansible.builtin.copy:
dest: "/etc/cups/.{{ labelprint_queue }}.fingerprint"
content: "{{ labelprint_queue_want | trim }}"
owner: root
group: root
mode: "0600"
when: labelprint_lpadmin is changed
tags: [labelprint, cups]
- name: accept and enable the label queue
become: true
ansible.builtin.command:
argv: ["{{ item }}", "{{ labelprint_queue }}"]
loop:
- cupsaccept
- cupsenable
changed_when: false
tags: [labelprint, cups]
# There is deliberately no `cupsctl` here. It is the obvious way to say
# "share on the LAN only", but cupsctl edits cupsd.conf through cupsd itself,
# which rewrites the file from its parsed form and drops every comment. That
# makes the template above differ on the next run, which re-templates and
# restarts cups, which lets cupsctl rewrite it again -- a deploy that reports
# changes forever and never converges.
#
# Nothing is lost. `cupsctl --share-printers` amounts to `Browsing On` plus a
# per-queue shared flag, and both are already set -- the first in the template,
# the second by lpadmin's printer-is-shared=true above. `--no-remote-any` is the
# absence of `Allow from all` in <Location />, which is how the template is
# written. The driver's own install.sh runs `cupsctl --remote-any`, which would
# offer this printer to anything that can route to the Pi; that is exactly what
# we are not doing.
+147
View File
@@ -0,0 +1,147 @@
---
# Builds RunTheWall/tspl-cups-driver (MIT) from a pinned commit. The build is
# three files -- a CUPS raster filter, a backend and a PPD -- so it is cheap to
# do on the Pi itself and avoids trusting a prebuilt binary.
- name: fetch the tspl driver source
become: true
ansible.builtin.git:
repo: "{{ labelprint_driver_repo }}"
dest: "{{ labelprint_driver_src }}"
version: "{{ labelprint_driver_version }}"
force: true
register: labelprint_driver_checkout
tags: [labelprint, driver]
- name: check whether the filter is already built
become: true
ansible.builtin.stat:
path: /usr/lib/cups/filter/rastertotspl
register: labelprint_filter
tags: [labelprint, driver]
# `make` alone is not idempotent enough to report honestly -- it prints a
# recipe line on a rebuild and nothing on a no-op -- so the decision to build is
# made from the checkout state instead.
- name: build the tspl raster filter
become: true
community.general.make:
chdir: "{{ labelprint_driver_src }}"
when: labelprint_driver_checkout.changed or not labelprint_filter.stat.exists
tags: [labelprint, driver]
- name: install the tspl raster filter
become: true
ansible.builtin.copy:
src: "{{ labelprint_driver_src }}/src/rastertotspl"
dest: /usr/lib/cups/filter/rastertotspl
remote_src: true
owner: root
group: root
mode: "0755"
notify: restart cups
tags: [labelprint, driver]
# 0700 and root-owned on purpose: cupsd refuses to run a backend that is group-
# or world-writable, and runs it as an unprivileged user if it is not 0700.
# Writing to the printer's usblp node needs the privileged path.
- name: install the tspl backend
become: true
ansible.builtin.copy:
src: "{{ labelprint_driver_src }}/backend/tspl"
dest: /usr/lib/cups/backend/tspl
remote_src: true
owner: root
group: root
mode: "0700"
notify: restart cups
tags: [labelprint, driver]
- name: create the PPD directory
become: true
ansible.builtin.file:
path: "{{ labelprint_ppd_dir }}"
state: directory
owner: root
group: root
mode: "0755"
tags: [labelprint, driver]
- name: install the tspl PPD
become: true
ansible.builtin.copy:
src: "{{ labelprint_driver_src }}/ppd/tspl-label.ppd"
dest: "{{ labelprint_ppd_dir }}/tspl-label.ppd"
remote_src: true
owner: root
group: root
mode: "0644"
tags: [labelprint, driver]
# ---------------------------------------------------------------------------
# The 4x6-only PPD
# ---------------------------------------------------------------------------
# Regenerated from the driver's PPD every run rather than kept as a fork, so a
# driver bump brings its new PPD through the same filter.
- name: create the helper directory
become: true
ansible.builtin.file:
path: /usr/local/share/labelprint
state: directory
owner: root
group: root
mode: "0755"
when: labelprint_media_only_4x6
tags: [labelprint, driver]
- name: install the PPD media trim filter
become: true
ansible.builtin.copy:
src: trim-ppd-media.awk
dest: /usr/local/share/labelprint/trim-ppd-media.awk
owner: root
group: root
mode: "0644"
when: labelprint_media_only_4x6
tags: [labelprint, driver]
- name: render the 4x6-only PPD
become: true
ansible.builtin.command:
argv:
- awk
- -v
- "keep={{ labelprint_ppd_pagesize }}"
- -f
- /usr/local/share/labelprint/trim-ppd-media.awk
- "{{ labelprint_ppd_dir }}/tspl-label.ppd"
register: labelprint_ppd_trim
changed_when: false
when: labelprint_media_only_4x6
tags: [labelprint, driver]
- name: install the 4x6-only PPD
become: true
ansible.builtin.copy:
content: "{{ labelprint_ppd_trim.stdout }}\n"
dest: "{{ labelprint_ppd_dir }}/tspl-label-4x6.ppd"
owner: root
group: root
mode: "0644"
validate: cupstestppd -q %s
when: labelprint_media_only_4x6
tags: [labelprint, driver]
# Gives the printer a stable /dev/usb/tspl-label symlink across USB
# re-enumeration. Harmless if the PM246's id is not in the shipped rules -- the
# backend still finds it by walking /dev/usb/lp*.
- name: install the tspl udev rules
become: true
ansible.builtin.copy:
src: "{{ labelprint_driver_src }}/udev/99-tspl-label.rules"
dest: /etc/udev/rules.d/99-tspl-label.rules
remote_src: true
owner: root
group: root
mode: "0644"
notify: reload udev rules
tags: [labelprint, driver]
@@ -0,0 +1,20 @@
---
- name: install the nftables ruleset
become: true
ansible.builtin.template:
src: nftables.conf.j2
dest: /etc/nftables.conf
owner: root
group: root
mode: "0755"
validate: /usr/sbin/nft -c -f %s
notify: reload nftables
tags: [labelprint, firewall]
- name: enable nftables
become: true
ansible.builtin.systemd:
name: nftables.service
enabled: true
state: started
tags: [labelprint, firewall]
+6
View File
@@ -0,0 +1,6 @@
---
- import_tasks: base.yml
- import_tasks: driver.yml
- import_tasks: cups.yml
- import_tasks: wifi.yml
- import_tasks: firewall.yml
+133
View File
@@ -0,0 +1,133 @@
---
# WPA2-PSK takes an 8-63 character passphrase, or exactly 64 hex characters as a
# precomputed key. NetworkManager stores a shorter one without complaint --
# psk-flags stays 0 and the value sits in the keyfile -- and then refuses to
# activate with "Secrets were required, but not provided", which reads like a
# missing password rather than an invalid one.
#
# For the rescue AP that failure is invisible until the day the home network is
# down and this is the only way in, so it is checked here instead. Only lengths
# are reported, never the values.
- name: check the Wi-Fi secrets are usable as WPA2-PSK
ansible.builtin.assert:
that:
- (vars[item] | length >= 8 and vars[item] | length <= 63)
or (vars[item] is match('^[0-9a-fA-F]{64}$'))
fail_msg: >-
{{ item }} is {{ vars[item] | length }} characters, which WPA2 will not
accept. Use an 8-63 character passphrase, or a 64-character hex
precomputed key. Fix it with `make vault`.
quiet: true
# The loop carries the variable NAME, never its value: a failed assert prints
# the item it was iterating over, so looping over the secrets themselves would
# dump both passwords to the terminal on any failure.
loop:
- stickah_psk
- stickah_psk_rescue
tags: [labelprint, wifi]
# The watchdog can pull the radio out from under this very play if it decides
# the Pi is offline while we are mid-deploy over the rescue AP. The hold expires
# on its own after {{ labelprint_hold_max_age_secs }}s, so an aborted run cannot
# leave the watchdog disabled.
- name: hold the radio for the duration of this deploy
become: true
ansible.builtin.file:
path: /run/wifi-rescue.hold
state: touch
owner: root
group: root
mode: "0644"
changed_when: false
tags: [labelprint, wifi]
- name: install the home Wi-Fi profile
become: true
ansible.builtin.template:
src: home-wifi.nmconnection.j2
dest: /etc/NetworkManager/system-connections/home-wifi.nmconnection
owner: root
group: root
mode: "0600"
notify: reload networkmanager connections
tags: [labelprint, wifi]
- name: install the rescue access point profile
become: true
ansible.builtin.template:
src: rescue-ap.nmconnection.j2
dest: /etc/NetworkManager/system-connections/rescue-ap.nmconnection
owner: root
group: root
mode: "0600"
notify: reload networkmanager connections
tags: [labelprint, wifi]
# ---------------------------------------------------------------------------
# Coexisting with netplan
# ---------------------------------------------------------------------------
# The hand-built image configures Wi-Fi through cloud-init's network-config, and
# on Raspberry Pi OS trixie netplan's NetworkManager integration turns that into
# a persistent profile of its own at /etc/netplan/90-NM-<uuid>.yaml -- not the
# /etc/netplan/50-cloud-init.yaml you would expect, and not something a
# cloud-init clean removes.
#
# That profile is deliberately left in place. It carries the same SSID as
# home-wifi, and autoconnect-priority decides between them: 100 here against
# netplan's 0, so NetworkManager picks ours every time. What netplan's copy buys
# is a fallback that predates anything in this role -- if home-wifi is ever
# rendered wrong, the Pi still comes back on the LAN instead of stranding itself
# on the rescue AP. Its PSK goes stale when the home password changes; that
# costs nothing, because a stale profile simply fails and ours is tried first.
#
# What is worth stopping is cloud-init rewriting the network on a future
# re-instance, which would put a third opinion in play.
- name: stop cloud-init from rewriting the network config
become: true
ansible.builtin.copy:
dest: /etc/cloud/cloud.cfg.d/99-disable-network-config.cfg
content: |
network: {config: disabled}
owner: root
group: root
mode: "0644"
tags: [labelprint, wifi]
# ---------------------------------------------------------------------------
# The watchdog
# ---------------------------------------------------------------------------
- name: install the Wi-Fi rescue watchdog
become: true
ansible.builtin.template:
src: wifi-rescue.sh.j2
dest: /usr/local/sbin/wifi-rescue
owner: root
group: root
mode: "0755"
tags: [labelprint, wifi]
- name: install the Wi-Fi rescue units
become: true
ansible.builtin.template:
src: "{{ item }}.j2"
dest: "/etc/systemd/system/{{ item }}"
owner: root
group: root
mode: "0644"
loop:
- wifi-rescue.service
- wifi-rescue.timer
notify: reload systemd
tags: [labelprint, wifi]
- name: apply pending unit changes
ansible.builtin.meta: flush_handlers
tags: [labelprint, wifi]
- name: enable the Wi-Fi rescue watchdog
become: true
ansible.builtin.systemd:
name: wifi-rescue.timer
enabled: true
state: started
tags: [labelprint, wifi]
@@ -0,0 +1,11 @@
# {{ ansible_managed }} -- rendered by `make bootfs`
#
# cloud-init applies user-data once per INSTANCE, not once per boot: rewrite
# user-data on a card that has already booted and nothing happens, because
# cloud-init recognises the instance id and skips straight to per-boot modules.
#
# Deriving the id from a hash of user-data itself fixes that. Re-running
# `make bootfs` with no changes leaves the id alone, so a card keeps its
# identity; change anything in user-data and the id changes with it, and the
# Pi re-applies the new config on its next boot without a reflash.
instance-id: {{ labelprint_hostname }}-{{ labelprint_user_data_id }}
@@ -0,0 +1,18 @@
# {{ ansible_managed }} -- rendered by `make bootfs`
#
# Ethernet only, on purpose. Wi-Fi is NOT configured here: cloud-init renders
# this through netplan into a persistent profile of netplan's own, which is not
# a file this repo manages or can readily update. The Wi-Fi profiles are written
# straight into /etc/NetworkManager/system-connections by user-data instead, so
# the image and Ansible manage the same files.
#
# A card that has already booted with Wi-Fi in network-config keeps netplan's
# profile. That is fine, and useful -- see the "Coexisting with netplan" note in
# roles/labelprint/tasks/wifi.yml.
network:
version: 2
ethernets:
eth0:
dhcp4: true
dhcp6: true
optional: true
@@ -0,0 +1,125 @@
#cloud-config
# {{ ansible_managed }} -- rendered by `make bootfs BOOTFS=...`
#
# First boot of the label print proxy. The job of this file is to make the Pi
# REACHABLE and nothing more: a login that works, a network that works, and a
# rescue access point for when it does not. The printer driver and the CUPS
# queue are Ansible's job -- see roles/labelprint/.
#
# The Wi-Fi profiles and the rescue watchdog below are rendered from the very
# same templates the role uses, so the first `make deploy-labelprint` reports no
# change on any of them. That is the point: the Pi can already rescue itself
# before Ansible has ever run.
#
# Applies on FIRST BOOT ONLY. Rewriting this file on a card that has already
# booted does nothing -- reflash the image.
hostname: {{ labelprint_hostname }}
manage_etc_hosts: true
manage_resolv_conf: false
timezone: America/New_York
keyboard:
model: pc105
layout: "us"
apt:
preserve_sources_list: true
users:
- name: {{ labelprint_user }}
shell: /bin/bash
lock_passwd: false
# Deliberately trivial, and deliberately not hashed: this password is
# documented in roles/labelprint/README.md, so a hash of it would protect
# nothing while pretending otherwise. It exists so that a Pi which will not
# take the SSH key -- wrong key, reflashed card, someone else's laptop -- is
# still reachable from the LAN or the rescue AP, which is the only place it
# can be reached from at all (see templates/nftables.conf.j2).
plain_text_passwd: {{ labelprint_user_password }}
sudo: "ALL=(ALL) NOPASSWD:ALL"
groups: [sudo, adm, lpadmin, plugdev, dialout, netdev, users]
ssh_authorized_keys:
- {{ lookup('file', labelprint_authorized_key_file) }}
# The key is what Ansible actually uses; the password is the fallback.
ssh_pwauth: true
disable_root: true
chpasswd:
expire: false
# Pre-installing what the role needs makes the first deploy quick, and means a
# Pi that comes up on the rescue AP with no internet still has cups and
# NetworkManager. Ansible installs the same list, so nothing here is load-bearing.
package_update: true
packages:
{% for pkg in labelprint_deps %}
- {{ pkg }}
{% endfor %}
write_files:
# Stop cloud-init rewriting the network on a future re-instance.
- path: /etc/cloud/cloud.cfg.d/99-disable-network-config.cfg
owner: root:root
permissions: '0644'
content: |
network: {config: disabled}
# Belt and braces. netplan can hand NetworkManager an "only manage what I
# listed" policy, and what this image lists is eth0; a stock Raspberry Pi OS
# leaves that policy empty, but an unmanaged wlan0 would take the rescue AP
# down with it, so it is not worth depending on.
- path: /etc/NetworkManager/conf.d/10-labelprint.conf
owner: root:root
permissions: '0644'
content: |
[keyfile]
unmanaged-devices=none
[device]
# A randomised MAC would give the Pi a different DHCP lease on every
# association, which makes it hard to find on the UDM Pro's client list.
wifi.scan-rand-mac-address=no
- path: /etc/NetworkManager/system-connections/home-wifi.nmconnection
owner: root:root
permissions: '0600'
content: |
{{ lookup('template', 'roles/labelprint/templates/home-wifi.nmconnection.j2') | indent(6) }}
- path: /etc/NetworkManager/system-connections/rescue-ap.nmconnection
owner: root:root
permissions: '0600'
content: |
{{ lookup('template', 'roles/labelprint/templates/rescue-ap.nmconnection.j2') | indent(6) }}
- path: /usr/local/sbin/wifi-rescue
owner: root:root
permissions: '0755'
content: |
{{ lookup('template', 'roles/labelprint/templates/wifi-rescue.sh.j2') | indent(6) }}
- path: /etc/systemd/system/wifi-rescue.service
owner: root:root
permissions: '0644'
content: |
{{ lookup('template', 'roles/labelprint/templates/wifi-rescue.service.j2') | indent(6) }}
- path: /etc/systemd/system/wifi-rescue.timer
owner: root:root
permissions: '0644'
content: |
{{ lookup('template', 'roles/labelprint/templates/wifi-rescue.timer.j2') | indent(6) }}
runcmd:
# RPi OS soft-blocks the radio until a regulatory domain is known. cmdline.txt
# carries cfg80211.ieee80211_regdom=US, but unblock anyway -- a blocked radio
# is the one failure that takes the rescue AP down with it.
- [rfkill, unblock, wifi]
- [systemctl, enable, --now, ssh]
- [systemctl, enable, --now, avahi-daemon]
# Let NetworkManager pick up the keyfiles written above. The profile netplan
# renders from network-config is left alone on purpose -- see the "Coexisting
# with netplan" note in roles/labelprint/tasks/wifi.yml.
- [nmcli, connection, reload]
- [systemctl, daemon-reload]
- [systemctl, enable, --now, wifi-rescue.timer]
@@ -0,0 +1,97 @@
# {{ ansible_managed }}
#
# CUPS on the label print proxy. The Pi does all the rendering, so macOS,
# Windows and Linux clients add the queue driverless over IPP Everywhere /
# AirPrint and never install a Phomemo driver.
#
# Reachable from the LAN and from the rescue access point only. Everything else
# is refused here and dropped again at nftables.
LogLevel warn
PageLogFormat
MaxLogSize 1m
# A label job is worth retrying: the printer is often powered off or out of
# stock when the job is submitted, and the default is to bin the job outright.
ErrorPolicy retry-job
# Only trusted, local networks reach this port -- see the Location blocks below
# and roles/labelprint/templates/nftables.conf.j2. Listening on all interfaces
# rather than a fixed address so the queue is still reachable at
# {{ labelprint_ap_addr }} when the Pi has fallen back to its rescue AP.
Listen 631
Listen /run/cups/cups.sock
# Advertise over Bonjour/mDNS so clients discover the queue by themselves.
Browsing On
BrowseLocalProtocols dnssd
# Dropping the _cups subtype is what makes macOS and iOS offer a driverless
# "AirPrint" add instead of guessing at a Generic PostScript driver.
BrowseDNSSDSubTypes _print,_universal
DefaultAuthType Basic
WebInterface Yes
<Location />
Order allow,deny
Allow from {{ labelprint_lan_cidr }}
Allow from {{ labelprint_ap_cidr }}
Allow from localhost
</Location>
<Location /admin>
AuthType Default
Require user @SYSTEM
Order allow,deny
Allow from {{ labelprint_lan_cidr }}
Allow from {{ labelprint_ap_cidr }}
</Location>
<Location /admin/conf>
AuthType Default
Require user @SYSTEM
Order allow,deny
Allow from {{ labelprint_lan_cidr }}
Allow from {{ labelprint_ap_cidr }}
</Location>
<Location /admin/log>
AuthType Default
Require user @SYSTEM
Order allow,deny
Allow from {{ labelprint_lan_cidr }}
Allow from {{ labelprint_ap_cidr }}
</Location>
<Policy default>
JobPrivateAccess default
JobPrivateValues default
SubscriptionPrivateAccess default
SubscriptionPrivateValues default
# Anyone on the LAN may print and manage their own jobs -- this is a label
# printer in a workshop, not a shared office device with quotas.
<Limit Create-Job Print-Job Print-URI Validate-Job>
Order deny,allow
</Limit>
<Limit Send-Document Send-URI Hold-Job Release-Job Restart-Job Purge-Jobs Set-Job-Attributes Create-Job-Subscription Renew-Subscription Cancel-Subscription Get-Notifications Reprocess-Job Cancel-Current-Job Suspend-Current-Job Resume-Job Cancel-Jobs CUPS-Authenticate-Job Close-Job CUPS-Move-Job Cancel-My-Jobs CUPS-Get-Document>
Order deny,allow
</Limit>
# Changing the printer itself needs a local admin.
<Limit Pause-Printer Resume-Printer Enable-Printer Disable-Printer Pause-Printer-After-Current-Job Hold-New-Jobs Release-Held-New-Jobs Deactivate-Printer Activate-Printer Restart-Printer Shutdown-Printer Startup-Printer Promote-Job Schedule-Job-After Cancel-Job CUPS-Accept-Jobs CUPS-Reject-Jobs>
AuthType Default
Require user @SYSTEM
Order deny,allow
</Limit>
<Limit CUPS-Add-Modify-Printer CUPS-Delete-Printer CUPS-Add-Modify-Class CUPS-Delete-Class CUPS-Set-Default>
AuthType Default
Require user @SYSTEM
Order deny,allow
</Limit>
<Limit All>
Order deny,allow
</Limit>
</Policy>
@@ -0,0 +1,32 @@
# {{ ansible_managed }}
#
# The normal, everyday Wi-Fi link. autoconnect-priority outranks anything
# cloud-init/netplan rendered for the same SSID, so this is the profile
# NetworkManager picks. autoconnect-retries=0 means retry forever rather than
# giving up after four attempts and leaving the Pi off the network until
# someone power-cycles it -- the rescue AP is the fallback, not the
# destination.
[connection]
id=home-wifi
uuid={{ (stickah_ssid ~ '-home-wifi') | to_uuid }}
type=wifi
interface-name=wlan0
autoconnect=true
autoconnect-priority=100
autoconnect-retries=0
[wifi]
mode=infrastructure
ssid={{ stickah_ssid }}
[wifi-security]
key-mgmt=wpa-psk
# Accepts either a passphrase or the 64-hex precomputed PSK.
psk={{ stickah_psk }}
[ipv4]
method=auto
[ipv6]
method=auto
addr-gen-mode=default
@@ -0,0 +1,63 @@
#!/usr/sbin/nft -f
# {{ ansible_managed }}
#
# The print proxy answers to the LAN and to its own rescue access point, and to
# nothing else. This is the outer half of the same rule that cupsd enforces in
# its Location blocks -- both are here on purpose, so a mistake in one is not
# the only thing standing between the printer and the rest of the world.
# Declare-then-delete rather than `flush ruleset`: NetworkManager's shared mode
# keeps its own table for the rescue AP's dnsmasq, and a global flush would take
# that with it every time this file is reloaded.
table inet labelprint
delete table inet labelprint
table inet labelprint {
set trusted {
type ipv4_addr
flags interval
elements = { {{ labelprint_lan_cidr }}, {{ labelprint_ap_cidr }} }
}
chain input {
type filter hook input priority filter; policy drop;
ct state established,related accept
ct state invalid drop
iif lo accept
icmp type { echo-request, destination-unreachable, time-exceeded, parameter-problem } accept
icmpv6 type { echo-request, destination-unreachable, packet-too-big, time-exceeded, parameter-problem, nd-neighbor-solicit, nd-neighbor-advert, nd-router-advert } accept
# DHCP replies to our own client. Broadcast, so conntrack does not see
# them as related to the request we sent.
udp dport 68 accept
# ssh and IPP, from the LAN or from a machine on the rescue AP.
ip saddr @trusted tcp dport { 22, 631 } accept
ip saddr @trusted udp dport 631 accept
# mDNS: how every client finds this printer, since it has no DNS record.
ip saddr @trusted udp dport 5353 accept
# DHCP for whoever joins the rescue AP. Deliberately not restricted by
# source address: a client asking for its first lease has no address
# yet and sends DHCPDISCOVER from 0.0.0.0, so a source-matched rule
# would mean the rescue network never hands out a lease at all. Only a
# machine already associated to our own AP can reach this port.
iifname "wlan0" udp dport 67 accept
# DNS, once they have an address.
iifname "wlan0" ip saddr {{ labelprint_ap_cidr }} udp dport 53 accept
iifname "wlan0" ip saddr {{ labelprint_ap_cidr }} tcp dport 53 accept
}
# The rescue AP is a way in to this Pi, not a route to anywhere else.
chain forward {
type filter hook forward priority filter; policy drop;
}
chain output {
type filter hook output priority filter; policy accept;
}
}
@@ -0,0 +1,38 @@
# {{ ansible_managed }}
#
# Rescue access point. Never autoconnects -- wifi-rescue brings it up only when
# the home SSID is unreachable, so that a changed password or a dead router
# leaves a way back in to reconfigure the Pi.
#
# Join this SSID and the Pi is at {{ labelprint_ap_addr }}: ssh pi@{{ labelprint_ap_addr }},
# or print to it directly at ipp://{{ labelprint_ap_addr }}:631/printers/{{ labelprint_queue }}.
[connection]
id=rescue-ap
uuid={{ (stickah_ssid_rescue ~ '-rescue-ap') | to_uuid }}
type=wifi
interface-name=wlan0
autoconnect=false
[wifi]
mode=ap
ssid={{ stickah_ssid_rescue }}
# The 3B+ radio does 5 GHz, but 2.4 GHz AP mode is what brcmfmac is reliable
# at, and a rescue network only has to carry an SSH session.
band=bg
channel=6
[wifi-security]
key-mgmt=wpa-psk
proto=rsn
pairwise=ccmp
group=ccmp
psk={{ stickah_psk_rescue }}
# "shared" makes NetworkManager run a dnsmasq for DHCP and DNS on this
# interface, so a laptop that joins gets an address without any further setup.
[ipv4]
method=shared
address1={{ labelprint_ap_addr }}/24
[ipv6]
method=ignore
@@ -0,0 +1,9 @@
# {{ ansible_managed }}
[Unit]
Description=Wi-Fi rescue access point watchdog
After=NetworkManager.service
Requires=NetworkManager.service
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/wifi-rescue
@@ -0,0 +1,132 @@
#!/bin/sh
# {{ ansible_managed }}
#
# Keeps the label print proxy reachable.
#
# Normally the Pi is a station on the home SSID. If that network is gone -- the
# password changed, the AP died, the Pi was carried somewhere else -- it brings
# up its own rescue access point so there is still a way in to reconfigure it.
# It keeps checking, and hands the radio back the moment the home SSID returns.
#
# The 3B+ has one radio and brcmfmac will not hold an AP and a station link at
# the same time, so this cannot listen for the home SSID while the AP is up.
# Instead the AP drops for a few seconds every {{ labelprint_ap_rescan_secs }}s
# to scan, and comes straight back if the home network is still missing.
#
# If you are working over the rescue AP and do not want the radio pulled out
# from under your SSH session:
#
# touch /run/wifi-rescue.hold
#
# The hold expires by itself after {{ (labelprint_hold_max_age_secs / 60) | int }} minutes, so a forgotten hold file
# cannot strand the Pi permanently.
set -eu
HOME_CON=home-wifi
AP_CON=rescue-ap
SSID='{{ stickah_ssid }}'
AP_SSID='{{ stickah_ssid_rescue }}'
IFACE=wlan0
HOLD=/run/wifi-rescue.hold
STAMP=/run/wifi-rescue.ap-since
RESCAN_SECS={{ labelprint_ap_rescan_secs }}
HOLD_MAX_AGE={{ labelprint_hold_max_age_secs }}
log() { logger -t wifi-rescue -- "$@"; }
con_active() { nmcli -t -f NAME connection show --active | grep -qxF "$1"; }
# A default IPv4 route is the honest test for "on a real network". The rescue
# AP uses NetworkManager's shared mode, which hands out addresses but installs
# no default route, so the AP can never make this look true.
online() { [ -n "$(ip -4 route show default 2>/dev/null)" ]; }
held() {
[ -e "$HOLD" ] || return 1
age=$(( $(date +%s) - $(stat -c %Y "$HOLD" 2>/dev/null || echo 0) ))
if [ "$age" -lt "$HOLD_MAX_AGE" ]; then
return 0
fi
log "hold file is ${age}s old; expiring it"
rm -f "$HOLD"
return 1
}
start_ap() {
con_active "$AP_CON" && return 0
log "starting rescue AP '$AP_SSID' on {{ labelprint_ap_addr }}"
if nmcli --wait 20 connection up "$AP_CON" >/dev/null 2>&1; then
date +%s > "$STAMP"
else
log "ERROR: rescue AP failed to start"
fi
}
join_home() {
# Bounded: the default 90s wait outlives the watchdog interval, and an
# SSID that is not there is not going to appear in the next minute.
#
# --wait is a GLOBAL nmcli option and has to precede the subcommand. Written
# as `connection up <id> --wait 20` it is rejected outright with "invalid
# extra argument" -- and only for a connection that exists, so it looks fine
# against a typo'd name. That failure mode is silent and total: every join
# returns failure and the rescue AP never starts.
nmcli --wait 20 connection up "$HOME_CON" >/dev/null 2>&1 || return 1
online
}
held && exit 0
if online; then
if con_active "$AP_CON"; then
log "back on the network; shutting the rescue AP down"
nmcli connection down "$AP_CON" >/dev/null 2>&1 || true
rm -f "$STAMP"
fi
exit 0
fi
if con_active "$AP_CON"; then
since=$(cat "$STAMP" 2>/dev/null || echo 0)
[ $(( $(date +%s) - since )) -lt "$RESCAN_SECS" ] && exit 0
log "rescue AP up for ${RESCAN_SECS}s; dropping it to scan for '$SSID'"
nmcli connection down "$AP_CON" >/dev/null 2>&1 || true
sleep 2
nmcli device wifi rescan ifname "$IFACE" >/dev/null 2>&1 || true
sleep 5
scan=$(nmcli -t -f SSID device wifi list ifname "$IFACE" 2>/dev/null || true)
if printf '%s\n' "$scan" | grep -qxF "$SSID"; then
log "'$SSID' is back; rejoining"
try=yes
elif [ -z "$(printf '%s' "$scan" | tr -d '[:space:]')" ]; then
# Nothing at all came back. Either the radio has not settled after
# dropping the AP, or the home SSID is hidden and will never show up in
# a scan. Worth 20 seconds to find out.
log "scan came back empty; trying '$SSID' anyway"
try=yes
else
# Other networks are visible and ours is not, so it really is gone.
# Straight back to the AP -- no point spending the join timeout.
try=no
fi
if [ "$try" = yes ]; then
if join_home; then
log "rejoined '$SSID'"
rm -f "$STAMP"
exit 0
fi
log "join failed; returning to the rescue AP"
fi
start_ap
exit 0
fi
log "offline and no rescue AP; trying '$SSID'"
if join_home; then
log "joined '$SSID'"
exit 0
fi
start_ap
@@ -0,0 +1,13 @@
# {{ ansible_managed }}
[Unit]
Description=Run the Wi-Fi rescue watchdog every {{ labelprint_watchdog_interval_secs }}s
[Timer]
# Waits for NetworkManager to have had a fair go at the home SSID before the
# first check, so a slow DHCP lease at boot does not trip the rescue AP.
OnBootSec=90
OnUnitActiveSec={{ labelprint_watchdog_interval_secs }}
AccuracySec=5
[Install]
WantedBy=timers.target
+26 -1
View File
@@ -27,6 +27,12 @@ hass_path: "{{ podman_volumes }}/hass"
partsy_path: "{{ podman_volumes }}/partsy"
partsy_skudak_path: "{{ podman_volumes }}/partsy-skudak"
photos_path: "{{ podman_volumes }}/photos"
# rsvp.debyl.io -- invite-only RSVP app (~/src/rsvp-debylio). Both listeners are
# published on loopback only: Caddy proxies the public one, and the admin one
# only for LAN clients. Caddy runs with network: host, so 127.0.0.1 reaches them.
rsvp_path: "{{ podman_volumes }}/rsvp"
rsvp_public_port: 9080
rsvp_admin_port: 9081
# Named rather than hardcoded so the ML service can be stood up beside a broken
# one without editing tasks. On 2026-08-28 a worker thread wedged in
# uninterruptible sleep (D state) in exit_mmap, which made the container
@@ -80,7 +86,12 @@ zomboid_max_ram: 24g
# The modlist itself lives in vars/zomboid_sophie_mods.yml -- 283 mod IDs and
# 256 workshop items vendored from the upstream Sophie 42 preset.
zomboid_preset_version: "sophie-42 @ 2026-07-31"
#
# The modlist is still the upstream preset. The world settings no longer are:
# files/zomboid/sophie/SandboxVars.lua is a snapshot of the running debbzoid
# world taken 2026-08-31, not the preset as shipped. See the seed tasks in
# containers/home/zomboid.yml.
zomboid_preset_version: "sophie-42 mods @ 2026-07-31 / live world settings @ 2026-08-31"
# Mods subtracted from the vendored Sophie lists before they reach the INI.
# Explicit and reasoned so that re-vendoring a newer preset cannot silently
@@ -139,6 +150,19 @@ zomboid_mods_renamed:
# between world regenerations. Set true to overwrite it from the templates.
zomboid_config_force: false
# Same switch, narrowed to spawnregions.lua and spawnpoints.lua. They are the
# only config a running world re-reads (at every server start, unlike
# SandboxVars, which is read only when the world is created), so a spawn edit
# can be deployed to the live world without a wipe. It is separate from
# zomboid_config_force so that landing one is not an excuse to force the INI and
# SandboxVars over the top of whatever the admins have set in-game.
#
# make deploy TAGS=zomboid-conf -e zomboid_spawn_force=true
#
# The server still has to be restarted, with players offline, before the new
# spawn config is read.
zomboid_spawn_force: false
sshpass_cron_path: "{{ podman_volumes }}/sshpass_cron"
caddy_path: "{{ podman_volumes }}/caddy"
@@ -167,6 +191,7 @@ home_server_name_io: home.debyl.io
parts_server_name_io: parts.debyl.io
photos_server_name_io: photos.debyl.io
gitea_debyl_server_name: git.debyl.io
rsvp_server_name: rsvp.debyl.io
# skudak.com domains (migration from skudakrennsport.com)
parts_skudak_server_name: parts.skudak.com
File diff suppressed because it is too large Load Diff
@@ -1,3 +1,28 @@
-- Deployed as Server/debbzoid_spawnpoints.lua, and currently unused: nothing
-- reads it, because no entry in spawnregions.lua names it with
-- serverfile = "debbzoid_spawnpoints.lua". The contents below are the stock
-- sample and mean nothing until such an entry exists.
--
-- This is the file to use if a spawn point has to move rather than disappear.
-- Point a region at it by serverfile instead of file, and put that region's
-- whole spawn table here -- the serverfile replaces the map's list, it does not
-- merge with it, so anything left out is gone.
--
-- Keeping "Echo Creek, KY" as a spawn choice while freeing the diner would look
-- like this, with an outdoor square -- the parking lot or the road, not another
-- building, or that building becomes the unclaimable one instead:
--
-- unemployed = {
-- { posX = <outdoor X>, posY = <outdoor Y>, posZ = 0 },
-- }
--
-- Coordinates are absolute world tiles when worldX/worldY are omitted; with
-- them, the world position is worldX * 300 + posX (same for Y).
--
-- The filename is not free-form: serverfile resolves inside the server's
-- Server/ directory, and the deploy names this copy after zomboid_server_name.
-- Rename the server and the serverfile reference has to follow, or the region
-- loses its points and drops out of the list.
function SpawnPoints()
return {
unemployed = {
@@ -1,7 +1,35 @@
-- Spawn regions offered to new characters, deployed as
-- Server/debbzoid_spawnregions.lua. PZ reads it at every server start, so a
-- change here takes effect on the live world after a restart -- no wipe.
--
-- Each entry either points at a map's own list (file = ...) or at a file in the
-- server's Server/ directory (serverfile = ...). Both are loaded by
-- media/lua/shared/SpawnRegions.lua; a region whose file is missing logs
-- "spawn points may be broken" to the server console and is dropped from the
-- list, taking its spawn choices with it.
--
-- This list is also what makes buildings unclaimable. SafeHouse.canBeSafehouse
-- walks every point of every region here, resolves each to a grid square, and
-- refuses the claim with "Spawn location. Cannot be claimed." if the square
-- being claimed is inside that square's building. It does not care which
-- profession the point belongs to, whether the region is reachable, or whether
-- the player is an admin. A point on an outdoor square blocks nothing, because
-- an outdoor square has no building.
--
-- 2026-08-31: "Echo Creek, KY" is commented out below. Its spawn list holds
-- exactly one point -- posX 3573, posY 10899, unemployed only -- and that point
-- sits inside the Echo Creek gas station/diner, which is why the diner cannot
-- be claimed as a safehouse. There is no server option that lifts the
-- restriction, so removing the region is the only way to free the building
-- short of moving the point somewhere else (see spawnpoints.lua).
--
-- The cost of removing it: unemployed characters lose Echo Creek as a spawn
-- choice. Nothing else changes -- the town, its loot and its buildings stay in
-- the world, and no existing character is affected.
function SpawnRegions()
return {
{ name = "Brandenburg, KY", file = "media/maps/Brandenburg, KY/spawnpoints.lua" },
{ name = "Echo Creek, KY", file = "media/maps/Echo Creek, KY/spawnpoints.lua" },
-- { name = "Echo Creek, KY", file = "media/maps/Echo Creek, KY/spawnpoints.lua" },
{ name = "Ekron, KY", file = "media/maps/Ekron, KY/spawnpoints.lua" },
{ name = "Fallas Lake, KY", file = "media/maps/Fallas Lake, KY/spawnpoints.lua" },
{ name = "Irvington, KY", file = "media/maps/Irvington, KY/spawnpoints.lua" },
@@ -0,0 +1,51 @@
---
# The image runs as uid 10001 (scratch, no passwd file). Rootless podman maps
# container uid N to host uid subuid_start + N - 1, so own the directory as
# that host uid directly. Setting the podman user here and chowning back
# afterwards would flip ownership on every deploy and briefly lock the running
# app out of its own database directory.
- name: create rsvp host directory volumes
become: true
ansible.builtin.file:
path: "{{ item }}"
state: directory
owner: "{{ podman_subuid.stdout | int + 10000 }}"
group: "{{ podman_subuid.stdout | int + 10000 }}"
mode: 0750
notify: restorecon podman
loop:
- "{{ rsvp_path }}/data"
- name: flush handlers
ansible.builtin.meta: flush_handlers
- import_tasks: podman/podman-check.yml
vars:
container_name: rsvp
container_image: "{{ image }}"
- name: create rsvp container
become: true
become_user: "{{ podman_user }}"
containers.podman.podman_container:
name: rsvp
image: "{{ image }}"
restart_policy: on-failure:3
log_driver: journald
env:
RSVP_DB_PATH: /data/rsvp.db
RSVP_BASE_URL: "https://{{ rsvp_server_name }}"
RSVP_TZ: America/New_York
# Client IPs for the invite-link miss limiter come from Caddy's
# X-Forwarded-For. Safe only because both ports are loopback-only.
RSVP_TRUST_FORWARDED: "1"
volumes:
- "{{ rsvp_path }}/data:/data"
ports:
- "127.0.0.1:{{ rsvp_public_port }}:8080"
- "127.0.0.1:{{ rsvp_admin_port }}:8081"
- name: create systemd startup job for rsvp
include_tasks: podman/systemd-generate.yml
vars:
container_name: rsvp
@@ -34,6 +34,10 @@
loop:
- config-template
- config-backup
# Holds the world each restore displaces, so a restore is always undoable.
# Deliberately outside data/ -- that is the container's bind mount, and a
# stray world directory inside Saves/Multiplayer/ is something PZ would see.
- restore-backup
- name: create zomboid host-side log directory
become: true
@@ -82,6 +86,35 @@
mode: '0644'
notify: reload zomboid systemd
- name: deploy zomboid restore script
become: true
ansible.builtin.template:
src: zomboid/zomboid-restore.sh.j2
dest: "{{ podman_home }}/bin/zomboid-restore.sh"
owner: "{{ podman_user }}"
group: "{{ podman_user }}"
mode: '0755'
- name: deploy zomboid restore path unit
become: true
ansible.builtin.template:
src: zomboid/zomboid-restore.path.j2
dest: "{{ podman_home }}/.config/systemd/user/zomboid-restore.path"
owner: "{{ podman_user }}"
group: "{{ podman_user }}"
mode: '0644'
notify: reload zomboid systemd
- name: deploy zomboid restore service unit
become: true
ansible.builtin.template:
src: zomboid/zomboid-restore.service.j2
dest: "{{ podman_home }}/.config/systemd/user/zomboid-restore.service"
owner: "{{ podman_user }}"
group: "{{ podman_user }}"
mode: '0644'
notify: reload zomboid systemd
- name: deploy zomboid stats script
become: true
ansible.builtin.template:
@@ -152,8 +185,12 @@
mode: 0644
notify: restorecon podman
# Server config is seeded from the vendored Sophie 42 preset *before* first
# boot, so the server never generates a vanilla INI we then have to patch.
# Server config is seeded *before* first boot, so the server never generates a
# vanilla INI we then have to patch.
#
# The INI and SandboxVars started as the vendored Sophie 42 preset and were
# re-synced from the running debbzoid world on 2026-08-31 -- see the header of
# templates/zomboid/server.ini.j2 for the INI keys that moved and why.
#
# force is off by design: these files are a starting point, not a managed
# state. Stop the server, hand-edit them, regenerate the world -- Ansible will
@@ -170,7 +207,7 @@
notify: restorecon podman
tags: zomboid-conf
- name: seed zomboid server ini from sophie preset
- name: seed zomboid server ini
become: true
ansible.builtin.template:
src: zomboid/server.ini.j2
@@ -183,17 +220,42 @@
tags: zomboid-conf
# Copied, not templated: these are Lua and must not go through Jinja.
- name: seed zomboid sandbox and spawn config from sophie preset
- name: seed zomboid sandbox settings
become: true
ansible.builtin.copy:
src: "zomboid/sophie/{{ item }}.lua"
dest: "{{ zomboid_path }}/data/Server/{{ zomboid_server_name }}_{{ item }}.lua"
src: zomboid/sophie/SandboxVars.lua
dest: "{{ zomboid_path }}/data/Server/{{ zomboid_server_name }}_SandboxVars.lua"
force: "{{ zomboid_config_force | bool }}"
owner: "{{ podman_subuid.stdout }}"
group: "{{ podman_user }}"
mode: 0644
notify: restorecon podman
tags: zomboid-conf
# Spawn config gets its own force switch, and it is not pedantry: these two
# files are the only part of the server config that a *running* world will pick
# up. PZ reads <name>_spawnregions.lua (and any serverfile it names) at every
# server start; SandboxVars is read once, when the world is created. So a spawn
# change is deployable on the live world -- push, restart, done -- while a
# SandboxVars change is not, and needs a wipe to mean anything.
#
# Sharing zomboid_config_force between them would mean force-pushing the whole
# preset to land a one-line spawnregions edit, which also rewrites the INI (a
# ResetID mismatch tells every client to reroll) and the world's SandboxVars.
#
# make deploy TAGS=zomboid-conf -e zomboid_spawn_force=true
#
# then restart the server -- with players offline -- for it to take effect.
- name: seed zomboid spawn config
become: true
ansible.builtin.copy:
src: "zomboid/sophie/{{ item }}.lua"
dest: "{{ zomboid_path }}/data/Server/{{ zomboid_server_name }}_{{ item }}.lua"
force: "{{ zomboid_spawn_force | bool }}"
owner: "{{ podman_subuid.stdout }}"
group: "{{ podman_user }}"
mode: 0644
loop:
- SandboxVars
- spawnregions
- spawnpoints
notify: restorecon podman
@@ -231,7 +293,7 @@
state: started
daemon_reload: true
# A pristine copy of the Sophie preset, kept where the server cannot overwrite it.
# The settings this repo would deploy, kept where the server cannot overwrite them.
#
# Reference only -- nothing applies this automatically. A wipe deliberately leaves
# the live settings alone so hand tuning survives it.
@@ -242,12 +304,21 @@
# much as an input, and that is how a Sophie world quietly became an Apocalypse
# one -- 139 values reverted, loot from 0.35 back to 0.9, CharacterFreePoints
# from 0 to 60. When that happens again, this is what to copy back from.
#
# What it holds changed on 2026-08-31. It was the pristine Sophie preset; it is
# now a snapshot of the live debbzoid world, because the admins' tuning had by
# then diverged from the preset in the same 139 values and re-seeding from the
# preset would have thrown that tuning away rather than restored it.
#
# force is on here, and only here. Nothing on the host writes this directory --
# the server cannot see it -- so it has no hand edits to protect, and if it does
# not track the repo it is not a restore point, just an older world's settings.
- name: seed canonical zomboid world settings
become: true
ansible.builtin.copy:
src: "zomboid/sophie/{{ item }}.lua"
dest: "{{ zomboid_path }}/config-template/{{ item }}.lua"
force: "{{ zomboid_config_force | bool }}"
force: true
owner: "{{ podman_user }}"
group: "{{ podman_user }}"
mode: 0644
@@ -723,3 +794,15 @@
enabled: true
state: started
daemon_reload: true
# Restore is triggered the same way -- Discord bot -> trigger file -> path unit.
# See zomboid-restore.path and zomboid-restore.service.
- name: enable zomboid restore path unit
become: true
become_user: "{{ podman_user }}"
ansible.builtin.systemd:
name: zomboid-restore.path
scope: user
enabled: true
state: started
daemon_reload: true
+8 -1
View File
@@ -123,9 +123,16 @@
- import_tasks: containers/home/gregtime.yml
vars:
image: localhost/greg-time-bot:3.16.5
image: localhost/greg-time-bot:3.17.3
tags: gregtime
# Built and loaded by `make deploy-remote` in ~/src/rsvp-debylio; bump this to
# the VERSION it loaded. The Caddy vhost ships with the caddy-config tag.
- import_tasks: containers/home/rsvp.yml
vars:
image: localhost/rsvpd:1.0.1
tags: rsvp
# Gated on zomboid_enabled (roles/podman/defaults/main.yml) so it can be taken
# down for a CI-heavy stretch without losing the world.
- import_tasks: containers/home/zomboid.yml
@@ -12,6 +12,23 @@
when: container.containers[0]["ImageName"] != container_image
ignore_errors: true
# Pull the new image BEFORE the old container is removed, so a tag that does
# not exist (or a registry that is down) fails the play here and leaves the
# running container untouched. Locally built images (localhost/...) are never
# pulled - they must already be in the podman user's storage. A container that
# does not exist yet has nothing to protect (and no containers[0] to compare),
# so the length check comes first; `when` list items stop at the first false.
- name: pull new image before replacing container
become: true
become_user: "{{ podman_user }}"
containers.podman.podman_image:
name: "{{ container_image }}"
state: present
when:
- container.containers | length > 0
- container.containers[0]["ImageName"] != container_image
- not container_image.startswith("localhost/")
- name: delete container if necessary
become: true
become_user: "{{ podman_user }}"
@@ -19,4 +36,4 @@
name: "{{ container_name }}"
state: absent
when: container.containers[0]["ImageName"] != container_image
ignore_errors: true
ignore_errors: true
@@ -23,6 +23,26 @@
}
format {{ caddy_log_format }}
level {{ caddy_log_level }}
# See rsvp-errors below.
exclude http.log.error.rsvp
}
# {{ rsvp_server_name }}: when a proxied request fails, Caddy's error log
# records the raw request URI -- which for /i/<token> is a working invite
# link. Send those through the same redaction as the site's access log
# instead of the default log above.
log rsvp-errors {
output file /var/log/caddy/rsvp.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format filter {
wrap json
request>uri regexp ^/i/[^/?#]+ /i/REDACTED
request>headers>Referer delete
}
include http.log.error.rsvp
}
}
@@ -193,6 +213,56 @@
}
}
# RSVP - {{ rsvp_server_name }} (public invite links, LAN-only /admin)
#
# Does NOT import common_headers: its Referrer-Policy "same-origin" would
# replace the app's "no-referrer", and invite tokens live in the URL path.
{{ rsvp_server_name }} {
header {
Strict-Transport-Security "max-age=31536000; includeSubDomains"
X-Robots-Tag "noindex, nofollow"
Referrer-Policy "no-referrer"
Content-Security-Policy "default-src 'self'; style-src 'self'; script-src 'self'; img-src 'self' data:; form-action 'self'; frame-ancestors 'none'"
X-Content-Type-Options "nosniff"
-Server
}
# Admin is a second listener with no login. It is only published on
# loopback, and only proxied here for LAN clients -- two independent controls.
@admin_local {
path /admin /admin/*
remote_ip {{ caddy_local_networks | join(' ') }}
}
handle @admin_local {
reverse_proxy 127.0.0.1:{{ rsvp_admin_port }}
}
handle /admin* {
respond "Not found" 404
}
handle {
reverse_proxy 127.0.0.1:{{ rsvp_public_port }}
}
# The invite token is the credential and it is in the URL path, so redact
# it from the request URI and drop the redirect Location (POSTs 303 back to
# /i/<token>) before anything is written. Named "rsvp" so its error logger
# (http.log.error.rsvp) can be redirected in the global options.
log rsvp {
output file /var/log/caddy/rsvp.log {
roll_size {{ caddy_log_roll_size }}
roll_keep {{ caddy_log_roll_keep }}
roll_keep_for {{ caddy_log_roll_keep_for }}
}
format filter {
wrap json
request>uri regexp ^/i/[^/?#]+ /i/REDACTED
request>headers>Referer delete
resp_headers>Location delete
}
}
}
# Uptime Kuma (Debyltech) - {{ uptime_kuma_server_name }}
{{ uptime_kuma_server_name }} {
{{ ip_restricted_site() }}
@@ -8,11 +8,30 @@
Only the keys this deployment actually owns are templated -- credentials,
RCON, Discord, backup cadence, and the mod lists. Every gameplay setting
(PVP, safety system, safehouses, sleep, MaxPlayers, PauseEmpty, ResetID,
the welcome message) is left exactly as the preset author shipped it.
the welcome message) is a literal, and was the preset author's value until
the live-config sync below.
Seeded once, then left alone: the deploy task uses force=false so the
server and admins can hand-edit this file between world regenerations.
Re-push it deliberately with -e zomboid_config_force=true.
Synced from the running debbzoid server on 2026-08-31, so a deliberate
force-push cannot silently revert what the admins set in-game. Eight keys
had drifted from the preset and now carry the live values:
PlayerSafehouse false -> true
SafehouseAllowNonResidential false -> true (the diner/gas-station case)
SafehouseAllowRespawn false -> true
SafehouseAllowLoot true -> false
SafehouseAllowFire true -> false
TrashDeleteAll false -> true
MapRemotePlayerVisibility 1 -> 4
ResetID 6953472 -> 826046
ResetID is in that list on purpose. It is the world's soft-reset token: if
the file's value ever differs from the one the live world was created with,
every connected client is told to roll a new character. Carrying the live
value means a force-push is a no-op instead of a server-wide wipe prompt.
-#}
{% set rename_map = {} %}
{% for r in zomboid_mods_renamed %}{% set _ = rename_map.update({r.old: r.new}) %}{% endfor %}
@@ -83,7 +102,7 @@ DefaultPort=16261
UDPPort=16262
# Reset ID determines if the server has undergone a soft-reset. If this number does match the client, the client must create a new character. Used in conjunction with PlayerServerID. It is strongly advised that you backup these IDs somewhere Min: 0 Max: 2147483647 Default: 985847558
ResetID=6953472
ResetID=826046
# Enter the mod loading ID here. It can be found in \Steam\steamapps\workshop\modID\mods\modName\info.txt
Mods={{ resolved_mod_ids | reject('in', excluded_mod_ids) | join(';') }}
@@ -131,7 +150,7 @@ AnnounceAnimalDeath=false
SaveWorldEveryMinutes=10
# Both admins and players can claim safehouses
PlayerSafehouse=false
PlayerSafehouse=true
# Only admins can claim safehouses
AdminSafehouse=false
@@ -140,13 +159,13 @@ AdminSafehouse=false
SafehouseAllowTrepass=true
# Allow fire to damage safehouses
SafehouseAllowFire=true
SafehouseAllowFire=false
# Allow non-members to take items from safehouses
SafehouseAllowLoot=true
SafehouseAllowLoot=false
# Players will respawn in a safehouse that they were a member of before they died
SafehouseAllowRespawn=false
SafehouseAllowRespawn=true
# Players must have survived this number of in-game days before they are allowed to claim a safehouse Min: 0 Max: 2147483647 Default: 0
SafehouseDaySurvivedToClaim=0
@@ -155,7 +174,7 @@ SafehouseDaySurvivedToClaim=0
SafeHouseRemovalTime=144
# Governs whether players can claim non-residential buildings.
SafehouseAllowNonResidential=false
SafehouseAllowNonResidential=true
SafehouseDisableDisguises=true
@@ -328,7 +347,7 @@ BanKickGlobalSound=true
RemovePlayerCorpsesOnCorpseRemoval=false
# If true, player can use the "delete all" button on bins.
TrashDeleteAll=false
TrashDeleteAll=true
# If true, player can hit again when struck by another player.
PVPMeleeWhileHitReaction=false
@@ -352,7 +371,7 @@ CarEngineAttractionModifier=0.5
PlayerBumpPlayer=false
# Controls display of remote players on the in-game map. 1=Hidden 2=Friends 3=Friends and nearby players 4=Everyone Min: 1 Max: 4 Default: 1
MapRemotePlayerVisibility=1
MapRemotePlayerVisibility=4
# Min: 1 Max: 300 Default: 5
BackupsCount=10
@@ -86,7 +86,8 @@ fi
# every later wipe regenerates from them -- which is how a Sophie world became an
# Apocalypse one. Nothing here corrects that automatically any more, so the
# snapshot above is the way back: config-backup/ holds the settings as they were
# before each wipe, and config-template/ holds the pristine Sophie preset.
# before each wipe, and config-template/ holds what this repo deploys -- since
# 2026-08-31 a snapshot of Greg's live tuning, not the pristine Sophie preset.
# Start server
log "Starting zomboid service..."
@@ -0,0 +1,9 @@
[Unit]
Description=Watch for Zomboid backup restore trigger
[Path]
PathExists={{ podman_home }}/.local/share/volumes/gregtime/data/zomboid-restore.trigger
Unit=zomboid-restore.service
[Install]
WantedBy=default.target
@@ -0,0 +1,8 @@
[Unit]
Description=Zomboid Backup Restore Service
[Service]
Type=oneshot
ExecStart={{ podman_home }}/bin/zomboid-restore.sh
StandardOutput=journal
StandardError=journal
@@ -0,0 +1,222 @@
#!/bin/bash
# Zomboid Backup Restore Script
# Triggered by systemd path unit when the discord bot requests a restore.
#
# Sibling of world-reset.sh: same trigger-file -> path-unit -> oneshot shape, same
# podman unshare discipline. The difference is that a reset throws the world away
# and this puts an older one back, so it is a good deal more careful:
#
# - it resolves the target itself and never accepts a path from the trigger
# - it refuses and exits BEFORE stopping the server if anything is wrong
# - it moves the live world aside instead of deleting it, so every restore
# is undoable
set -e
VOL="{{ podman_home }}/.local/share/volumes"
LOGFILE="${VOL}/zomboid/logs/restore.log"
TRIGGER_FILE="${VOL}/gregtime/data/zomboid-restore.trigger"
RESULT_FILE="${VOL}/gregtime/data/zomboid-restore.result"
SERVER_NAME="{{ zomboid_server_name }}"
DATA="${VOL}/zomboid/data"
SAVES_PATH="${DATA}/Saves/Multiplayer/${SERVER_NAME}"
DB_PATH="${DATA}/db/${SERVER_NAME}.db"
BACKUPS="${DATA}/backups"
SNAP_ROOT="${VOL}/zomboid/restore-backup"
STAGING="${VOL}/zomboid/restore-staging"
KEEP_SNAPSHOTS=3
log() {
local msg="[$(date '+%Y-%m-%d %H:%M:%S')] $1"
echo "$msg"
# Same reasoning as world-reset.sh: stdout is already in the journal, and an
# unwritable log file must never be what aborts a restore under set -e.
echo "$msg" >> "$LOGFILE" 2>/dev/null || true
}
# The bot reads this back to report into #zomboid. Written as the container uid
# so the bot (uid 1000 inside its own container) can actually open it.
result() {
local ok="$1" detail="$2"
printf '{"ok":%s,"detail":%s,"at":"%s"}\n' \
"$ok" "$(printf '%s' "$detail" | sed 's/\\/\\\\/g; s/"/\\"/g; s/^/"/; s/$/"/')" \
"$(date -Is)" > /tmp/zomboid-restore.result.$$ 2>/dev/null || return 0
podman unshare cp /tmp/zomboid-restore.result.$$ "$RESULT_FILE" 2>/dev/null || true
podman unshare chown 1000:1000 "$RESULT_FILE" 2>/dev/null || true
rm -f /tmp/zomboid-restore.result.$$ 2>/dev/null || true
}
die() {
log "ABORT: $1"
result false "$1"
exit 1
}
export XDG_RUNTIME_DIR="/run/user/$(id -u)"
log "Restore triggered"
# ---------------------------------------------------------------------------
# 1. Read and validate the trigger
# ---------------------------------------------------------------------------
podman unshare test -f "$TRIGGER_FILE" || die "no trigger file"
TRIGGER_BODY="$(podman unshare cat "$TRIGGER_FILE")"
podman unshare rm -f "$TRIGGER_FILE"
field() { printf '%s\n' "$TRIGGER_BODY" | sed -n "s/^$1=//p" | head -1 | tr -d '\r'; }
ACTION="$(field action)"
SET="$(field set)"
INDEX="$(field index)"
MTIME="$(field mtime)"
REQUESTER="$(field requester)"
log "Requested by: ${REQUESTER:-unknown} (action=${ACTION:-restore})"
[[ "$ACTION" == "restore" || "$ACTION" == "undo" ]] || die "bad action '${ACTION}'"
# ---------------------------------------------------------------------------
# 2. Resolve the source -- entirely from our own filesystem, never from the
# trigger. The trigger only ever gets to *describe* a target.
# ---------------------------------------------------------------------------
SOURCE_DESC=""
SRC_ZIP=""
UNDO_DIR=""
if [[ "$ACTION" == "restore" ]]; then
[[ "$SET" == "period" || "$SET" == "startup" || "$SET" == "version" ]] \
|| die "bad backup set '${SET}'"
[[ "$MTIME" =~ ^[0-9]+$ ]] || die "bad mtime '${MTIME}'"
# Resolve by mtime, NOT by index. PZ rotates these every BackupsPeriod
# minutes -- backup_7.zip becomes backup_8.zip and so on -- so the index the
# bot showed a human 90 seconds ago may already point at a different world.
# The mtime is the only stable identity a PZ backup has.
for f in "${BACKUPS}/period"/backup_*.zip "${BACKUPS}/startup"/backup_*.zip "${BACKUPS}/version"/backup_*.zip; do
podman unshare test -f "$f" || continue
base="$(basename "$f")"
[[ "$base" =~ ^backup_[0-9]+\.zip$ ]] || continue
m="$(podman unshare stat -c %Y "$f" 2>/dev/null || echo 0)"
if [[ "$m" == "$MTIME" ]]; then
SRC_ZIP="$f"
break
fi
done
[[ -n "$SRC_ZIP" ]] || die "backup from $(date -d "@${MTIME}" '+%Y-%m-%d %H:%M:%S' 2>/dev/null || echo "$MTIME") has rotated out -- run the list again"
podman unshare unzip -l "$SRC_ZIP" "Saves/Multiplayer/${SERVER_NAME}/*" >/dev/null 2>&1 \
|| die "archive $(basename "$SRC_ZIP") has no ${SERVER_NAME} world in it"
SOURCE_DESC="$(basename "$SRC_ZIP") (${SET}, $(date -d "@${MTIME}" '+%Y-%m-%d %H:%M:%S' 2>/dev/null || echo "$MTIME"))"
log "Resolved target: $SRC_ZIP -> $SOURCE_DESC"
else
# Newest snapshot that actually holds a world.
for d in $(ls -1dt "${SNAP_ROOT}"/*/ 2>/dev/null); do
if podman unshare test -d "${d}${SERVER_NAME}"; then
UNDO_DIR="${d%/}"
break
fi
done
[[ -n "$UNDO_DIR" ]] || die "no snapshot to undo to"
SOURCE_DESC="snapshot $(basename "$UNDO_DIR")"
log "Resolved undo target: $UNDO_DIR"
fi
# ---------------------------------------------------------------------------
# 3. Stop the server. Nothing above this line touches it, so every failure mode
# up to here leaves a running world completely alone.
# ---------------------------------------------------------------------------
# Disarm the wipe trigger for the duration. A reset landing midway through a
# restore would delete the half-restored world and leave nothing coherent.
log "Disarming world-reset path unit"
systemctl --user stop zomboid-world-reset.path || true
log "Stopping zomboid service..."
systemctl --user stop zomboid.service || true
sleep 5
# ---------------------------------------------------------------------------
# 4. Move the live world aside. Never delete -- this is the undo point.
# ---------------------------------------------------------------------------
SNAP_DIR="${SNAP_ROOT}/$(date '+%Y-%m-%d_%H-%M-%S')"
mkdir -p "$SNAP_DIR" 2>/dev/null || die "could not create $SNAP_DIR"
if podman unshare test -d "$SAVES_PATH"; then
podman unshare mv "$SAVES_PATH" "${SNAP_DIR}/${SERVER_NAME}"
log "Live world moved to $(basename "$SNAP_DIR")"
fi
if podman unshare test -f "$DB_PATH"; then
podman unshare cp "$DB_PATH" "${SNAP_DIR}/${SERVER_NAME}.db"
fi
# ---------------------------------------------------------------------------
# 5. Put the target in place
# ---------------------------------------------------------------------------
podman unshare rm -rf "$STAGING"
podman unshare mkdir -p "$STAGING"
if [[ "$ACTION" == "restore" ]]; then
log "Extracting ${SOURCE_DESC}..."
podman unshare unzip -q "$SRC_ZIP" \
"Saves/Multiplayer/${SERVER_NAME}/*" "db/${SERVER_NAME}.db" -d "$STAGING" \
|| die "extract failed"
podman unshare mv "${STAGING}/Saves/Multiplayer/${SERVER_NAME}" "$SAVES_PATH"
if podman unshare test -f "${STAGING}/db/${SERVER_NAME}.db"; then
podman unshare mv "${STAGING}/db/${SERVER_NAME}.db" "$DB_PATH"
fi
# PZ's own backups do not capture everything under the save directory --
# mod state such as blam/ is excluded from every archive it writes. Those
# files were equally present at the moment we are restoring to, so dropping
# them would land the world slightly *behind* the target rather than on it.
# Carry across anything the snapshot had that the archive did not.
for entry in $(podman unshare ls -1 "${SNAP_DIR}/${SERVER_NAME}" 2>/dev/null); do
if ! podman unshare test -e "${SAVES_PATH}/${entry}"; then
podman unshare cp -a "${SNAP_DIR}/${SERVER_NAME}/${entry}" "${SAVES_PATH}/${entry}" 2>/dev/null || true
log "Carried forward un-backed-up entry: ${entry}"
fi
done
else
log "Restoring ${SOURCE_DESC}..."
podman unshare cp -a "${UNDO_DIR}/${SERVER_NAME}" "$SAVES_PATH"
if podman unshare test -f "${UNDO_DIR}/${SERVER_NAME}.db"; then
podman unshare cp -a "${UNDO_DIR}/${SERVER_NAME}.db" "$DB_PATH"
fi
fi
podman unshare rm -rf "$STAGING"
# ---------------------------------------------------------------------------
# 6. Ownership, permissions, labels
# ---------------------------------------------------------------------------
# Container uid 1000; on the host that lands on the subuid base + 999. Same call
# zomboid.yml makes after seeding config.
log "Fixing ownership and permissions..."
podman unshare chown -R 1000:1000 "$SAVES_PATH" "$DB_PATH"
podman unshare chmod -R u+rwX,g+rwX "$SAVES_PATH"
podman unshare chmod 0664 "$DB_PATH"
# Targeted, and -x so it can never wander off this filesystem. The role-wide
# restorecon handler exists in -x form for a reason; see roles/podman/handlers.
if command -v restorecon >/dev/null 2>&1; then
restorecon -Frx "$SAVES_PATH" "$DB_PATH" 2>/dev/null || true
fi
# ---------------------------------------------------------------------------
# 7. Prune old snapshots
# ---------------------------------------------------------------------------
for old in $(ls -1dt "${SNAP_ROOT}"/*/ 2>/dev/null | tail -n +$((KEEP_SNAPSHOTS + 1))); do
podman unshare rm -rf "$old" || true
log "Pruned old snapshot $(basename "${old%/}")"
done
# ---------------------------------------------------------------------------
# 8. Back up
# ---------------------------------------------------------------------------
log "Starting zomboid service..."
systemctl --user start zomboid.service
systemctl --user start zomboid-world-reset.path || true
log "Restore complete: ${SOURCE_DESC}"
result true "Restored ${SOURCE_DESC}. Previous world kept as $(basename "$SNAP_DIR")."
Binary file not shown.