fix(gitea-actions): serve CI images from the Gitea registry #12

Merged
bastian merged 12 commits from feat/ci-images-registry into master 2026-09-28 12:19:57 -04:00
Owner

Why

CI job images existed only as localhost/gitea-ci* in the gitea-runner store. The nightly CI prune (podman_prune_ci_until: 48h) deletes any image older than 48h that no container is using. A local-only image can't be pulled back, so after every idle stretch all CI failed in 0-1s (docker pull localhost/gitea-ci:latest → tls: internal error) until make deploy TAGS=gitea-actions rebuilt everything. This happened on 2026-09-10 and again on 2026-09-14, when the ESP-IDF image had been pruned too.

What

  • Registry-backed images: built as git.debyl.io/gitbot/gitea-ci:latest, gitea-ci-espidf:v5.4.1, and gitea-ci-platformio:7.0.1, and pushed on every role run.
  • Pull before build: the role pulls from the registry when the image is missing and rebuilds only when the Containerfile changed, so there's no more long rebuild just to restore an image.
  • Runner login: gitea-runner logs in to the registry with the existing gitea_registry_username/gitea_registry_token, written to ~/.docker/config.json. act_runner reads that file for job image pulls, and podman uses it as its fallback auth file.
  • Prune exemption: base images get the label io.debyl.ci-base=true, and the CI prune skips it (label!=). The prune's until counts from build time, not pull time, so without this a re-pulled 7.8 GB ESP-IDF image would be deleted again the next night. Superseded tags are kept too and need manual cleanup.
  • The three images are now one gitea_ci_images list instead of copy-pasted tasks.

Rollout

  1. Run make deploy TAGS=gitea-actions,podman-prune. The first run tags from the build cache, so it should be quick, then pushes about 11 GB to the registry.
  2. Update workflows that pin container: image: localhost/... before the old local images are pruned (about 48h):
    • debyltech/esp-mg-tpms .gitea/workflows/package.yml (use a chore(ci): commit, since fix:/feat: publishes an OTA)
    • skudak/esp32-stm32-vcu .gitea/workflows/package.yml
    • skudak/esp32-web-interface .gitea/workflows/build.yml

Checks

  • Checked: ansible-playbook --syntax-check, make lint, and a rendered prune script with bash -n.
  • Checked on the host: podman 5.7.1 accepts label!=io.debyl.ci-base as an image filter.
  • Not verified: whether the gitbot token has write:package. If the push task fails with 401/403, the token needs that scope.

🤖 Generated with Claude Code

## Why CI job images existed only as `localhost/gitea-ci*` in the gitea-runner store. The nightly CI prune (`podman_prune_ci_until: 48h`) deletes any image older than 48h that no container is using. A local-only image can't be pulled back, so after every idle stretch all CI failed in 0-1s (`docker pull localhost/gitea-ci:latest` → `tls: internal error`) until `make deploy TAGS=gitea-actions` rebuilt everything. This happened on 2026-09-10 and again on 2026-09-14, when the ESP-IDF image had been pruned too. ## What - **Registry-backed images:** built as `git.debyl.io/gitbot/gitea-ci:latest`, `gitea-ci-espidf:v5.4.1`, and `gitea-ci-platformio:7.0.1`, and pushed on every role run. - **Pull before build:** the role pulls from the registry when the image is missing and rebuilds only when the Containerfile changed, so there's no more long rebuild just to restore an image. - **Runner login:** gitea-runner logs in to the registry with the existing `gitea_registry_username`/`gitea_registry_token`, written to `~/.docker/config.json`. act_runner reads that file for job image pulls, and podman uses it as its fallback auth file. - **Prune exemption:** base images get the label `io.debyl.ci-base=true`, and the CI prune skips it (`label!=`). The prune's `until` counts from build time, not pull time, so without this a re-pulled 7.8 GB ESP-IDF image would be deleted again the next night. Superseded tags are kept too and need manual cleanup. - The three images are now one `gitea_ci_images` list instead of copy-pasted tasks. ## Rollout 1. Run `make deploy TAGS=gitea-actions,podman-prune`. The first run tags from the build cache, so it should be quick, then pushes about 11 GB to the registry. 2. Update workflows that pin `container: image: localhost/...` before the old local images are pruned (about 48h): - `debyltech/esp-mg-tpms` `.gitea/workflows/package.yml` (use a `chore(ci):` commit, since `fix:`/`feat:` publishes an OTA) - `skudak/esp32-stm32-vcu` `.gitea/workflows/package.yml` - `skudak/esp32-web-interface` `.gitea/workflows/build.yml` ## Checks - Checked: `ansible-playbook --syntax-check`, `make lint`, and a rendered prune script with `bash -n`. - Checked on the host: podman 5.7.1 accepts `label!=io.debyl.ci-base` as an image filter. - **Not verified:** whether the `gitbot` token has `write:package`. If the push task fails with 401/403, the token needs that scope. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
bastian added 1 commit 2026-09-14 16:44:18 -04:00
The CI job images only existed under localhost/ in the gitea-runner store,
and the nightly CI prune deletes any image older than 48h that no container
holds. After every idle stretch CI failed in 0-1s pulling
localhost/gitea-ci:latest until the role was re-run and the images rebuilt.

- Build under git.debyl.io/gitbot/..., push after every run, and pull from the
  registry instead of rebuilding when the Containerfile is unchanged.
- Log gitea-runner in via ~/.docker/config.json, which both act_runner (job
  image pulls) and podman read.
- Label the base images io.debyl.ci-base and skip that label in the CI prune;
  its `until` counts from build time, so a re-pulled image would otherwise be
  deleted again the next night.

Workflows pinning `container: image: localhost/gitea-ci-*` must move to the
registry paths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
bastian added 1 commit 2026-09-14 19:57:52 -04:00
bastian added 5 commits 2026-09-22 11:37:31 -04:00
The runner's job images were built by ansible into localhost/ only, so the
nightly CI prune deleted them and every idle stretch ended with CI failing in
under a second on `docker pull localhost/gitea-ci:latest` until someone re-ran
the role and waited out a rebuild. The previous commit moved them to the Gitea
registry; this moves the *build* off the deploy path entirely.

- .gitea/workflows/ci-images.yml builds files/Containerfile.* and pushes to
  git.debyl.io/gitbot/. Per-image change detection, so an ESP-IDF pin bump does
  not rebuild the other two; weekly schedule for base-image updates; a
  workflow_dispatch selector. PRs build under a throwaway :pr-<n> tag and drop
  it -- the build lands in the live runner's store, and act_runner will not
  re-pull a tag it already has, so a PR using the real tag would hand every
  later job on this host an unmerged image.
- The Containerfiles stop being ansible templates: their version vars are now
  --build-arg, read by the workflow out of the same defaults/main.yml the role
  interpolates, so CI and ansible build the same bytes from one set of pins.
- LABEL io.debyl.ci-base moves into each Containerfile so neither builder can
  forget the prune exemption; the workflow re-checks it before pushing.
- roles/gitea-actions pulls instead of building. gitea_ci_build_local=true
  restores the local build+push for seeding a cold registry or when CI is
  down -- the workflow that builds gitea-ci runs in gitea-ci.
- Lint .gitea/ alongside ansible/, and document the flow in the role README.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The daily quote keeps a ledger of what it has posted and re-rolls ZenQuotes
until it finds something the channel has not read, with the header framing and
the offline fallback pool drawn against that same ledger.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1.26.1 -> 1.27.3 picks up the security fixes in 1.27.0 through 1.27.3.
The image was one shared variable, so the two instances could only move
together; split it into gitea_debyl_image / gitea_skudak_image and tag
the debyl tasks gitea-debyl so each can be upgraded and verified on its
own. Skudak stays on 1.26.1 in this commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
chore(gitea): bump git.skudak.com to 1.27.3
CI Images / Plan (pull_request) Successful in 37s
CI Images / Build ${{ matrix.key }} (pull_request) Failing after 58s
15a8ec693e
Same security fixes as git.debyl.io, deployed and verified after it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
bastian added 1 commit 2026-09-25 12:42:42 -04:00
fix(hass): driveway lights ignore TV mode, ramped evening dimming, bump to 2026.9.3
CI Images / Plan (pull_request) Failing after 2s
CI Images / Build ${{ matrix.key }} (pull_request) Skipped
fd985e014c
The sunset automation required TV mode off, so an afternoon of TV skipped
the driveway string lights entirely (Sep 22 and 23) - not an outage. The
driveway now has its own sunset -1h automation, with catch-up on restart
or switch reconnect before 23:00.

Evening brightness lives in one script (evening_lights_apply) that blends
between the old step levels; a 5-minute ramp from 20:30 eases lights that
are on and leaves alone any a person has changed by hand. The Dining Hall
no longer bumps to 50% at 21:30.

TV off after 23:30 now only turns off the living room glow instead of
bringing the whole house back to full brightness, and TV on/off leave the
lights alone in daylight. Lights-out moves to 23:30, and a 01:00 sweep
catches anything switched back on at the wall.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bastian added 1 commit 2026-09-28 11:32:30 -04:00
fix(hass): don't fire on-off automations when a device reconnects
CI Images / Plan (pull_request) Failing after 2s
CI Images / Build ${{ matrix.key }} (pull_request) Skipped
f674f61b8d
The Bedroom Light HS200 dropped off Wi-Fi for 5 s at 03:01 and came back
reporting "on"; the Bedroom On device trigger treated unavailable -> on as
someone flipping the switch and lit the bedroom Hue lamps at 100%.

Bedroom On/Off and TV On/Off now use state triggers with not_from
unavailable/unknown, so reconnects and HA restarts no longer count as a
flip. The driveway's switch-reconnect catch-up is dropped for the same
reason (it would undo a manual off); the HA-restart catch-up stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bastian added 3 commits 2026-09-28 12:19:39 -04:00
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
docs(claude): work directly on master, no branches or PRs
CI Images / Plan (pull_request) Failing after 2s
CI Images / Build ${{ matrix.key }} (pull_request) Skipped
73c552303b
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bastian merged commit 0eca63d4b7 into master 2026-09-28 12:19:57 -04:00
bastian deleted branch feat/ci-images-registry 2026-09-28 12:19:57 -04:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: bastian/deploy_home#12