Files
deploy_home/ansible/roles/gitea-actions
Bastian de BylandClaude Opus 5 d0e76bd6cf feat(gitea-actions): build the CI job images in Gitea CI
The runner's job images were built by ansible into localhost/ only, so the
nightly CI prune deleted them and every idle stretch ended with CI failing in
under a second on `docker pull localhost/gitea-ci:latest` until someone re-ran
the role and waited out a rebuild. The previous commit moved them to the Gitea
registry; this moves the *build* off the deploy path entirely.

- .gitea/workflows/ci-images.yml builds files/Containerfile.* and pushes to
  git.debyl.io/gitbot/. Per-image change detection, so an ESP-IDF pin bump does
  not rebuild the other two; weekly schedule for base-image updates; a
  workflow_dispatch selector. PRs build under a throwaway :pr-<n> tag and drop
  it -- the build lands in the live runner's store, and act_runner will not
  re-pull a tag it already has, so a PR using the real tag would hand every
  later job on this host an unmerged image.
- The Containerfiles stop being ansible templates: their version vars are now
  --build-arg, read by the workflow out of the same defaults/main.yml the role
  interpolates, so CI and ansible build the same bytes from one set of pins.
- LABEL io.debyl.ci-base moves into each Containerfile so neither builder can
  forget the prune exemption; the workflow re-checks it before pushing.
- roles/gitea-actions pulls instead of building. gitea_ci_build_local=true
  restores the local build+push for seeding a cold registry or when CI is
  down -- the workflow that builds gitea-ci runs in gitea-ci.
- Lint .gitea/ alongside ansible/, and document the flow in the role README.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:10:43 -04:00
..

gitea-actions

Runs the Gitea Actions runners on home.debyl.io. One act_runner process per Gitea instance (git.debyl.io, git.skudak.com), both as the gitea-runner user, both backed by the same rootless podman image store.

CI job images

Jobs do not run on the host. Each one gets an ephemeral container from one of three images:

runs-on / container: Image Used by
fedora, ubuntu-latest, ubuntu-22.04 git.debyl.io/gitbot/gitea-ci:latest Go / node / web jobs, docker build
container: image: git.debyl.io/gitbot/gitea-ci-espidf:<esp_idf_version> esp-mg-tpms, skudak/esp32-stm32-vcu
container: image: git.debyl.io/gitbot/gitea-ci-platformio:<pio_espressif32_version> skudak/esp32-web-interface

This role does not build them. .gitea/workflows/ci-images.yml builds files/Containerfile.* and pushes to the Gitea registry; the role logs gitea-runner in and pulls. Version pins live in defaults/main.yml and are read by both the role and the workflow, so a bump moves the image tag in one place.

Why the registry

The images used to exist only as localhost/gitea-ci* in the runner's store. The nightly prune (roles/podman, podman_prune_ci_until: 48h) deletes any CI-user image older than that which no container holds, so after an idle weekend every job failed in under a second on docker pull localhost/gitea-ci:latest, and the only fix was re-running this role and waiting out a full rebuild.

Two things now keep that from happening:

  • A registry copy. force_pull stays false, which in act_runner means pull only when missing — so a present image is never re-fetched, and a pruned one is restored by the next job without anyone noticing.
  • A prune exemption. Each Containerfile declares LABEL io.debyl.ci-base="true", and the prune skips that label (podman_prune_ci_keep_label). Its until counts from build time, not pull time, so without this a re-pulled image would be deleted again the same night — a 7.8 GB ESP-IDF download every single day.

The label is declared in the Containerfile rather than passed as --label so neither builder can omit it; the workflow re-checks it with docker inspect before pushing.

Authentication

Both the role and act_runner read /home/gitea-runner/.docker/config.json. act_runner uses it for the job-image pull it performs when a label's image is missing; podman falls back to the same file. The role writes it from gitea_registry_username / gitea_registry_token (vault), so one login covers both. The skudak runner pulls from git.debyl.io too — same host, same user, same file.

The workflow pushes with a REGISTRY_TOKEN secret on bastian/deploy_home, belonging to the same gitbot user: Gitea authorises a package push by the token's owner, not by the path, so pushing to gitbot/ means logging in as gitbot.

Rebuilding

Normally nothing to do — edit a files/Containerfile.* or a version pin, push to master, and the workflow rebuilds only the affected images. It also rebuilds everything weekly so base-image updates land without a commit, and takes a workflow_dispatch with an image selector.

Pull requests build but do not push, under a throwaway :pr-<n> tag that is deleted afterwards. The build runs in the live runner's image store, so a PR tagged with the real name would hand every later job on this host an unmerged image.

Bootstrap / CI is down

The workflow that builds gitea-ci runs in gitea-ci, so a registry that has never held it cannot bootstrap itself. Build on the host instead:

make deploy TAGS=gitea-actions EXTRA_VARS="gitea_ci_build_local=true"

That builds all three from the same Containerfiles and pushes them. One run is enough even on a cold registry: tasks/main.yml imports images.yml before runner.yml, so the images are published before the runner labels are flipped to point at them.

The alternative first-time path is to merge the workflow and dispatch it while the deployed labels still say localhost/ — the job then builds inside the old local image and seeds the registry — then run a plain make deploy TAGS=gitea-actions to switch the labels over.