fix(immich): keep the ML models resident, and pin the ML URL explicitly

Smart search, face detection and OCR were all failing with

  Machine learning request to "http://immich-machine-learning:3003" failed

while the container still reported Up. Podman only sees PID 1: gunicorn's
master was alive and holding the listening socket, but its worker had died at a
WORKER TIMEOUT and was never respawned -- hence a connect timeout rather than a
refusal. The dead worker was a zombie whose remaining thread was stuck in
uninterruptible sleep in exit_mmap, so it survived SIGKILL, podman rm -f and
rm -f -t 0, and kept the container name and network alias until the host was
rebooted.

MACHINE_LEARNING_MODEL_TTL=0 addresses the cause rather than the symptom. The
default unloads models after 300s idle, so every search following a gap
reloaded four of them (CLIP, buffalo_l detection + recognition, PP-OCRv5) on a
CPU-only 4-core box and then tore those mappings back down -- and that teardown
is what wedged. Keeping them resident costs ~1-2 GB and removes the path.

IMMICH_MACHINE_LEARNING_URL is now set explicitly instead of relying on
immich's implicit default, so the ML container can be renamed without silently
losing search, and the name is a variable so a replacement can be stood up
beside a broken one without editing tasks.

Verified after deploy: zero ML failures, "in-memory cache with unloading
disabled", ping 200 from immich-server, 26/26 containers healthy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Bastian de Byl
2026-08-28 11:39:57 -04:00
co-authored by Claude Opus 5
parent 2a6390c0ad
commit ae99ba415e
2 changed files with 24 additions and 3 deletions
+8
View File
@@ -27,6 +27,14 @@ hass_path: "{{ podman_volumes }}/hass"
partsy_path: "{{ podman_volumes }}/partsy" partsy_path: "{{ podman_volumes }}/partsy"
partsy_skudak_path: "{{ podman_volumes }}/partsy-skudak" partsy_skudak_path: "{{ podman_volumes }}/partsy-skudak"
photos_path: "{{ podman_volumes }}/photos" photos_path: "{{ podman_volumes }}/photos"
# Named rather than hardcoded so the ML service can be stood up beside a broken
# one without editing tasks. On 2026-08-28 a worker thread wedged in
# uninterruptible sleep (D state) in exit_mmap, which made the container
# unremovable -- it survived SIGKILL and `podman rm -f` and kept its name and
# network alias until the host was rebooted. The reboot cleared it, so this
# stays on the canonical name; see MACHINE_LEARNING_MODEL_TTL in photos.yml for
# the change that stops it recurring.
immich_ml_container: immich-machine-learning
uptime_kuma_path: "{{ podman_volumes }}/uptime-kuma" uptime_kuma_path: "{{ podman_volumes }}/uptime-kuma"
uptime_kuma_personal_path: "{{ podman_volumes }}/uptime-kuma-personal" uptime_kuma_personal_path: "{{ podman_volumes }}/uptime-kuma-personal"
zomboid_path: "{{ podman_volumes }}/zomboid" zomboid_path: "{{ podman_volumes }}/zomboid"
@@ -94,26 +94,35 @@
- import_tasks: podman/podman-check.yml - import_tasks: podman/podman-check.yml
vars: vars:
container_name: immich-machine-learning container_name: "{{ immich_ml_container }}"
container_image: "{{ ml_image }}" container_image: "{{ ml_image }}"
# MACHINE_LEARNING_MODEL_TTL=0 keeps the models resident instead of unloading
# them after 300s idle. The default made every search following an idle gap
# reload four models (CLIP, buffalo_l detection + recognition, PP-OCRv5) on a
# CPU-only 4-core box, then tear those mappings down again -- and it is that
# teardown, in exit_mmap, that wedged a worker thread in uninterruptible sleep
# on 2026-08-27 and took smart search, face detection and OCR down with it.
# Costs ~1-2 GB resident to remove the code path entirely.
- name: create immich-ml container - name: create immich-ml container
become: true become: true
become_user: "{{ podman_user }}" become_user: "{{ podman_user }}"
containers.podman.podman_container: containers.podman.podman_container:
name: immich-machine-learning name: "{{ immich_ml_container }}"
image: "{{ ml_image }}" image: "{{ ml_image }}"
restart_policy: on-failure:3 restart_policy: on-failure:3
log_driver: journald log_driver: journald
network: network:
- shared - shared
env:
MACHINE_LEARNING_MODEL_TTL: "0"
volumes: volumes:
- "{{ photos_path }}/mlcache:/cache" - "{{ photos_path }}/mlcache:/cache"
- name: create systemd startup job for immich-machine-learning - name: create systemd startup job for immich-machine-learning
include_tasks: podman/systemd-generate.yml include_tasks: podman/systemd-generate.yml
vars: vars:
container_name: immich-machine-learning container_name: "{{ immich_ml_container }}"
- import_tasks: podman/podman-check.yml - import_tasks: podman/podman-check.yml
vars: vars:
@@ -186,6 +195,10 @@
DB_USERNAME: photos DB_USERNAME: photos
DB_PASSWORD: "{{ photos_db_pass }}" DB_PASSWORD: "{{ photos_db_pass }}"
IMMICH_PORT: 8088 IMMICH_PORT: 8088
# Set explicitly rather than relying on immich's default of
# http://immich-machine-learning:3003, so the ML container can be renamed
# without silently losing search, faces and OCR.
IMMICH_MACHINE_LEARNING_URL: "http://{{ immich_ml_container }}:3003"
volumes: volumes:
- "{{ photos_path }}/storage:/mnt/media/originals" - "{{ photos_path }}/storage:/mnt/media/originals"
- "{{ photos_path }}/immich:/usr/src/app/upload" - "{{ photos_path }}/immich:/usr/src/app/upload"