harden TrueNAS CIFS mounts so immich self-heals
TrueNAS was power-cycled, the CIFS mounts failed, and systemd never retried
-- mount units are not restarted on failure. SMB came back, nothing
remounted, and immich-server served an empty library for days while its
database still listed 14,181 assets pointing at /mnt/media/originals.
fstab carried no _netdev, no nofail and no automount, so there was no path
back without a human. Now:
x-systemd.automount any access re-attempts the mount; failure stops
being terminal
_netdev / nofail ordered after network-online, dead NAS cannot block boot
soft I/O errors instead of blocking forever, so the
container can be restarted rather than wedging in
uninterruptible sleep
idle-timeout unmount when unused, clearing stale handles
resilienthandles SMB3 rides out brief blips
ansible.posix.mount mounts directly and never starts the generated
.automount unit, leaving the on-access trigger inactive -- enable it
explicitly, or the headline fix silently does nothing.
The containers are systemd USER units while the mounts are SYSTEM units, so
RequiresMountsFor= is unavailable. cifs-watchdog bridges the scopes: checks
health, recovers, and restarts ONLY immich-server (the sole consumer of both
paths; postgres/redis/ML use local volumes).
Two bugs the umount test caught, both worth knowing:
- `ls` cannot test mountedness. An unmounted mount point is an ordinary
empty directory, so ls succeeds and recovery was skipped entirely.
- A drop repaired within a single run leaves prev=healthy, so keying the
restart solely on the stored state skipped it while the container still
held its stale view.
Also moves the SMB password out of /etc/fstab, which is 0644 and was
readable by every local user, into a 0600 credentials file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -17,13 +17,56 @@
|
||||
- name: flush handlers
|
||||
ansible.builtin.meta: flush_handlers
|
||||
|
||||
- name: create cifs credentials directory
|
||||
become: true
|
||||
ansible.builtin.file:
|
||||
path: /etc/cifs
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: 0700
|
||||
|
||||
# /etc/fstab is 0644 by design, so an inline `password=` is readable by every
|
||||
# local user. Keep the credential in a 0600 file and reference it instead.
|
||||
- name: deploy photos cifs credentials
|
||||
become: true
|
||||
ansible.builtin.template:
|
||||
src: cifs-credentials.j2
|
||||
dest: /etc/cifs/photos.creds
|
||||
owner: root
|
||||
group: root
|
||||
mode: 0600
|
||||
vars:
|
||||
cifs_username: photos
|
||||
cifs_password: "{{ photos_cifs_pass }}"
|
||||
no_log: true
|
||||
|
||||
# These mounts previously failed permanently whenever TrueNAS was power-cycled:
|
||||
# a systemd .mount unit does NOT retry after a failed attempt, so the share came
|
||||
# back and nothing remounted, leaving immich serving an empty library while its
|
||||
# database still referenced 14k assets.
|
||||
#
|
||||
# x-systemd.automount the actual fix -- any ACCESS to the path re-attempts the
|
||||
# mount, so a failure stops being terminal
|
||||
# _netdev / nofail order after network-online; a dead NAS must not block boot
|
||||
# soft fail I/O with an error instead of blocking forever, so
|
||||
# immich-server can be restarted during an outage rather
|
||||
# than wedging in uninterruptible sleep. Accepted trade-off:
|
||||
# a write interrupted mid-flight fails and is retried.
|
||||
# idle-timeout unmount when unused, which clears stale handles instead
|
||||
# of nursing a half-dead connection
|
||||
# resilienthandles SMB3 rides out brief server blips transparently
|
||||
#
|
||||
# See also cifs-watchdog.sh.j2, which restarts immich-server when a mount that
|
||||
# was unhealthy becomes healthy again -- the automount restores the FILESYSTEM,
|
||||
# but the container still holds the old, empty view until it is bounced.
|
||||
- name: mount photos cifs
|
||||
become: true
|
||||
ansible.posix.mount:
|
||||
src: "{{ photos_cifs_src }}"
|
||||
path: "{{ photos_path }}/storage"
|
||||
fstype: cifs
|
||||
opts: "username=photos,password={{ photos_cifs_pass }},uid={{ podman_subuid.stdout }},gid={{ podman_subuid.stdout }}"
|
||||
opts: "{{ cifs_mount_opts }}"
|
||||
state: mounted
|
||||
|
||||
- name: mount immich cifs
|
||||
@@ -32,9 +75,23 @@
|
||||
src: "{{ immich_cifs_src }}"
|
||||
path: "{{ photos_path }}/immich"
|
||||
fstype: cifs
|
||||
opts: "username=photos,password={{ photos_cifs_pass }},uid={{ podman_subuid.stdout }},gid={{ podman_subuid.stdout }}"
|
||||
opts: "{{ cifs_mount_opts }}"
|
||||
state: mounted
|
||||
|
||||
# systemd-fstab-generator creates the .automount unit from x-systemd.automount,
|
||||
# but ansible.posix.mount mounts the path directly and never starts it, leaving
|
||||
# the on-access trigger INACTIVE -- so a dropped mount stayed dropped, which is
|
||||
# the exact failure this work exists to fix. Enable it explicitly.
|
||||
- name: enable cifs automount units
|
||||
become: true
|
||||
ansible.builtin.systemd:
|
||||
name: "{{ item }}"
|
||||
enabled: true
|
||||
state: started
|
||||
daemon_reload: true
|
||||
loop: "{{ cifs_watchdog_mounts | map('regex_replace', '^/', '') | map('regex_replace', '/', '-') | map('regex_replace', '$', '.automount') | list }}"
|
||||
failed_when: false
|
||||
|
||||
- import_tasks: podman/podman-check.yml
|
||||
vars:
|
||||
container_name: immich-machine-learning
|
||||
|
||||
Reference in New Issue
Block a user