- frigate_api._get_faces_data now validates the /api/faces response is a
dict before returning it, so a malformed body no longer crashes
data.items() in get_all_frigate_person_files (and the equivalent
data.get() in get_frigate_person_files).
- upload_tracker now takes an exclusive flock on a DATA_DIR lock file for
every tracker load-mutate-save cycle, so a scheduled run and a manual
docker exec against the same DATA_DIR can no longer race a
read-modify-write and silently drop the loser's marks.
- begin_batch()/flush_batch() now flush to disk every 10 marks instead of
deferring the whole per-person upload loop, bounding how many uploaded
marks a crash mid-batch can lose.
data/ is the DATA_DIR the README tells you to mount, and it was not ignored: it holds 1,593
face-embedding .npy files plus upload trackers keyed by person name. A single 'git add -A'
would have published them. Never committed to date, so history is clean.
- immich_api: .get("items") or [] handles {"items": null} without crashing len()
- immich_api: catch TypeError alongside ValueError in filter_recent_assets for
timezone-naive fileCreatedAt comparisons
- scheduler: catch SystemExit in addition to KeyboardInterrupt so cli.main()
cannot kill the long-running scheduler process
- scheduler: reseed croniter from wall-clock time after each run so overrunning
jobs don't schedule an immediate back-to-back rerun
- config: reject negative YEARS_FILTER values with a warning, reset to default 10
- frigate_api: return True (not False) for empty filenames list — callers cannot
distinguish no-op from network failure on False
- jobs: warn on unrecognised STRATEGY value instead of silently falling back
- jobs: casefold ONLY_PEOPLE / SKIP_PEOPLE matching so "john doe" matches "John Doe"
- executor: <= → < so a same-score candidate can fill a freed replacement slot
- embeddings: set _insightface_loaded=True on GPU+CPU double-failure to prevent
N re-init attempts (one per asset) when InsightFace is broken for a whole run
Use full_body for the 500 'could not process' permanent-rejection check,
consistent with the 400 'face' check on the line above. error_detail is
truncated to 100 chars via the fallback path, which could silently miss
the phrase in a long response body.
Remove | None from process_face_mode return type — every code path returns
tuple[int,int] or str; None is unreachable. Update docstring to match.
1. if saved: → if isinstance(saved, tuple): so string skip-reasons from
process_face_mode no longer register as successes and create phantom
asset_map entries with no JPEG on disk. Dead reason/fallback code in
the else branch now correctly handles str and None returns.
2. "could not process" permanent-rejection check now uses error_detail
(json message field, falling back to body[:100]) instead of full_body,
keeping the match consistent with what is displayed to the user.
3. CHANGELOG [Unreleased] breaking-change note for MAX_AUTO_IMAGES 20→5
so upgrading users know to set the env var if they want the old cap.
Smaller default cap is more conservative for new installs and better
reflects the minimum viable training set for Frigate face recognition.
Users who need more can set MAX_AUTO_IMAGES explicitly.
process_face_mode now returns a descriptive string instead of None for
filtered-out faces ("face too small 45x38px, min 90px", "no face
metadata"), so the executor can print a useful reason rather than the
generic "no usable face data".
Also suppresses the InsightFace norm_crop FutureWarning about deprecated
estimate usage, which was noisy at INFO level on every aligned crop.
Frigate returns HTTP 500 with 'Could not process' when its face detector
cannot find or embed a face in the uploaded crop — this will never succeed
on retry. Previously these were silently logged at DEBUG and retried on
every future run.
- Surface 500 error details inline (same display path as 400)
- Mark 500 + 'could not process' as a permanent rejection so the asset
is skipped on future runs instead of retried indefinitely
dev is a protected branch requiring PRs. The previous direct push caused
the lockfile update CI job to fail with 'protected branch hook declined'.
When the triggering branch is dev, the workflow now creates a side branch
and opens a PR; all other branches continue to push directly.
insightface depends on onnxruntime (CPU) as a direct dependency. During
uv sync --extra gpu, both onnxruntime (CPU, 24.6 MB binary) and
onnxruntime-gpu (GPU, 24.7 MB binary) are installed in parallel — both
claim onnxruntime/capi/onnxruntime_pybind11_state.so. The last writer wins,
which is non-deterministic in uv's parallel installer.
On GitHub Actions (no GPU, different scheduler ordering), the CPU binary
consistently wins, leaving onnxruntime-gpu's pybind11_state.so as the CPU
version. CUDAExecutionProvider then silently disappears because the CPU
binary's provider registration code has no CUDA EP.
Fix: after uv sync, reinstall onnxruntime-gpu explicitly using the already-
cached wheel. Since uv pip install is synchronous and runs after the parallel
sync completes, the GPU binary is guaranteed to be on disk when the build
layer commits.
cli.py:
- _smaller_duplicate_ids: walrus operator eliminates double p.get("id")
per element; truthiness check replaces dead "is not None" guard (all
persons in by_name are guaranteed to have a truthy id after the
line-75 gate)
- Extract _excl() helper inside _handle_duplicate_people — replaces 4
identical [p for p in lst if p.get("id") not in skip_ids] expressions
across all return paths
jobs.py:
- Extract _valid_people() — shared filter for interactive_configure and
auto_configure; uses (p.get("name") or "").strip() to match cli.py's
whitespace-strip gate, preventing whitespace-only Immich names from
reaching _build_job and creating blank Frigate person labels
- Hoist queued_ids set before the display loop in interactive_configure:
O(N) set lookup per render instead of O(N×|jobs|) linear scan
jobs.py:
- Add p.get("id") guard to valid_people filter in both
interactive_configure and auto_configure — id-less named persons
passed through by _handle_duplicate_people are now excluded before
any bare-subscript access in the configure paths
- Fix bare p["id"] → p.get("id") in the queued-marker check at line 214
(runs unconditionally on all valid_people during menu display, before
any user selection or fetch_all_assets guard)
cli.py:
- Remove dead-code survivor_id and merge_ids guards: after the by_name
fix (line 75 requires p.get("id")), all persons in any ordered list
have ids, so neither guard can ever fire; removing them prevents
misleading readers about what states are reachable
cli.py:
- Filter id-less persons from by_name at construction (root fix for all
bare-subscript crashes downstream — persons with a name but no id are
excluded from duplicate detection entirely)
- Belt-and-suspenders on warning-path display: p['id'] → p.get('id')
- Extract survivor_id with .get(); skip group if survivor has no id
- Guard merge_ids: skip API call when list is empty after id filtering
- Walrus operator in merge_ids comprehension: p.get("id") called once
per item instead of twice
executor.py:
- Add cross-reference comment at success-path reset so the for/else
rollback pairing is explicit for future maintainers
- executor.py: clear min_quality_score_for_slot alongside effective_count
restore in for/else block; leaving the stale floor from the deleted
file's score blocked the next candidate from filling the restored slot
- cli.py: guard merge_ids with p.get('id') is not None, consistent with
the _smaller_duplicate_ids fix; bare p['id'] raised KeyError on any
person dict missing the id field in the auto-merge path
- immich_api.py: replace bare data['major'/'minor'/'patch'] subscripts
with .get() in get_immich_version; KeyError was silently swallowed by
except Exception, causing version-gated flags to disable without warning
ORT 1.26.0 changed provider loading to gate on the presence of required
nvidia pip packages before attempting to load libonnxruntime_providers_cuda.so.
Without nvidia-cuda-runtime-cu12, nvidia-cufft-cu12, and nvidia-curand-cu12
installed as Python packages, ORT silently skips the CUDA EP plugin entirely
(confirmed via /proc/maps: the .so was never dlopen'd despite existing on disk
and all system CUDA libs being present in ldconfig).
nvidia-nvjitlink-cu12 pulled in as a transitive dependency.
- executor: restore effective_count when replacement upload fails all retries
(delete succeeded but slot was never filled, leaving cap undercount)
- diversity: skip zero-norm embeddings before dedup/FPS selection
(InsightFace zeros pass dedup with similarity 0 and score distance 1.0,
getting selected first as maximally diverse)
- cli: exclude None from skip_ids in _smaller_duplicate_ids
(p.get('id') without None guard lets None into the set, silently
dropping every other id-less person from the processed list)
- embeddings: select face nearest crop centre instead of largest by area
(25% margin can pull a bigger neighbouring face into the crop;
largest-face selection then embeds the wrong person)