ORT 1.26.0 changed provider loading to gate on the presence of required
nvidia pip packages before attempting to load libonnxruntime_providers_cuda.so.
Without nvidia-cuda-runtime-cu12, nvidia-cufft-cu12, and nvidia-curand-cu12
installed as Python packages, ORT silently skips the CUDA EP plugin entirely
(confirmed via /proc/maps: the .so was never dlopen'd despite existing on disk
and all system CUDA libs being present in ldconfig).
nvidia-nvjitlink-cu12 pulled in as a transitive dependency.
- executor: restore effective_count when replacement upload fails all retries
(delete succeeded but slot was never filled, leaving cap undercount)
- diversity: skip zero-norm embeddings before dedup/FPS selection
(InsightFace zeros pass dedup with similarity 0 and score distance 1.0,
getting selected first as maximally diverse)
- cli: exclude None from skip_ids in _smaller_duplicate_ids
(p.get('id') without None guard lets None into the set, silently
dropping every other id-less person from the processed list)
- embeddings: select face nearest crop centre instead of largest by area
(25% margin can pull a bigger neighbouring face into the crop;
largest-face selection then embeds the wrong person)
Round 11's replace_all missed two occurrences:
- _smaller_duplicate_ids inner comprehension (line 84): p["id"] →
p.get("id") so a named person with a missing "id" field does not
crash skip_ids computation before any return path is reached
- all-merges-failed fallback return (line 157): same fix; the outer
indentation prevented replace_all from matching this occurrence
The intentional p["id"] in merge_ids (line 119) is kept: that ID is
passed directly to merge_people() where None would be a caller bug,
not a silent data corruption.
- upload_tracker: revert data[flat_key] = [] from round 8; clearing the
entire shared legacy flat list on a corrupt value wipes all persons'
IDs, not just the one being reset; since a corrupt non-list value is
already unreadable by load_uploaded_ids, leaving it in place is safer
than a mass-wipe; update warning message to note the field is unaffected
but unreadable so the corruption is still observable
- cli: use p.get("id") instead of p["id"] in both people-list fallback
returns (_handle_duplicate_people lines 144 and 157) for consistency
with the success path at line 149; bare subscript crashes on malformed
unnamed persons that bypass _smaller_duplicate_ids
- immich_api: use 'or []' instead of .get("people", []) in get_people
so {"people": null} responses (some Immich versions with zero people
enrolled) return [] rather than None; .get() default only fires when
the key is absent, not when its value is null
- embeddings: log OSError from os.dup2 restore at DEBUG rather than
silently swallowing it; if a C extension (CUDA/onnxruntime) invalidates
the saved fd, the restore fails silently and stdout stays wired to
/dev/null — logging makes the event observable without changing the
swallow-and-continue semantics
- cache: remove MemoryError re-raise from EmbeddingCache.get(); a cache
read OOM aborted the entire diversity-selection batch for the person
rather than falling back to a fresh embedding computation, which is
the more appropriate OOM gate; broadening back to except Exception
restores the pre-round-5 fallback behavior
A KeyboardInterrupt raised inside the saved_out cleanup block would
propagate past the saved_err and devnull_fd blocks, leaking those fds.
In CPython this race is not realistically triggerable — KI is delivered
between bytecodes and os.dup2 is a single atomic C syscall — so we
accept the theoretical risk rather than silencing BaseException in a
finally block.
- executor: revert person_has_fscores=True back into try/except else
branch; moving it outside in round 8 was a regression — when the
tracker write fails on the first-ever upload (no prior frigate_scores
in tracker), setting the flag True prematurely switches at-cap
replacement into fscore mode, get_most_redundant_mapped_file returns
None (no entries), and all replacements are silently skipped;
the flag must only be set when the score is actually written
- diversity: remove dead face_crop-None guard; any face that passes
assess_quality (≥90 px MIN_FACE_WIDTH) produces a crop ≥135 px
(face + 25% margin), which is always above the 30 px crop minimum,
making the guard unreachable; _crop_face_from_thumbnail also calls
_get_face_bbox internally, so face_bbox is not None guarantees the
inner bbox check also passes
- diversity: fix hard_count regression from round 7 — revert to
'is not None and < 0.85' so only images that actually receive a
FPS boost (confirmed low confidence) are counted as hard examples;
None-confidence images use conf_array=1.0 (no boost) and should
not appear in the hard-example log count
- diversity: fix _scale_bbox_to_thumbnail to use explicit zero-guard
for imageWidth/imageHeight (meta_w or 0; scale = img_w/meta_w if
meta_w else 1.0) — mirrors image_processing.py pattern; prevents
`or img_w` from silently treating imageWidth=0 as missing and
returning scale=1.0 without surfacing the zero-metadata case
- embeddings: wrap all three os.close calls in _suppress_output
finally block with try/except OSError: pass so a failed close
in one branch cannot abort the outer finally and leak devnull_fd
or the saved_err/saved_out fds
- upload_tracker: clear corrupt flat-list key (data[flat_key] = [])
after the isinstance warning instead of leaving the corrupt value
in place — prevents stale IDs persisting across reset_person calls
and future load_uploaded_ids() from seeing a non-list value
- executor: move 'if pre_fscore is not None: person_has_fscores = True'
out of the try/except else branch so it fires even when mark_uploaded
raises; Frigate scores exist once measured regardless of tracker
write success, and replacement strategy should reflect that
- diversity: revert conf_array default from 0.5 back to 1.0 (np.ones);
the 0.5 default caused None-confidence images to receive a 1.7× FPS
boost and beat high-confidence detections — counter-productive for
Frigate training data quality
- diversity: fix hard_count to include None-confidence images (count
images where score is None or < 0.85, not only confirmed < 0.85);
the previous check systematically undercounted boosted images when the
Immich faces API omits the score field
- executor: fix garbled comment fragment "Skipped on / skipped when"
left by a partial edit in round 4; merge into a single coherent sentence
- executor: expand actually_uploaded trade-off comment to document all
three consequences of a tracker write failure (Frigate duplicate,
quality-replacement exclusion, cap-slot consumption), not only the
duplicate risk mentioned previously
- upload_tracker: add person_ids guard to reset_person isinstance check
so the non-list warning only fires when cleanup would actually have run,
not on no-op calls where person_ids is empty
- executor: snapshot has_frigate_model = effective_count > 0 before the
upload loop; use it in the recognize_face gate instead of the live
effective_count, which is incremented mid-loop and would otherwise
trigger recognize_face calls against an empty Frigate model on first run
- jobs: restore if already_uploaded > 0 guard before limit = capacity so
first-run auto-strategy jobs keep limit="auto" and the FPS adaptive
early-stop can fire instead of always filling MAX_AUTO_IMAGES slots
- cli: retry get_people() once after a post-merge empty response before
falling back to the pre-merge list; improve warning to name expired API
key as a possible cause alongside transient network errors
- diversity: hoist hard_weight = np.where(...) above the FPS while loop
since conf_array is constant; eliminates one O(n) numpy pass per
selected image
- embeddings: move os.close into try/finally so saved_out/saved_err are
always closed even when os.dup2 restore raises, preventing fd leak
- cache: replace narrow except tuple with except MemoryError: raise /
except Exception: return None so struct.error and other np.load failures
return None without masking OOM
- executor: fix first-run advisory message to check effective_count == 0
(post-stale-cleanup) instead of pre_run_count; remove now-unused
pre_run_count variable entirely
- jobs: remove dead "skip" entry from strategy_map (unreachable since the
early-return at the top of _resolve_strategy fires first)
- upload_tracker: log a warning when reset_person encounters a non-list
flat_key value instead of silently skipping the cleanup
- diversity: remove erroneous break outside if-faces in _scale_bbox_to_thumbnail
(broke people-loop early for first unannotated person, defeating scale fix)
- diversity: increment quality_filtered for face-too-small crop skips so the
summary log counts them alongside assess_quality failures
- diversity: fix hard-example log count to use original confidence_scores[i]
instead of synthetic conf_array default (0.5), eliminating false 100%
hard-example reports for persons with no Immich confidence data
- cache: add EOFError to except tuple in EmbeddingCache.get() so truncated
.npy files return None instead of crashing the embedding pipeline
- embeddings: wrap each os.dup2 restore in its own try/except OSError in
_suppress_output finally block so stderr is always restored even if the
stdout restore raises
- executor: gate recognize_face on effective_count > 0 (post-stale-cleanup)
instead of pre_run_count > 0 so recognize_face is not called against an
untrained Frigate model after the user manually deletes all training files
- executor: document actually_uploaded trade-off in comment (appending
unconditionally on tracker failure risks a Frigate duplicate but prevents
permanent filename unmapping which breaks quality-replacement scoring)
- jobs: check strategy == "skip" before the has_embedding and custom_limit
early-returns in _resolve_strategy so STRATEGY=skip is always honoured
- diversity: scale face bbox to thumbnail space before quality check so
check_face_size uses actual thumbnail pixels, not original-image coords
- diversity: skip asset when face bbox exists but crop guard rejects it,
preventing InsightFace from picking the wrong person in a group photo
- diversity: add _scale_bbox_to_thumbnail helper (extracted from crop logic)
- diversity: use set for medoid membership test in _kmedoids (O(n) not O(n*k))
- diversity: remove dead np.unique in _select_time_spread (linspace produces
strictly increasing indices; unique is a no-op and implies wrong semantics)
- embeddings: move os.open/os.dup calls inside try in _suppress_output so
EMFILE during setup does not leak already-allocated fds
- immich_api: count and log assets with missing/unparseable fileCreatedAt in
filter_recent_assets instead of silently discarding them
- executor: capture pre_run_count before stale-mapping cleanup so the
"first run" coaching message doesn't fire after manual file deletion
- cli: use p['id'] (KeyError-safe) instead of p.get('id') in fallback path
to match all other access sites on the same people list
- cache: narrow except to (OSError, ValueError) in EmbeddingCache.get so
MemoryError propagates instead of converting OOM to a silent cache miss
- immich_api: guard resp.json() with isinstance(dict) check in get_people and
fetch_all_assets so AttributeError doesn't escape on proxy/CDN non-dict responses
- executor: move actually_uploaded.append outside try/else so Frigate filename→asset_id
mapping is created via reconcile even when the tracker write fails
- cli: fall back to pre-merge people list when re-fetch after merge returns empty
(transient error) instead of silently dropping all people
- cli: treat ENABLE_FRIGATE_SCORES=false / BLUR_THRESHOLD=0 as not-set in
the unsupported-vars warning (falsy string check replaces raw truthiness)
- upload_tracker: guard set(data[flat_key]) with isinstance(list) check in
reset_person so a corrupted non-iterable legacy field doesn't crash mid-reset
- upload_tracker: guard dims[0]/dims[1] in find_by_crop_dimension with a
length check so a truncated crop_dims entry doesn't raise IndexError
- cache: wrap os.remove() in clear() with try/except OSError to handle
TOCTOU race with concurrent put() calls
- diversity: default conf_array to 0.5 (was 1.0) for faces with missing
confidence so they receive a moderate diversity boost instead of being
treated as high-confidence
- diversity: sort assets in the fast path (len <= limit) so return order is
consistent with the sorted-by-fileCreatedAt path
Correctness:
- jobs: cap auto-diversity limit for brand-new people (was never capped,
could exceed MAX_AUTO_IMAGES on first run)
- image_processing: separate None/0 guard for imageWidth/imageHeight so
missing field is explicit rather than silently aliased to img_w
- upload_tracker (_mark, update_frigate_count): copy-before-mutate so
exceptions between cache access and _save don't corrupt in-process state
- jobs: reject LIMIT=0 on no-embedding path (was silently empty run)
- jobs: add STRATEGY=skip to strategy_map so env var is honoured
- embeddings: convert to RGB before cvtColor so RGBA/grayscale thumbnails
don't raise cv2.error and silently drop from diversity selection
- config: use falsy guard for OUTPUT_DIR so blank env var falls through
to config file value
- reconcile: _ts() returns float("inf") on parse failure so unrecognised
filenames sort last instead of collapsing to 0.0 and corrupting FIFO mapping
- diversity: remove dead selected_set (never read; -np.inf sentinel already
prevents re-selection)
Lint (ruff):
- executor: sort upload_tracker import block (I001)
- executor: replace lambda is_better_than with operator.lt/gt (E731 x2)
- executor, upload_tracker: wrap long logger.warning calls (E501 x4)
- record_frigate_files_batch: copy-before-mutate so a write failure
doesn't leave cache ahead of disk (same fix as remove_frigate_files_batch)
- executor: replace tracker_ok boolean with try/else
- jobs: collapse duplicate custom_limit is not None checks into one guard
Bump version to 0.6.3.
- Drop flat list as primary storage; derive uploaded/rejected IDs from by_person
(single source of truth). Legacy flat lists in existing files still read for
backward compat. Removes dual-representation sync hazard.
- Add begin_batch/flush_batch: per-person upload loop now does 1 os.replace
instead of N (one per mark_uploaded call). Benefit on slow storage.
- reset_all_people(): RESET_PERSON=* is now O(1) disk writes instead of O(P^2).
- blur_score_from_image inlines cv2.Laplacian directly, removing assess_quality
call overhead and decoupling from the full quality pipeline.
* revert: replace SQLite tracker with JSON backend (v0.6.0)
The SQLite migration (v0.5.0) spawned 21 bug-fix releases in two days:
data-loss risk in the migration layer, schema PK conflicts on per-person
tracking, tracker isolation races under concurrent runs, and a disk-full
error that triggered duplicate Frigate uploads. The complexity cost
outweighs the benefit.
Restored the pre-SQL JSON tracker (frigate_uploaded_ids.json /
frigate_rejected_ids.json in DATA_DIR). Public API is identical — all
callers in executor.py, jobs.py, cli.py, and reconcile.py work unchanged.
Existing JSON files are read automatically; frigate_tracker.db can be
deleted once verified.
* fix: narrow corrupt-thumbnail exception to UnidentifiedImageError; restore IMMICH_URL empty-string fallback
* docs: rewrite v0.6.0 changelog, strip v0.5.x entries, fix README SQLite references
* chore: remove dead get_frigate_filename_for_asset (orphaned since FRIGATE_SCORE_THRESHOLD removal in v0.4.0)
* fix: sort imports in executor.py (ruff I001)
- diversity: cap k-medoids seed count at target so _cluster_aware_selection
never returns more images than requested (violated MAX_AUTO_IMAGES when
remaining capacity was 1-4 slots); add early return for limit=0 to
prevent k-medoids from running with a zero budget; slice return to target
as a final guard
- executor: mark_rejected() when fetch_full_image returns None so assets
that can't be fetched (both original and preview) aren't retried every run
- executor: wrap mark_uploaded() in its own try/except so a SQLite disk-full
error after a successful HTTP 200 doesn't retry the Frigate POST (duplicate
upload) — the upload succeeded; only the tracker write failed
- cli: apply skip_ids deduplication to the re-fetched people list after a
partial merge (some groups succeed, some fail) so unmerged duplicates
don't produce two jobs for the same Frigate folder
- jobs: strip whitespace from SKIP_PEOPLE/ONLY_PEOPLE elements on split
so "Alice, Bob" (space after comma) correctly matches "Bob"
- upload_tracker: replace executescript() in _migrate_schema_v2 with
individual execute() calls inside a transaction so a crash between DROP
and RENAME rolls back instead of permanently destroying tracked_assets
- frigate_api: _get_frigate_url now strips leading/trailing whitespace
before rstrip('/') so whitespace-only FRIGATE_URL is treated as unset
- executor: upload_to_frigate now uses _get_frigate_url() eliminating
double-slash upload paths when FRIGATE_URL has a trailing slash
- executor: corrupt thumbnail (resp.ok=True, Image.open fails) now calls
mark_rejected() so permanently broken assets are not retried forever
- upload_tracker: reset_person now uses _get_frigate_url() instead of
inline os.environ.get('FRIGATE_URL', '').strip()
- image_processing: _save_jpeg writes to a .tmp file and calls
os.replace() so a disk-full error never leaves a truncated JPEG
- cli: _handle_duplicate_people falls back to local deduplication when
all Immich merges fail, preventing two jobs from overwriting the same
Frigate folder
- config: _getenv_optional_float now delegates to _getenv_num() like
_getenv_optional_int, eliminating the inconsistent duplicate
- reconcile: _ts() uses rsplit('.', 1)[0] instead of .replace('.webp','')
so FIFO mapping works with any Frigate training-file extension
Correctness:
- fetch_face_data: only fall back to faces[0] when person_id is absent;
previously a missing person match injected a different person's bbox
- upload_tracker: change PK from (asset_id, status) to
(asset_id, person_name, status); old PK allowed INSERT OR REPLACE to
silently overwrite person_name when the same photo appeared in two
people's jobs, breaking quality-replacement JOINs; auto-migrates DBs
- filter_recent_assets: treat years=0 as "no age filter" instead of
falling through to Config.YEARS_FILTER via falsy `or`
- _is_module_available: return find_spec(...) is not None; find_spec
returns None (not raises) for absent top-level modules, so the
previous code always returned True
- execute_jobs error handler: use asset.get("id", "<unknown>") to avoid
a secondary KeyError propagating out of execute_jobs on malformed dicts
- upload_to_frigate: also mark_rejected on HTTP 422, not only HTTP 400
with "face" in body; other permanent errors left assets untracked and
retried forever
- reconcile_frigate_mappings: sort key lambda f: (_ts(f), f) makes order
deterministic when timestamps are equal or 0.0; set iteration order is
hash-randomised, stable sort preserves it
Reuse / cleanup:
- config.py: add _getenv_optional_int delegating to _getenv_num(name, None, int)
- jobs.py: _resolve_strategy uses _getenv_optional_int("LIMIT") instead
of inline os.environ.get + int() + warning duplicate of _getenv_num
- frigate_api.py: add _get_frigate_url() helper; eliminates 4× copy of
os.environ.get("FRIGATE_URL", "").rstrip("/")
- quality.py: extract blur_score_from_image(img, max_dim=1440) helper;
executor.py time-spread blur fallback now uses it instead of inlining
the resize+RGB+assess_quality sequence, keeping scale logic in one place
scripts/benchmark.py retained the old os.getenv inline pattern after
_getenv_bool was introduced in v0.5.16. Now uses a deferred local
import of _getenv_bool, consistent with the script's pattern of keeping
all winnow imports inside function bodies rather than at the top level.
- _getenv_num: add raw.strip() + empty-string guard so numeric vars set
to "" (common Compose pattern for "use default") return the default
silently instead of warning "not a valid int/float"
- _getenv_bool: same guard so True-defaulted flags set to "" return
the configured default instead of silently returning False
- embeddings.py: replace inline FORCE_CPU bool parse with _getenv_bool
- cache.py: fix np.save extension bug from v0.5.13 — tmp path used
final+".tmp" (abc.npy.tmp) but np.save auto-appends .npy to paths not
ending in .npy, writing to abc.npy.tmp.npy instead; os.replace then
raised FileNotFoundError silently, making every cache write a no-op
and leaking *.npy.tmp.npy files. Fixed by inserting .tmp before .npy:
tmp = final[:-4] + ".tmp.npy"
- config.py: remove str(default) round-trip in _getenv_int/_getenv_float
— use raw = os.getenv(name); return default if raw is None else int(raw)
so a future float default can't cause a spurious "not a valid integer"
warning and return the wrong type
- executor.py: consolidate 4 progress.remove_task calls into one
try/finally around the per-job body; continue inside try/finally
executes the finally before the next iteration, making the invariant
structurally enforced rather than relying on discipline across 4 sites
YEARS_FILTER, MIN_FACE_WIDTH, MIN_FACE_COUNT, MAX_AUTO_IMAGES,
BLUR_THRESHOLD, MIN_CONFIDENCE, and FACE_MARGIN used bare int()/float()
calls with no error handler. A typo (trailing space, non-numeric value)
raised ValueError inside __getattr__, producing a cryptic traceback on
the first config access rather than at the validate() step. Values are
now parsed by _getenv_int/_getenv_float helpers that warn and fall back
to the documented default, matching the existing FRIGATE_SCORE_CEILING
pattern.
- executor.py: wrap shutil.rmtree/os.makedirs in try/except OSError so a
permission failure logs and skips the job rather than aborting the run
- cache.py: write embeddings to a .tmp file and atomically rename into place
via os.replace so a process kill can't leave a corrupted .npy cache slot
- immich_api.py: guard fileCreatedAt with isinstance(str) check before calling
.replace() so a non-string timestamp doesn't raise AttributeError and kill
the entire filter_recent_assets pass
- upload_tracker.py: raise SQLite busy timeout from 5 s to 30 s to handle
concurrent cron+manual run overlap without dropping upload-tracking records
progress.add_task() fires unconditionally at the top of the job loop;
both continue paths (ValueError from _safe_person_dir and the symlink
TOCTOU guard) skipped remove_task(), leaving orphaned 0% rows in the
terminal for the rest of the run.
The v0.5.10 compound guard 'isdir and not islink' silently skipped the
rmtree when person_dir was a symlink-to-directory, then let makedirs
follow the symlink — allowing crop writes outside output_dir with no
diagnostic. Replace with an explicit islink pre-check that logs an error
and continues, matching the ValueError path from _safe_person_dir.
- reconcile.py: re-escalate the < target branch from INFO to WARNING and
add 'permanently unmapped' label. Both post-loop branches produce identical
permanent mapping loss; v0.5.9 incorrectly treated the timeout case as
recoverable.
- executor.py: guard shutil.rmtree with 'not os.path.islink(person_dir)'
so a race-replaced symlink-to-directory is skipped rather than raising
an unhandled OSError that aborts all remaining jobs. Correct comment:
rmtree raises OSError, not NotADirectoryError.
- reconcile.py: swap log levels — external-upload path (permanent mapping
loss) escalated to WARNING; timeout path (transient, retries next cycle)
downgraded to INFO. Also extend the warning message to note the files are
permanently unmapped.
- immich_api.py: extend fetch_all_assets docstring to document that
all-garbage page termination (in addition to network errors) makes
total_raw a lower bound.
- executor.py: add comment above shutil.rmtree noting that POSIX rmtree
raises NotADirectoryError on a top-level symlink, documenting why the
removed islink guard is safe to omit.
- upload_tracker.py: replace setdefault with explicit guard in _entry() —
setdefault evaluates its default-dict argument before checking key
presence, allocating and discarding a dict on every already-present call.
- executor.py: fix _safe_person_dir docstring — realpath+startswith is the
load-bearing traversal guard; islink is a supplementary early-exit for the
symlink sub-case only. The previous comment "checking after realpath would be
too late" implied islink was the primary guard, which is backwards.
- immich_api.py: move total_raw accumulation to after the dead-end-page break
so all-garbage pages don't inflate the count and produce misleading
"N total, 0 recent" output. Mixed pages (some valid, some non-dict) still
count page_count so transient schema issues don't shrink MIN_FACE_COUNT below
threshold. Add warning when a RequestException interrupts pagination mid-way
so operators know total_raw is a lower bound.
- config.py: eliminate residual TOCTOU — change `if config_file.exists():` to
`if _data_cfg_exists or config_file.exists():` so _data_cfg is never
stat'd twice (the v0.5.7 fix cached the first check but not the second).