Compare commits

...
Author SHA1 Message Date
dependabot[bot] 60fb7be55e chore(deps): bump the python-deps group across 1 directory with 5 updates
Bumps the python-deps group with 5 updates in the / directory:

| Package | From | To |
| --- | --- | --- |
| [croniter](https://github.com/pallets-eco/croniter) | `6.2.2` | `6.2.3` |
| [numpy](https://github.com/numpy/numpy) | `2.5.0` | `2.5.1` |
| [opencv-python-headless](https://github.com/opencv/opencv-python) | `4.13.0.92` | `5.0.0.93` |
| [pillow](https://github.com/python-pillow/Pillow) | `12.2.0` | `12.3.0` |
| [nvidia-cudnn-cu12](https://developer.nvidia.com/cuda-zone) | `9.23.2.1` | `9.24.0.43` |



Updates `croniter` from 6.2.2 to 6.2.3
- [Release notes](https://github.com/pallets-eco/croniter/releases)
- [Changelog](https://github.com/pallets-eco/croniter/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pallets-eco/croniter/compare/6.2.2...6.2.3)

Updates `numpy` from 2.5.0 to 2.5.1
- [Release notes](https://github.com/numpy/numpy/releases)
- [Changelog](https://github.com/numpy/numpy/blob/main/doc/RELEASE_WALKTHROUGH.rst)
- [Commits](https://github.com/numpy/numpy/compare/v2.5.0...v2.5.1)

Updates `opencv-python-headless` from 4.13.0.92 to 5.0.0.93
- [Release notes](https://github.com/opencv/opencv-python/releases)
- [Commits](https://github.com/opencv/opencv-python/commits)

Updates `pillow` from 12.2.0 to 12.3.0
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.2.0...12.3.0)

Updates `nvidia-cudnn-cu12` from 9.23.2.1 to 9.24.0.43

---
updated-dependencies:
- dependency-name: croniter
  dependency-version: 6.2.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: python-deps
- dependency-name: numpy
  dependency-version: 2.5.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: python-deps
- dependency-name: opencv-python-headless
  dependency-version: 5.0.0.93
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: python-deps
- dependency-name: pillow
  dependency-version: 12.3.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: python-deps
- dependency-name: nvidia-cudnn-cu12
  dependency-version: 9.24.0.43
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: python-deps
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-09 18:55:44 +00:00
flan eba50f8a4a Merge pull request #48 from sudolulo/dependabot/docker/nvidia/cuda-12.9.2-cudnn-runtime-ubuntu24.04
chore(deps): bump nvidia/cuda from 12.8.1-cudnn-runtime-ubuntu24.04 to 12.9.2-cudnn-runtime-ubuntu24.04
2026-06-26 18:14:53 -04:00
flan 95d5e91a67 Merge pull request #49 from sudolulo/dependabot/uv/python-deps-c7af4feaef
chore(deps): bump the python-deps group with 3 updates
2026-06-26 18:14:22 -04:00
dependabot[bot] 9e904d9f34 chore(deps): bump the python-deps group with 3 updates
Bumps the python-deps group with 3 updates: [numpy](https://github.com/numpy/numpy), [pytest](https://github.com/pytest-dev/pytest) and [ruff](https://github.com/astral-sh/ruff).


Updates `numpy` from 2.4.6 to 2.5.0
- [Release notes](https://github.com/numpy/numpy/releases)
- [Changelog](https://github.com/numpy/numpy/blob/main/doc/RELEASE_WALKTHROUGH.rst)
- [Commits](https://github.com/numpy/numpy/compare/v2.4.6...v2.5.0)

Updates `pytest` from 9.1.0 to 9.1.1
- [Release notes](https://github.com/pytest-dev/pytest/releases)
- [Changelog](https://github.com/pytest-dev/pytest/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pytest-dev/pytest/compare/9.1.0...9.1.1)

Updates `ruff` from 0.15.18 to 0.15.20
- [Release notes](https://github.com/astral-sh/ruff/releases)
- [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md)
- [Commits](https://github.com/astral-sh/ruff/compare/0.15.18...0.15.20)

---
updated-dependencies:
- dependency-name: numpy
  dependency-version: 2.5.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: python-deps
- dependency-name: pytest
  dependency-version: 9.1.1
  dependency-type: direct:development
  update-type: version-update:semver-patch
  dependency-group: python-deps
- dependency-name: ruff
  dependency-version: 0.15.20
  dependency-type: direct:development
  update-type: version-update:semver-patch
  dependency-group: python-deps
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-06-25 18:55:02 +00:00
dependabot[bot] 18305b5b06 chore(deps): bump nvidia/cuda
Bumps nvidia/cuda from 12.8.1-cudnn-runtime-ubuntu24.04 to 12.9.2-cudnn-runtime-ubuntu24.04.

---
updated-dependencies:
- dependency-name: nvidia/cuda
  dependency-version: 12.9.2-cudnn-runtime-ubuntu24.04
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-06-25 18:53:34 +00:00
flan 9b68f62f69 Merge pull request #44 from sudolulo/dependabot/github_actions/actions/checkout-7.0.0
chore(deps): bump actions/checkout from 6.0.3 to 7.0.0
2026-06-20 15:06:54 -04:00
flan b5d03a2695 Merge branch 'main' into dependabot/github_actions/actions/checkout-7.0.0 2026-06-20 15:05:37 -04:00
flan fce46d409e Merge pull request #45 from sudolulo/dependabot/uv/python-deps-ae41630b22
chore(deps): bump the python-deps group across 1 directory with 5 updates
2026-06-20 15:05:13 -04:00
dependabot[bot] 5044343a99 chore(deps): bump the python-deps group across 1 directory with 5 updates
Bumps the python-deps group with 5 updates in the / directory:

| Package | From | To |
| --- | --- | --- |
| [onnxruntime-gpu](https://github.com/microsoft/onnxruntime) | `1.26.0` | `1.27.0` |
| [nvidia-cudnn-cu12](https://developer.nvidia.com/cuda-zone) | `9.23.1.3` | `9.23.2.1` |
| [onnxruntime](https://github.com/microsoft/onnxruntime) | `1.26.0` | `1.27.0` |
| [pytest](https://github.com/pytest-dev/pytest) | `9.0.3` | `9.1.0` |
| [ruff](https://github.com/astral-sh/ruff) | `0.15.17` | `0.15.18` |



Updates `onnxruntime-gpu` from 1.26.0 to 1.27.0
- [Release notes](https://github.com/microsoft/onnxruntime/releases)
- [Changelog](https://github.com/microsoft/onnxruntime/blob/main/docs/ReleaseManagement.md)
- [Commits](https://github.com/microsoft/onnxruntime/commits)

Updates `nvidia-cudnn-cu12` from 9.23.1.3 to 9.23.2.1

Updates `onnxruntime` from 1.26.0 to 1.27.0
- [Release notes](https://github.com/microsoft/onnxruntime/releases)
- [Changelog](https://github.com/microsoft/onnxruntime/blob/main/docs/ReleaseManagement.md)
- [Commits](https://github.com/microsoft/onnxruntime/commits)

Updates `pytest` from 9.0.3 to 9.1.0
- [Release notes](https://github.com/pytest-dev/pytest/releases)
- [Changelog](https://github.com/pytest-dev/pytest/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pytest-dev/pytest/compare/9.0.3...9.1.0)

Updates `ruff` from 0.15.17 to 0.15.18
- [Release notes](https://github.com/astral-sh/ruff/releases)
- [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md)
- [Commits](https://github.com/astral-sh/ruff/compare/0.15.17...0.15.18)

---
updated-dependencies:
- dependency-name: nvidia-cudnn-cu12
  dependency-version: 9.23.2.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: python-deps
- dependency-name: onnxruntime
  dependency-version: 1.27.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: python-deps
- dependency-name: onnxruntime-gpu
  dependency-version: 1.27.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: python-deps
- dependency-name: pytest
  dependency-version: 9.1.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
  dependency-group: python-deps
- dependency-name: ruff
  dependency-version: 0.15.18
  dependency-type: direct:development
  update-type: version-update:semver-patch
  dependency-group: python-deps
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-06-18 20:42:03 +00:00
dependabot[bot] 2a673fcd9d chore(deps): bump actions/checkout from 6.0.3 to 7.0.0
Bumps [actions/checkout](https://github.com/actions/checkout) from 6.0.3 to 7.0.0.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/df4cb1c069e1874edd31b4311f1884172cec0e10...9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-06-18 20:40:32 +00:00
flan 04151057c9 Merge pull request #46 from sudolulo/dev
release: v0.6.6
2026-06-18 16:38:28 -04:00
flan 3ea6b9d566 Merge pull request #47 from sudolulo/chore/simplify-lockfile-ci
chore: replace lockfile auto-update with uv lock --check
2026-06-18 16:36:10 -04:00
flan e281ac01e1 chore: add pre-commit hook to auto-update lockfile on pyproject.toml changes 2026-06-18 20:33:58 +00:00
flan 0ac02c3108 chore: update lockfile for v0.6.6 2026-06-18 20:31:36 +00:00
flan db30449b09 chore: replace lockfile auto-update with uv lock --check 2026-06-18 20:29:43 +00:00
flan 1e4425bfd7 release: v0.6.6 2026-06-18 20:23:58 +00:00
flan 2be9f400dd Merge pull request #42 from sudolulo/fix/frigate-500-permanent-rejection
fix: Frigate 500 permanent rejection + 13 correctness bugs
2026-06-17 15:17:20 -04:00
flan f3b5bd9334 fix: update MAX_AUTO_IMAGES default assertion to 5 2026-06-17 19:16:43 +00:00
flan 0906342edf fix: 10 correctness bugs from full codebase audit
- immich_api: .get("items") or [] handles {"items": null} without crashing len()
- immich_api: catch TypeError alongside ValueError in filter_recent_assets for
  timezone-naive fileCreatedAt comparisons
- scheduler: catch SystemExit in addition to KeyboardInterrupt so cli.main()
  cannot kill the long-running scheduler process
- scheduler: reseed croniter from wall-clock time after each run so overrunning
  jobs don't schedule an immediate back-to-back rerun
- config: reject negative YEARS_FILTER values with a warning, reset to default 10
- frigate_api: return True (not False) for empty filenames list — callers cannot
  distinguish no-op from network failure on False
- jobs: warn on unrecognised STRATEGY value instead of silently falling back
- jobs: casefold ONLY_PEOPLE / SKIP_PEOPLE matching so "john doe" matches "John Doe"
- executor: <= → < so a same-score candidate can fill a freed replacement slot
- embeddings: set _insightface_loaded=True on GPU+CPU double-failure to prevent
  N re-init attempts (one per asset) when InsightFace is broken for a whole run
2026-06-17 19:12:58 +00:00
flan 5e0a871314 fix: 2 low findings from audit — consistent 500 match source, accurate return type
Use full_body for the 500 'could not process' permanent-rejection check,
consistent with the 400 'face' check on the line above. error_detail is
truncated to 100 chars via the fallback path, which could silently miss
the phrase in a long response body.

Remove | None from process_face_mode return type — every code path returns
tuple[int,int] or str; None is unreachable. Update docstring to match.
2026-06-17 17:47:31 +00:00
flan 86caffc8d7 fix: 3 audit findings — truthy skip-reason bug, 500 match consistency, CHANGELOG note
1. if saved: → if isinstance(saved, tuple): so string skip-reasons from
   process_face_mode no longer register as successes and create phantom
   asset_map entries with no JPEG on disk. Dead reason/fallback code in
   the else branch now correctly handles str and None returns.

2. "could not process" permanent-rejection check now uses error_detail
   (json message field, falling back to body[:100]) instead of full_body,
   keeping the match consistent with what is displayed to the user.

3. CHANGELOG [Unreleased] breaking-change note for MAX_AUTO_IMAGES 20→5
   so upgrading users know to set the env var if they want the old cap.
2026-06-17 17:31:20 +00:00
flan f6e494071e fix: lower MAX_AUTO_IMAGES default from 20 to 5
Smaller default cap is more conservative for new installs and better
reflects the minimum viable training set for Frigate face recognition.
Users who need more can set MAX_AUTO_IMAGES explicitly.
2026-06-17 17:19:09 +00:00
flan b69f776378 fix: surface skip reasons and suppress norm_crop FutureWarning
process_face_mode now returns a descriptive string instead of None for
filtered-out faces ("face too small 45x38px, min 90px", "no face
metadata"), so the executor can print a useful reason rather than the
generic "no usable face data".

Also suppresses the InsightFace norm_crop FutureWarning about deprecated
estimate usage, which was noisy at INFO level on every aligned crop.
2026-06-17 17:18:43 +00:00
flan c4910d82ef fix: treat Frigate 500 'Could not process' as permanent rejection
Frigate returns HTTP 500 with 'Could not process' when its face detector
cannot find or embed a face in the uploaded crop — this will never succeed
on retry. Previously these were silently logged at DEBUG and retried on
every future run.

- Surface 500 error details inline (same display path as 400)
- Mark 500 + 'could not process' as a permanent rejection so the asset
  is skipped on future runs instead of retried indefinitely
2026-06-17 17:16:04 +00:00
flan 5b8b3ab736 Merge pull request #41 from sudolulo/dev
fix: lockfile workflow opens PR on dev instead of direct push
2026-06-16 23:23:44 -04:00
flan dc0c431e6b chore: update lockfile 2026-06-17 03:12:20 +00:00
flan 3296940806 fix: lockfile workflow opens PR on dev instead of direct push
dev is a protected branch requiring PRs. The previous direct push caused
the lockfile update CI job to fail with 'protected branch hook declined'.
When the triggering branch is dev, the workflow now creates a side branch
and opens a PR; all other branches continue to push directly.
2026-06-17 03:12:00 +00:00
flan fdcb4efac3 Merge pull request #40 from sudolulo/dev
release: v0.6.5
2026-06-16 23:08:08 -04:00
flan c1f04be15b release: v0.6.5 2026-06-17 03:06:27 +00:00
flan 001dd2c575 fix: reinstall onnxruntime-gpu after uv sync to guarantee GPU binary wins
insightface depends on onnxruntime (CPU) as a direct dependency. During
uv sync --extra gpu, both onnxruntime (CPU, 24.6 MB binary) and
onnxruntime-gpu (GPU, 24.7 MB binary) are installed in parallel — both
claim onnxruntime/capi/onnxruntime_pybind11_state.so. The last writer wins,
which is non-deterministic in uv's parallel installer.

On GitHub Actions (no GPU, different scheduler ordering), the CPU binary
consistently wins, leaving onnxruntime-gpu's pybind11_state.so as the CPU
version. CUDAExecutionProvider then silently disappears because the CPU
binary's provider registration code has no CUDA EP.

Fix: after uv sync, reinstall onnxruntime-gpu explicitly using the already-
cached wheel. Since uv pip install is synchronous and runs after the parallel
sync completes, the GPU binary is guaranteed to be on disk when the build
layer commits.
2026-06-17 03:05:33 +00:00
flan edf576bc93 Merge pull request #39 from sudolulo/fix/codebase-audit-r2
fix: codebase audit rounds 2-5 (correctness, cli, jobs)
2026-06-16 23:02:42 -04:00
flan 561a1a3d72 fix: 5 findings from codebase audit round 5
cli.py:
- _smaller_duplicate_ids: walrus operator eliminates double p.get("id")
  per element; truthiness check replaces dead "is not None" guard (all
  persons in by_name are guaranteed to have a truthy id after the
  line-75 gate)
- Extract _excl() helper inside _handle_duplicate_people — replaces 4
  identical [p for p in lst if p.get("id") not in skip_ids] expressions
  across all return paths

jobs.py:
- Extract _valid_people() — shared filter for interactive_configure and
  auto_configure; uses (p.get("name") or "").strip() to match cli.py's
  whitespace-strip gate, preventing whitespace-only Immich names from
  reaching _build_job and creating blank Frigate person labels
- Hoist queued_ids set before the display loop in interactive_configure:
  O(N) set lookup per render instead of O(N×|jobs|) linear scan
2026-06-17 02:28:58 +00:00
flan 614542decd fix: 4 findings from codebase audit round 4
jobs.py:
- Add p.get("id") guard to valid_people filter in both
  interactive_configure and auto_configure — id-less named persons
  passed through by _handle_duplicate_people are now excluded before
  any bare-subscript access in the configure paths
- Fix bare p["id"] → p.get("id") in the queued-marker check at line 214
  (runs unconditionally on all valid_people during menu display, before
  any user selection or fetch_all_assets guard)

cli.py:
- Remove dead-code survivor_id and merge_ids guards: after the by_name
  fix (line 75 requires p.get("id")), all persons in any ordered list
  have ids, so neither guard can ever fire; removing them prevents
  misleading readers about what states are reachable
2026-06-17 02:20:38 +00:00
flan 3fccf9c8f9 fix: 6 findings from codebase audit round 3
cli.py:
- Filter id-less persons from by_name at construction (root fix for all
  bare-subscript crashes downstream — persons with a name but no id are
  excluded from duplicate detection entirely)
- Belt-and-suspenders on warning-path display: p['id'] → p.get('id')
- Extract survivor_id with .get(); skip group if survivor has no id
- Guard merge_ids: skip API call when list is empty after id filtering
- Walrus operator in merge_ids comprehension: p.get("id") called once
  per item instead of twice

executor.py:
- Add cross-reference comment at success-path reset so the for/else
  rollback pairing is explicit for future maintainers
2026-06-17 02:09:16 +00:00
flan eab3d9fe64 fix: 3 correctness bugs from codebase audit round 2
- executor.py: clear min_quality_score_for_slot alongside effective_count
  restore in for/else block; leaving the stale floor from the deleted
  file's score blocked the next candidate from filling the restored slot
- cli.py: guard merge_ids with p.get('id') is not None, consistent with
  the _smaller_duplicate_ids fix; bare p['id'] raised KeyError on any
  person dict missing the id field in the auto-merge path
- immich_api.py: replace bare data['major'/'minor'/'patch'] subscripts
  with .get() in get_immich_version; KeyError was silently swallowed by
  except Exception, causing version-gated flags to disable without warning
2026-06-17 01:54:13 +00:00
flan 5509be150e Merge pull request #38 from sudolulo/fix/codebase-audit-r1
fix: codebase audit r1 — correctness fixes, version banner, GPU dep
2026-06-16 21:51:09 -04:00
flan 0bd2eaf9fb fix: add missing nvidia CUDA pip packages for onnxruntime-gpu 1.26.0
ORT 1.26.0 changed provider loading to gate on the presence of required
nvidia pip packages before attempting to load libonnxruntime_providers_cuda.so.
Without nvidia-cuda-runtime-cu12, nvidia-cufft-cu12, and nvidia-curand-cu12
installed as Python packages, ORT silently skips the CUDA EP plugin entirely
(confirmed via /proc/maps: the .so was never dlopen'd despite existing on disk
and all system CUDA libs being present in ldconfig).

nvidia-nvjitlink-cu12 pulled in as a transitive dependency.
2026-06-17 01:47:49 +00:00
flan b80d26b36b feat: display version in startup banner 2026-06-17 01:18:45 +00:00
flan 068a8e675f fix: 4 correctness bugs from full-codebase audit
- executor: restore effective_count when replacement upload fails all retries
  (delete succeeded but slot was never filled, leaving cap undercount)
- diversity: skip zero-norm embeddings before dedup/FPS selection
  (InsightFace zeros pass dedup with similarity 0 and score distance 1.0,
  getting selected first as maximally diverse)
- cli: exclude None from skip_ids in _smaller_duplicate_ids
  (p.get('id') without None guard lets None into the set, silently
  dropping every other id-less person from the processed list)
- embeddings: select face nearest crop centre instead of largest by area
  (25% margin can pull a bigger neighbouring face into the crop;
  largest-face selection then embeds the wrong person)
2026-06-17 01:12:28 +00:00
flan c36e7bf28e Merge pull request #37 from sudolulo/dev
fix: exclude main from lockfile update workflow trigger
2026-06-16 21:03:35 -04:00
flan 0ffe08bc6f Merge pull request #36 from sudolulo/fix/lockfile-workflow-main-exclusion
fix: exclude main from lockfile update workflow trigger
2026-06-16 21:01:25 -04:00
flan 3423d41535 fix: exclude main from lockfile update trigger
main is protected and only receives merges from dev; pushing directly
to it from CI is blocked by branch protection rules.
2026-06-17 00:56:52 +00:00
flan 6e29407231 Merge pull request #35 from sudolulo/dev
release: v0.6.4
2026-06-16 20:21:19 -04:00
flan 5dcfde7c36 release: v0.6.4 2026-06-17 00:17:32 +00:00
flan 0914608bc8 fix: address 2 missed p[\"id\"] bare subscripts in cli.py (round 12)
Round 11's replace_all missed two occurrences:
- _smaller_duplicate_ids inner comprehension (line 84): p["id"] →
  p.get("id") so a named person with a missing "id" field does not
  crash skip_ids computation before any return path is reached
- all-merges-failed fallback return (line 157): same fix; the outer
  indentation prevented replace_all from matching this occurrence

The intentional p["id"] in merge_ids (line 119) is kept: that ID is
passed directly to merge_people() where None would be a caller bug,
not a silent data corruption.
2026-06-17 00:10:01 +00:00
flan 2182c87c40 fix: address 2 code review findings (round 11)
- upload_tracker: revert data[flat_key] = [] from round 8; clearing the
  entire shared legacy flat list on a corrupt value wipes all persons'
  IDs, not just the one being reset; since a corrupt non-list value is
  already unreadable by load_uploaded_ids, leaving it in place is safer
  than a mass-wipe; update warning message to note the field is unaffected
  but unreadable so the corruption is still observable
- cli: use p.get("id") instead of p["id"] in both people-list fallback
  returns (_handle_duplicate_people lines 144 and 157) for consistency
  with the success path at line 149; bare subscript crashes on malformed
  unnamed persons that bypass _smaller_duplicate_ids
2026-06-17 00:02:49 +00:00
flan 0236ed2d6b fix: address 3 code review findings (round 10)
- immich_api: use 'or []' instead of .get("people", []) in get_people
  so {"people": null} responses (some Immich versions with zero people
  enrolled) return [] rather than None; .get() default only fires when
  the key is absent, not when its value is null
- embeddings: log OSError from os.dup2 restore at DEBUG rather than
  silently swallowing it; if a C extension (CUDA/onnxruntime) invalidates
  the saved fd, the restore fails silently and stdout stays wired to
  /dev/null — logging makes the event observable without changing the
  swallow-and-continue semantics
- cache: remove MemoryError re-raise from EmbeddingCache.get(); a cache
  read OOM aborted the entire diversity-selection batch for the person
  rather than falling back to a fresh embedding computation, which is
  the more appropriate OOM gate; broadening back to except Exception
  restores the pre-round-5 fallback behavior
2026-06-16 23:51:21 +00:00
flan 34f7985357 docs: document BaseException limitation in _suppress_output finally block
A KeyboardInterrupt raised inside the saved_out cleanup block would
propagate past the saved_err and devnull_fd blocks, leaking those fds.
In CPython this race is not realistically triggerable — KI is delivered
between bytecodes and os.dup2 is a single atomic C syscall — so we
accept the theoretical risk rather than silencing BaseException in a
finally block.
2026-06-16 23:41:38 +00:00
flan 25880ded91 fix: address 2 code review findings (round 9)
- executor: revert person_has_fscores=True back into try/except else
  branch; moving it outside in round 8 was a regression — when the
  tracker write fails on the first-ever upload (no prior frigate_scores
  in tracker), setting the flag True prematurely switches at-cap
  replacement into fscore mode, get_most_redundant_mapped_file returns
  None (no entries), and all replacements are silently skipped;
  the flag must only be set when the score is actually written
- diversity: remove dead face_crop-None guard; any face that passes
  assess_quality (≥90 px MIN_FACE_WIDTH) produces a crop ≥135 px
  (face + 25% margin), which is always above the 30 px crop minimum,
  making the guard unreachable; _crop_face_from_thumbnail also calls
  _get_face_bbox internally, so face_bbox is not None guarantees the
  inner bbox check also passes
2026-06-16 23:40:05 +00:00
flan 461ceb7af4 fix: address 5 code review findings (round 8)
- diversity: fix hard_count regression from round 7 — revert to
  'is not None and < 0.85' so only images that actually receive a
  FPS boost (confirmed low confidence) are counted as hard examples;
  None-confidence images use conf_array=1.0 (no boost) and should
  not appear in the hard-example log count
- diversity: fix _scale_bbox_to_thumbnail to use explicit zero-guard
  for imageWidth/imageHeight (meta_w or 0; scale = img_w/meta_w if
  meta_w else 1.0) — mirrors image_processing.py pattern; prevents
  `or img_w` from silently treating imageWidth=0 as missing and
  returning scale=1.0 without surfacing the zero-metadata case
- embeddings: wrap all three os.close calls in _suppress_output
  finally block with try/except OSError: pass so a failed close
  in one branch cannot abort the outer finally and leak devnull_fd
  or the saved_err/saved_out fds
- upload_tracker: clear corrupt flat-list key (data[flat_key] = [])
  after the isinstance warning instead of leaving the corrupt value
  in place — prevents stale IDs persisting across reset_person calls
  and future load_uploaded_ids() from seeing a non-list value
- executor: move 'if pre_fscore is not None: person_has_fscores = True'
  out of the try/except else branch so it fires even when mark_uploaded
  raises; Frigate scores exist once measured regardless of tracker
  write success, and replacement strategy should reflect that
2026-06-16 23:24:11 +00:00
flan 7282c76b68 fix: address 5 code review findings (round 7)
- diversity: revert conf_array default from 0.5 back to 1.0 (np.ones);
  the 0.5 default caused None-confidence images to receive a 1.7× FPS
  boost and beat high-confidence detections — counter-productive for
  Frigate training data quality
- diversity: fix hard_count to include None-confidence images (count
  images where score is None or < 0.85, not only confirmed < 0.85);
  the previous check systematically undercounted boosted images when the
  Immich faces API omits the score field
- executor: fix garbled comment fragment "Skipped on / skipped when"
  left by a partial edit in round 4; merge into a single coherent sentence
- executor: expand actually_uploaded trade-off comment to document all
  three consequences of a tracker write failure (Frigate duplicate,
  quality-replacement exclusion, cap-slot consumption), not only the
  duplicate risk mentioned previously
- upload_tracker: add person_ids guard to reset_person isinstance check
  so the non-list warning only fires when cleanup would actually have run,
  not on no-op calls where person_ids is empty
2026-06-16 23:09:14 +00:00
flan 692d77ee9f fix: address 4 code review findings (round 6)
- executor: snapshot has_frigate_model = effective_count > 0 before the
  upload loop; use it in the recognize_face gate instead of the live
  effective_count, which is incremented mid-loop and would otherwise
  trigger recognize_face calls against an empty Frigate model on first run
- jobs: restore if already_uploaded > 0 guard before limit = capacity so
  first-run auto-strategy jobs keep limit="auto" and the FPS adaptive
  early-stop can fire instead of always filling MAX_AUTO_IMAGES slots
- cli: retry get_people() once after a post-merge empty response before
  falling back to the pre-merge list; improve warning to name expired API
  key as a possible cause alongside transient network errors
- diversity: hoist hard_weight = np.where(...) above the FPS while loop
  since conf_array is constant; eliminates one O(n) numpy pass per
  selected image
2026-06-16 22:22:51 +00:00
flan 2cb126a589 fix: address 5 code review findings (round 5)
- embeddings: move os.close into try/finally so saved_out/saved_err are
  always closed even when os.dup2 restore raises, preventing fd leak
- cache: replace narrow except tuple with except MemoryError: raise /
  except Exception: return None so struct.error and other np.load failures
  return None without masking OOM
- executor: fix first-run advisory message to check effective_count == 0
  (post-stale-cleanup) instead of pre_run_count; remove now-unused
  pre_run_count variable entirely
- jobs: remove dead "skip" entry from strategy_map (unreachable since the
  early-return at the top of _resolve_strategy fires first)
- upload_tracker: log a warning when reset_person encounters a non-list
  flat_key value instead of silently skipping the cleanup
2026-06-16 22:03:59 +00:00
flan 96099ed6e2 fix: address 8 code review findings (round 4)
- diversity: remove erroneous break outside if-faces in _scale_bbox_to_thumbnail
  (broke people-loop early for first unannotated person, defeating scale fix)
- diversity: increment quality_filtered for face-too-small crop skips so the
  summary log counts them alongside assess_quality failures
- diversity: fix hard-example log count to use original confidence_scores[i]
  instead of synthetic conf_array default (0.5), eliminating false 100%
  hard-example reports for persons with no Immich confidence data
- cache: add EOFError to except tuple in EmbeddingCache.get() so truncated
  .npy files return None instead of crashing the embedding pipeline
- embeddings: wrap each os.dup2 restore in its own try/except OSError in
  _suppress_output finally block so stderr is always restored even if the
  stdout restore raises
- executor: gate recognize_face on effective_count > 0 (post-stale-cleanup)
  instead of pre_run_count > 0 so recognize_face is not called against an
  untrained Frigate model after the user manually deletes all training files
- executor: document actually_uploaded trade-off in comment (appending
  unconditionally on tracker failure risks a Frigate duplicate but prevents
  permanent filename unmapping which breaks quality-replacement scoring)
- jobs: check strategy == "skip" before the has_embedding and custom_limit
  early-returns in _resolve_strategy so STRATEGY=skip is always honoured
2026-06-16 21:43:24 +00:00
flan 4af9da2550 fix: address 10 code review findings (round 3)
- diversity: scale face bbox to thumbnail space before quality check so
  check_face_size uses actual thumbnail pixels, not original-image coords
- diversity: skip asset when face bbox exists but crop guard rejects it,
  preventing InsightFace from picking the wrong person in a group photo
- diversity: add _scale_bbox_to_thumbnail helper (extracted from crop logic)
- diversity: use set for medoid membership test in _kmedoids (O(n) not O(n*k))
- diversity: remove dead np.unique in _select_time_spread (linspace produces
  strictly increasing indices; unique is a no-op and implies wrong semantics)
- embeddings: move os.open/os.dup calls inside try in _suppress_output so
  EMFILE during setup does not leak already-allocated fds
- immich_api: count and log assets with missing/unparseable fileCreatedAt in
  filter_recent_assets instead of silently discarding them
- executor: capture pre_run_count before stale-mapping cleanup so the
  "first run" coaching message doesn't fire after manual file deletion
- cli: use p['id'] (KeyError-safe) instead of p.get('id') in fallback path
  to match all other access sites on the same people list
- cache: narrow except to (OSError, ValueError) in EmbeddingCache.get so
  MemoryError propagates instead of converting OOM to a silent cache miss
2026-06-16 21:17:31 +00:00
flan 8bdce9253a fix: address 10 codebase audit findings — API guards, reconcile, merge fallback, tracker guards
- immich_api: guard resp.json() with isinstance(dict) check in get_people and
  fetch_all_assets so AttributeError doesn't escape on proxy/CDN non-dict responses
- executor: move actually_uploaded.append outside try/else so Frigate filename→asset_id
  mapping is created via reconcile even when the tracker write fails
- cli: fall back to pre-merge people list when re-fetch after merge returns empty
  (transient error) instead of silently dropping all people
- cli: treat ENABLE_FRIGATE_SCORES=false / BLUR_THRESHOLD=0 as not-set in
  the unsupported-vars warning (falsy string check replaces raw truthiness)
- upload_tracker: guard set(data[flat_key]) with isinstance(list) check in
  reset_person so a corrupted non-iterable legacy field doesn't crash mid-reset
- upload_tracker: guard dims[0]/dims[1] in find_by_crop_dimension with a
  length check so a truncated crop_dims entry doesn't raise IndexError
- cache: wrap os.remove() in clear() with try/except OSError to handle
  TOCTOU race with concurrent put() calls
- diversity: default conf_array to 0.5 (was 1.0) for faces with missing
  confidence so they receive a moderate diversity boost instead of being
  treated as high-confidence
- diversity: sort assets in the fast path (len <= limit) so return order is
  consistent with the sorted-by-fileCreatedAt path
2026-06-16 20:51:14 +00:00
flan 34fccf8839 fix: address 10 full-codebase audit findings + lint
Correctness:
- jobs: cap auto-diversity limit for brand-new people (was never capped,
  could exceed MAX_AUTO_IMAGES on first run)
- image_processing: separate None/0 guard for imageWidth/imageHeight so
  missing field is explicit rather than silently aliased to img_w
- upload_tracker (_mark, update_frigate_count): copy-before-mutate so
  exceptions between cache access and _save don't corrupt in-process state
- jobs: reject LIMIT=0 on no-embedding path (was silently empty run)
- jobs: add STRATEGY=skip to strategy_map so env var is honoured
- embeddings: convert to RGB before cvtColor so RGBA/grayscale thumbnails
  don't raise cv2.error and silently drop from diversity selection
- config: use falsy guard for OUTPUT_DIR so blank env var falls through
  to config file value
- reconcile: _ts() returns float("inf") on parse failure so unrecognised
  filenames sort last instead of collapsing to 0.0 and corrupting FIFO mapping
- diversity: remove dead selected_set (never read; -np.inf sentinel already
  prevents re-selection)

Lint (ruff):
- executor: sort upload_tracker import block (I001)
- executor: replace lambda is_better_than with operator.lt/gt (E731 x2)
- executor, upload_tracker: wrap long logger.warning calls (E501 x4)
2026-06-16 18:40:13 +00:00
flan 7a268d1ea2 chore: sync dev with main (v0.6.3) 2026-06-16 18:18:42 +00:00
flan 44cbedaf91 Merge branch 'main' of github.com:sudolulo/winnow 2026-06-16 18:16:27 +00:00
flan 3c2ce80282 Merge branch 'main' of github.com:sudolulo/winnow into dev 2026-06-16 18:16:14 +00:00
flan 3c0ef47fdc release: v0.6.3 2026-06-16 18:13:42 +00:00
flan cf7660595d chore: update lockfile 2026-06-16 18:13:42 +00:00
flan 14f759e960 fix: address 3 quality review findings — record_frigate_files_batch cache mutation, tracker_ok flag, LIMIT guard
- record_frigate_files_batch: copy-before-mutate so a write failure
  doesn't leave cache ahead of disk (same fix as remove_frigate_files_batch)
- executor: replace tracker_ok boolean with try/else
- jobs: collapse duplicate custom_limit is not None checks into one guard

Bump version to 0.6.3.
2026-06-16 18:13:36 +00:00
flan e8cb390fe4 fix: address 2 quality review findings — begin_batch dirty guard, LIMIT<=0 warning 2026-06-16 18:04:38 +00:00
flan 54b52b0a73 fix: address 3 quality review findings — batch reject tracker, skip flush when clean, hoist frigate url check 2026-06-16 17:55:34 +00:00
flan b622e58f1b fix: address 3 quality review findings — flush_batch finally guard, _laplacian_var helper, has_frigate_scores no-copy 2026-06-16 17:09:00 +00:00
flan f3622b8d41 fix: address 4 quality review findings — flush_batch order, batch finally guard, cache copy, LIMIT=0 fallthrough 2026-06-16 16:59:43 +00:00
flan 817fa17e41 fix: address 3 quality review findings — tracker_ok gate, LIMIT=0 warning, cache write log level 2026-06-16 16:41:32 +00:00
flan 8846a4f1df fix: address 3 post-fix audit findings — begin_batch flush guard, misleading debug log, shared asset_id score deletion 2026-06-16 16:19:54 +00:00
flan a6bae5da05 fix: address 10 audit findings — import bug, fscore stale flag, cache mutation, batch safety, falsy guards 2026-06-16 16:16:28 +00:00
flan 4cdd4657d6 fix: v0.6.2 — structural tracker refactor, batch writes, multi-instance prep
- Drop flat list as primary storage; derive uploaded/rejected IDs from by_person
  (single source of truth). Legacy flat lists in existing files still read for
  backward compat. Removes dual-representation sync hazard.
- Add begin_batch/flush_batch: per-person upload loop now does 1 os.replace
  instead of N (one per mark_uploaded call). Benefit on slow storage.
- reset_all_people(): RESET_PERSON=* is now O(1) disk writes instead of O(P^2).
- blur_score_from_image inlines cv2.Laplacian directly, removing assess_quality
  call overhead and decoupling from the full quality pipeline.
2026-06-16 15:44:41 +00:00
flan dc2efb5ac4 fix: v0.6.1 — tracker integrity, quality replacement correctness, code review fixes
- Catch OSError alongside PIL.UnidentifiedImageError for corrupt thumbnails
- Fix quality replacement mode flip mid-loop (person_has_fscores no longer re-evaluated)
- reset_person rebuilds flat list from remaining entries instead of subtracting
- _save cache updated only after os.replace succeeds (prevents cache/disk split-brain)
- Stale Frigate file cleanup uses remove_frigate_files_batch (N writes → 1)
- _migrate_entry deep-copies nested dicts so .pop() cannot mutate the cache
- find_by_crop_dimension and _pick_mapped_file consistent on duplicate asset→file mapping
- Atomic JSON write (tmp + os.replace) guards against truncated files on crash
- get_person_summary uses _migrate_entry instead of three isinstance guards
- Quality floor check allows None-scored candidates through (don't block freed slots)
- Fix comment-only if body (IndentationError on import) in full-res download path
- Merge duplicate if-stale guard into one block
- _flat_key uses constant equality instead of substring match
- remove_frigate_file returns early when person absent (no ghost entries)
- skip_ids extracted to _smaller_duplicate_ids() helper (was duplicated 3×)
- blur_score_from_image returns None on error instead of 0.0
2026-06-16 15:21:12 +00:00
flan 2de0c02c4e Merge pull request #34 from sudolulo/dev
release: v0.6.0 — revert SQLite tracker to JSON backend
2026-06-15 11:53:23 -04:00
flan 9e84e276da Merge remote-tracking branch 'origin/main' into dev 2026-06-15 15:52:29 +00:00
flan 794dbe2a1d revert: replace SQLite tracker with JSON backend (v0.6.0) (#32)
* revert: replace SQLite tracker with JSON backend (v0.6.0)

The SQLite migration (v0.5.0) spawned 21 bug-fix releases in two days:
data-loss risk in the migration layer, schema PK conflicts on per-person
tracking, tracker isolation races under concurrent runs, and a disk-full
error that triggered duplicate Frigate uploads. The complexity cost
outweighs the benefit.

Restored the pre-SQL JSON tracker (frigate_uploaded_ids.json /
frigate_rejected_ids.json in DATA_DIR). Public API is identical — all
callers in executor.py, jobs.py, cli.py, and reconcile.py work unchanged.
Existing JSON files are read automatically; frigate_tracker.db can be
deleted once verified.

* fix: narrow corrupt-thumbnail exception to UnidentifiedImageError; restore IMMICH_URL empty-string fallback

* docs: rewrite v0.6.0 changelog, strip v0.5.x entries, fix README SQLite references

* chore: remove dead get_frigate_filename_for_asset (orphaned since FRIGATE_SCORE_THRESHOLD removal in v0.4.0)

* fix: sort imports in executor.py (ruff I001)
2026-06-15 11:44:41 -04:00
flan 86c76ba6a7 Update README.md (#33) 2026-06-15 11:44:12 -04:00
flan 0602c4ac04 fix: v0.5.21 — diversity cap, fetch rejection, tracker isolation, merge dedup, env parsing
- diversity: cap k-medoids seed count at target so _cluster_aware_selection
  never returns more images than requested (violated MAX_AUTO_IMAGES when
  remaining capacity was 1-4 slots); add early return for limit=0 to
  prevent k-medoids from running with a zero budget; slice return to target
  as a final guard
- executor: mark_rejected() when fetch_full_image returns None so assets
  that can't be fetched (both original and preview) aren't retried every run
- executor: wrap mark_uploaded() in its own try/except so a SQLite disk-full
  error after a successful HTTP 200 doesn't retry the Frigate POST (duplicate
  upload) — the upload succeeded; only the tracker write failed
- cli: apply skip_ids deduplication to the re-fetched people list after a
  partial merge (some groups succeed, some fail) so unmerged duplicates
  don't produce two jobs for the same Frigate folder
- jobs: strip whitespace from SKIP_PEOPLE/ONLY_PEOPLE elements on split
  so "Alice, Bob" (space after comma) correctly matches "Bob"
2026-06-15 03:56:57 +00:00
flan 95fb39ed50 fix: v0.5.20 — migration safety, URL normalization, image rejection, atomic writes
- upload_tracker: replace executescript() in _migrate_schema_v2 with
  individual execute() calls inside a transaction so a crash between DROP
  and RENAME rolls back instead of permanently destroying tracked_assets
- frigate_api: _get_frigate_url now strips leading/trailing whitespace
  before rstrip('/') so whitespace-only FRIGATE_URL is treated as unset
- executor: upload_to_frigate now uses _get_frigate_url() eliminating
  double-slash upload paths when FRIGATE_URL has a trailing slash
- executor: corrupt thumbnail (resp.ok=True, Image.open fails) now calls
  mark_rejected() so permanently broken assets are not retried forever
- upload_tracker: reset_person now uses _get_frigate_url() instead of
  inline os.environ.get('FRIGATE_URL', '').strip()
- image_processing: _save_jpeg writes to a .tmp file and calls
  os.replace() so a disk-full error never leaves a truncated JPEG
- cli: _handle_duplicate_people falls back to local deduplication when
  all Immich merges fail, preventing two jobs from overwriting the same
  Frigate folder
- config: _getenv_optional_float now delegates to _getenv_num() like
  _getenv_optional_int, eliminating the inconsistent duplicate
- reconcile: _ts() uses rsplit('.', 1)[0] instead of .replace('.webp','')
  so FIFO mapping works with any Frigate training-file extension
2026-06-15 03:26:51 +00:00
flan 728b84dc8c fix: address 10 full-codebase audit findings (v0.5.19)
Correctness:
- fetch_face_data: only fall back to faces[0] when person_id is absent;
  previously a missing person match injected a different person's bbox
- upload_tracker: change PK from (asset_id, status) to
  (asset_id, person_name, status); old PK allowed INSERT OR REPLACE to
  silently overwrite person_name when the same photo appeared in two
  people's jobs, breaking quality-replacement JOINs; auto-migrates DBs
- filter_recent_assets: treat years=0 as "no age filter" instead of
  falling through to Config.YEARS_FILTER via falsy `or`
- _is_module_available: return find_spec(...) is not None; find_spec
  returns None (not raises) for absent top-level modules, so the
  previous code always returned True
- execute_jobs error handler: use asset.get("id", "<unknown>") to avoid
  a secondary KeyError propagating out of execute_jobs on malformed dicts
- upload_to_frigate: also mark_rejected on HTTP 422, not only HTTP 400
  with "face" in body; other permanent errors left assets untracked and
  retried forever
- reconcile_frigate_mappings: sort key lambda f: (_ts(f), f) makes order
  deterministic when timestamps are equal or 0.0; set iteration order is
  hash-randomised, stable sort preserves it

Reuse / cleanup:
- config.py: add _getenv_optional_int delegating to _getenv_num(name, None, int)
- jobs.py: _resolve_strategy uses _getenv_optional_int("LIMIT") instead
  of inline os.environ.get + int() + warning duplicate of _getenv_num
- frigate_api.py: add _get_frigate_url() helper; eliminates 4× copy of
  os.environ.get("FRIGATE_URL", "").rstrip("/")
- quality.py: extract blur_score_from_image(img, max_dim=1440) helper;
  executor.py time-spread blur fallback now uses it instead of inlining
  the resize+RGB+assess_quality sequence, keeping scale logic in one place
2026-06-15 02:56:51 +00:00
flan 3c4479a110 fix: replace inline FORCE_CPU check in benchmark.py with _getenv_bool
scripts/benchmark.py retained the old os.getenv inline pattern after
_getenv_bool was introduced in v0.5.16. Now uses a deferred local
import of _getenv_bool, consistent with the script's pattern of keeping
all winnow imports inside function bodies rather than at the top level.
2026-06-15 02:29:30 +00:00
flan 06c2c4a584 fix: strip empty env vars in _getenv_num/_getenv_bool; unify FORCE_CPU
- _getenv_num: add raw.strip() + empty-string guard so numeric vars set
  to "" (common Compose pattern for "use default") return the default
  silently instead of warning "not a valid int/float"
- _getenv_bool: same guard so True-defaulted flags set to "" return
  the configured default instead of silently returning False
- embeddings.py: replace inline FORCE_CPU bool parse with _getenv_bool
2026-06-15 02:20:34 +00:00
flan de6804226d refactor: unify env var helpers and remove magic slice in cache
- Add _getenv_num as shared core for _getenv_int/_getenv_float
- Add _getenv_optional_float for FRIGATE_SCORE_CEILING (replaces 9-line inline block)
- Add _getenv_bool; replace 11 inline .lower()-in-("true","1","yes") sites
  across config.py, jobs.py, and cli.py with single call site
- _resolve_strategy no-embedding branch: inline try/except → _getenv_int("LIMIT", 30)
- cache.py: final[:-4] → final.removesuffix(".npy") — assumption is now explicit
2026-06-15 02:10:00 +00:00
flan fe5e1574ac release: v0.5.15 — fix cache regression and structural cleanup
- cache.py: fix np.save extension bug from v0.5.13 — tmp path used
  final+".tmp" (abc.npy.tmp) but np.save auto-appends .npy to paths not
  ending in .npy, writing to abc.npy.tmp.npy instead; os.replace then
  raised FileNotFoundError silently, making every cache write a no-op
  and leaking *.npy.tmp.npy files. Fixed by inserting .tmp before .npy:
  tmp = final[:-4] + ".tmp.npy"

- config.py: remove str(default) round-trip in _getenv_int/_getenv_float
  — use raw = os.getenv(name); return default if raw is None else int(raw)
  so a future float default can't cause a spurious "not a valid integer"
  warning and return the wrong type

- executor.py: consolidate 4 progress.remove_task calls into one
  try/finally around the per-job body; continue inside try/finally
  executes the finally before the next iteration, making the invariant
  structurally enforced rather than relying on discipline across 4 sites
2026-06-15 01:45:43 +00:00
flan 5bcc5975bc release: v0.5.14 — graceful fallback for invalid numeric env vars
YEARS_FILTER, MIN_FACE_WIDTH, MIN_FACE_COUNT, MAX_AUTO_IMAGES,
BLUR_THRESHOLD, MIN_CONFIDENCE, and FACE_MARGIN used bare int()/float()
calls with no error handler. A typo (trailing space, non-numeric value)
raised ValueError inside __getattr__, producing a cryptic traceback on
the first config access rather than at the validate() step. Values are
now parsed by _getenv_int/_getenv_float helpers that warn and fall back
to the documented default, matching the existing FRIGATE_SCORE_CEILING
pattern.
2026-06-15 01:35:27 +00:00
flan 51f7ed3961 release: v0.5.13 — robustness fixes from full-project audit
- executor.py: wrap shutil.rmtree/os.makedirs in try/except OSError so a
  permission failure logs and skips the job rather than aborting the run
- cache.py: write embeddings to a .tmp file and atomically rename into place
  via os.replace so a process kill can't leave a corrupted .npy cache slot
- immich_api.py: guard fileCreatedAt with isinstance(str) check before calling
  .replace() so a non-string timestamp doesn't raise AttributeError and kill
  the entire filter_recent_assets pass
- upload_tracker.py: raise SQLite busy timeout from 5 s to 30 s to handle
  concurrent cron+manual run overlap without dropping upload-tracking records
2026-06-15 01:19:38 +00:00
flan 850ae2f9fc fix: remove progress task on skipped jobs (v0.5.12)
progress.add_task() fires unconditionally at the top of the job loop;
both continue paths (ValueError from _safe_person_dir and the symlink
TOCTOU guard) skipped remove_task(), leaving orphaned 0% rows in the
terminal for the rest of the run.
2026-06-15 01:05:49 +00:00
flan 796aded2da fix: log+skip on symlink TOCTOU in execute_jobs (v0.5.11)
The v0.5.10 compound guard 'isdir and not islink' silently skipped the
rmtree when person_dir was a symlink-to-directory, then let makedirs
follow the symlink — allowing crop writes outside output_dir with no
diagnostic. Replace with an explicit islink pre-check that logs an error
and continues, matching the ValueError path from _safe_person_dir.
2026-06-15 01:02:28 +00:00
flan 480bf80534 fix: reconcile < target severity and rmtree symlink guard (v0.5.10)
- reconcile.py: re-escalate the < target branch from INFO to WARNING and
  add 'permanently unmapped' label. Both post-loop branches produce identical
  permanent mapping loss; v0.5.9 incorrectly treated the timeout case as
  recoverable.

- executor.py: guard shutil.rmtree with 'not os.path.islink(person_dir)'
  so a race-replaced symlink-to-directory is skipped rather than raising
  an unhandled OSError that aborts all remaining jobs. Correct comment:
  rmtree raises OSError, not NotADirectoryError.
2026-06-15 00:58:22 +00:00
flan 2d39291fe7 fix: correct reconcile log severity, docstring gaps, and _entry allocation (v0.5.9)
- reconcile.py: swap log levels — external-upload path (permanent mapping
  loss) escalated to WARNING; timeout path (transient, retries next cycle)
  downgraded to INFO. Also extend the warning message to note the files are
  permanently unmapped.

- immich_api.py: extend fetch_all_assets docstring to document that
  all-garbage page termination (in addition to network errors) makes
  total_raw a lower bound.

- executor.py: add comment above shutil.rmtree noting that POSIX rmtree
  raises NotADirectoryError on a top-level symlink, documenting why the
  removed islink guard is safe to omit.

- upload_tracker.py: replace setdefault with explicit guard in _entry() —
  setdefault evaluates its default-dict argument before checking key
  presence, allocating and discarding a dict on every already-present call.
2026-06-15 00:51:21 +00:00
flan 1556d90bcc fix: correct path traversal docstring, total_raw inflation, and config re-stat (v0.5.8)
- executor.py: fix _safe_person_dir docstring — realpath+startswith is the
  load-bearing traversal guard; islink is a supplementary early-exit for the
  symlink sub-case only. The previous comment "checking after realpath would be
  too late" implied islink was the primary guard, which is backwards.

- immich_api.py: move total_raw accumulation to after the dead-end-page break
  so all-garbage pages don't inflate the count and produce misleading
  "N total, 0 recent" output. Mixed pages (some valid, some non-dict) still
  count page_count so transient schema issues don't shrink MIN_FACE_COUNT below
  threshold. Add warning when a RequestException interrupts pagination mid-way
  so operators know total_raw is a lower bound.

- config.py: eliminate residual TOCTOU — change `if config_file.exists():` to
  `if _data_cfg_exists or config_file.exists():` so _data_cfg is never
  stat'd twice (the v0.5.7 fix cached the first check but not the second).
2026-06-15 00:41:00 +00:00
flan 00e378b375 Merge remote-tracking branch 'origin/dev' 2026-06-15 00:27:25 +00:00
flan 9685c310af fix: symlink guard placement, fetch_all_assets raw count, config TOCTOU (#31)
executor.py:
- Move islink check into _safe_person_dir on the raw path, before realpath
  resolves it; the previous check at the rmtree site was unreachable dead code
  because realpath already followed any symlink

immich_api.py / jobs.py:
- fetch_all_assets now returns (assets, total_raw) where total_raw is the
  item count seen before non-dict filtering; callers use it for MIN_FACE_COUNT
  guard and display so transient non-dict API items can't incorrectly skip people
- Add WARNING when pagination stops because a page had items but all were non-dict

config.py:
- Cache _data_cfg.exists() in _data_cfg_exists so the dual-config warning
  and config_file selection always read from the same stat() result; previously
  two calls created a TOCTOU window where log and code could disagree

Bump version to 0.5.7
2026-06-14 20:27:16 -04:00
flan e5b9293861 Merge remote-tracking branch 'origin/dev' 2026-06-15 00:16:33 +00:00
flan 363190dbe5 fix: pagination runaway, double iteration, and reconciliation efficiency (#30)
immich_api.py:
- Move empty-page break after non-dict filtering — a page of all-null
  items no longer loops to MAX_PAGES without terminating
- Single-pass partition replaces two inverse isinstance scans per page
- Upgrade non-dict item log from DEBUG to WARNING (silent asset loss)

reconcile.py:
- Check Frigate before the first sleep so fast responses return
  immediately rather than always paying a 1 s delay
- Compute set difference once per poll iteration instead of twice

Bump version to 0.5.6
2026-06-14 20:16:27 -04:00
flan 78e01d9621 Merge remote-tracking branch 'origin/dev' 2026-06-15 00:04:29 +00:00
flan 3a052db1ce chore: quality cleanup — extract magic numbers, improve docs and naming (#29)
- diversity.py: extract 3000/20/32 to _POOL_CAP/_POOL_SCALE/_EMBEDDING_BATCH_SIZE
- reconcile.py: extract (1,2,4,8) poll delays to _RECONCILE_POLL_DELAYS with comment
- immich_api.py: document dual response shape; debug-log skipped non-dict items
- frigate_api.py: rename `encoded` → `encoded_name` for clarity
- upload_tracker.py: atomic-write note on record_frigate_files_batch docstring;
  get_person_summary() uses setdefault to eliminate four repeated default dicts;
  _VALID_SCORE_COLS comment explains SQL-injection guard intent

Bump version to 0.5.5
2026-06-14 20:04:24 -04:00
flan df98169398 Merge pull request #28 from sudolulo/dev
release: 0.5.4
2026-06-14 19:53:48 -04:00
flan 84ebd91929 fix: quality replacement slot floor uses deleted file's score not failed candidate's (#27)
* fix: quality replacement slot floor uses deleted file's score not failed candidate's

When a blur-score replacement deletes a low-quality Frigate file but the
subsequent upload fails, min_quality_score_for_slot was set to candidate_score
(the good file that failed to upload). This filtered out any subsequent
candidate that didn't beat the failed upload, even if it was better than
the file we just deleted — leaving the freed slot unfilled unnecessarily.

The comment on the guard already documented the correct intent: 'require
the next candidate to beat the deleted file's score'. Fix: use target_score
(the deleted file's blur score) as the floor instead of candidate_score.

* chore: bump version to 0.5.4
2026-06-14 19:53:33 -04:00
flan c77069f2e2 Merge pull request #26 from sudolulo/dev
release: 0.5.3
2026-06-14 19:39:51 -04:00
flan 2e08504682 fix: audit hardening — input validation, error handling, and robustness (#25)
* fix: audit hardening — input validation, error handling, and robustness

- immich_api: guard person["id"] with .get() + early return on missing field
- immich_api: include page number in pagination exception log
- immich_api: validate faces response is a list before indexing
- executor: wrap Image.open() in try/except for non-image HTTP responses
- executor: strip leading 'v' from Frigate version before parsing (v0.16.0 was misread)
- config: wrap FRIGATE_SCORE_CEILING float() parse in try/except with warning
- config: warn when both DATA_DIR and legacy CWD config files exist simultaneously
- scheduler: wrap PID file write in try/except so /tmp failures don't crash startup
- scheduler: clamp sleep to 60s max to bound recovery time after NTP clock jumps
- frigate_api: log unexpected non-list type in get_frigate_person_files at DEBUG

* fix: LIMIT env var crash and symlink guard on person output dir

- jobs: wrap int(LIMIT) parse in try/except — bad value (e.g. "30.5", "all")
  now logs a warning and falls back to the default instead of crashing
- executor: check for symlink before shutil.rmtree on person_dir — prevents
  following a symlink out of OUTPUT_DIR on a shared volume

* chore: bump version to 0.5.3
2026-06-14 19:39:35 -04:00
flan 8bf23edc85 Merge pull request #24 from sudolulo/dev
release: 0.5.2
2026-06-14 19:17:24 -04:00
flan 166729a17d Merge remote-tracking branch 'origin/main' into dev
# Conflicts:
#	CHANGELOG.md
#	pyproject.toml
#	uv.lock
2026-06-14 23:17:16 +00:00
flanandgithub-actions[bot] 8acf8b52b8 fix: Immich v2.7.5 compat, supply-chain hardening, and quality fixes (0.5.2) (#23)
* fix: Immich v2.7.5 compat, supply-chain hardening, and quality fixes (0.5.2)

- Remove assetCount pre-filter broken by Immich v2.7.5 API change; check
  MIN_FACE_COUNT after fetch_all_assets instead
- Replace curl|sh uv installer with COPY --from Docker stage (supply chain)
- Fix HEALTHCHECK to use kill -0 on PID file instead of static file test
- Fix CONFIG_FILE path to resolve inside DATA_DIR for volume persistence
- Fix EmbeddingCache singleton to re-init when cache_dir changes
- Fix fd leak in _suppress_output() with nested finally closes
- Fix silent exception on SQLite connection close in upload_tracker
- Log unexpected Frigate API keys at DEBUG in get_all_frigate_person_files
- Add reconcile FIFO-mapping debug log
- Pin all CI action SHAs; update setup-uv v8.2.0, upload/download-artifact,
  ruff-action v4.0.0

* chore: update lockfile

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-14 19:15:20 -04:00
flanandgithub-actions[bot] e795a42e20 refactor: rename CACHE_DIR → DATA_DIR, container path .if_cache → data (#21) (#22)
* refactor: rename CACHE_DIR to DATA_DIR, default path .if_cache → data

CACHE_DIR held both the embedding cache and the SQLite tracker DB, making
the name misleading. DATA_DIR is more accurate.

- Config reads DATA_DIR first; falls back to CACHE_DIR with a deprecation
  warning so existing setups don't break on upgrade
- Default local path: data (was .if_cache)
- Docker default path: /app/data (was /app/.if_cache)
- Internal references (embeddings.py, upload_tracker.py) updated to DATA_DIR
- compose.yml, .env.example, README, wiki, and changelog updated
- Version bumped to 0.5.1

* chore: update lockfile

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-14 18:00:08 -04:00
flanandgithub-actions[bot] d99d607fc8 refactor: rename CACHE_DIR → DATA_DIR, container path .if_cache → data (#21)
* refactor: rename CACHE_DIR to DATA_DIR, default path .if_cache → data

CACHE_DIR held both the embedding cache and the SQLite tracker DB, making
the name misleading. DATA_DIR is more accurate.

- Config reads DATA_DIR first; falls back to CACHE_DIR with a deprecation
  warning so existing setups don't break on upgrade
- Default local path: data (was .if_cache)
- Docker default path: /app/data (was /app/.if_cache)
- Internal references (embeddings.py, upload_tracker.py) updated to DATA_DIR
- compose.yml, .env.example, README, wiki, and changelog updated
- Version bumped to 0.5.1

* chore: update lockfile

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-14 17:59:48 -04:00
flan d30f2956e0 Merge pull request #20 from sudolulo/docs/sync-readme-to-main
docs: sync README accuracy fixes to main
2026-06-14 17:45:11 -04:00
flan 8f30379262 Merge pull request #19 from sudolulo/fix/readme-accuracy
fix: README accuracy for 0.5.0
2026-06-14 17:43:40 -04:00
flan 6c23755654 fix: README accuracy — remove object-mode reference, correct :latest arch to amd64-only 2026-06-14 21:42:33 +00:00
flan 98734d071a Merge pull request #18 from sudolulo/release/workflow-fix
fix: release workflow for main (0.5.0 Docker builds)
2026-06-14 17:30:31 -04:00
flan a77b7b1cc8 Merge pull request #17 from sudolulo/fix/release-workflow-and-docs
fix: release workflow and docs for 0.5.0
2026-06-14 17:25:36 -04:00
flan 4e0e8032ef fix: update release workflow and docs for 0.5.0 single-lockfile refactor
- release.yml: replace per-variant lockfile generation loop with a
  single 'uv lock'; add idempotent release creation (skip if tag
  already has a release so re-triggered runs don't 422)
- docker-publish.yml: remove stale uv-cpu/rocm/intel.lock entries
  from paths-ignore (those files no longer exist)
- README: CACHE_DIR description now names winnow_tracker.db; tracker
  description mentions SQLite
2026-06-14 21:24:06 +00:00
flan 08623088f3 Merge pull request #16 from sudolulo/dev
release: winnow 0.5.0
2026-06-14 17:18:11 -04:00
flan edd22407e9 chore: merge main into dev, resolve lockfile consolidation conflicts 2026-06-14 21:16:29 +00:00
flan d6cb9c6ab8 Merge pull request #15 from sudolulo/release/0.5.0
chore: release 0.5.0
2026-06-14 17:03:23 -04:00
github-actions[bot] fa4fc8984c chore: update lockfile 2026-06-14 21:02:02 +00:00
flan a4571f59f6 chore: release 0.5.0 — version bump, changelog, clean up .env.example 2026-06-14 21:01:41 +00:00
flan 51ea7bd43c Merge pull request #14 from sudolulo/refactor/code-quality
refactor: collapse Config proxy, SQLite tracker, split reconcile, consolidate pyproject
2026-06-14 16:58:52 -04:00
flan 2a2c6b0c47 Merge remote-tracking branch 'origin/dev' into refactor/code-quality 2026-06-14 20:57:31 +00:00
flan 043ebf85d7 fix: prevent silent reject-ID data loss on partial migration rename
Two fixes found during post-refactor audit:

1. upload_tracker: remove COUNT(*) guard from _maybe_migrate. The guard
   blocked re-migration when a previous run successfully committed both
   JSON files but a PermissionError on the second rename() left it on
   disk. On the next startup COUNT > 0 → early return → rejected IDs
   permanently unimported. INSERT OR IGNORE is idempotent so re-running
   migration is always safe; guard not needed.

   Also wrap each rename() in its own try/except so a failure on one
   file is logged and does not propagate uncaught.

2. reconcile: break early when new_count > target is detected in the
   poll loop. Previously the loop ran all four delay intervals (1+2+4+8s)
   before the post-loop > target branch fired, wasting up to 15 seconds
   when a concurrent external upload was visible on the first poll.
2026-06-14 20:54:39 +00:00
flan 886b51fdce Revert "chore: update example URLs to local instance addresses"
This reverts commit caa916a508.
2026-06-14 20:40:45 +00:00
flan caa916a508 chore: update example URLs to local instance addresses 2026-06-14 20:40:27 +00:00
flan cabdb9c0eb fix: address 9 code review findings
- upload_tracker: partial migration now rolls back atomically on failure;
  JSON renamed only after successful commit so failed runs retry cleanly
- upload_tracker: allowlist score_col in _pick_mapped_file to close
  latent SQL injection surface
- config: move load_dotenv() from module import into _load() so no I/O
  at import time and reset() fully resets env loading
- config: use is None checks for IMMICH_URL/OUTPUT_DIR config-file
  fallback so explicitly empty env vars are not overridden by the file
- executor: skip reconcile when Frigate API is unreachable at upload
  start — tracker baseline is incomplete and would mis-trigger the
  external-upload guard, permanently losing file mappings
- reconcile: change polling break condition from >= to == target so
  transient overshoots don't prematurely exit the loop and trigger
  the external-upload guard
- jobs: apply capacity cap as the selection limit rather than truncating
  post-selection by position, so the diversity algorithm works within
  the right budget from the start
- Dockerfile: explicit gpu branch + exit 1 on unknown VARIANT instead
  of silent fallback
2026-06-14 20:37:51 +00:00
flan 71f1924f1a fix: restore original dedup/prompt order in jobs.py, move time import to module level in reconcile.py
_build_job no longer calls filter_already_uploaded internally; callers pass
pre-filtered assets so there's no double DB hit and the interactive path
restores the original prompt order (retry_rejected asked before strategy,
so post-dedup count informs the choice). Skip-count rprint restored in
auto_configure. Late 'import time' inside reconcile_frigate_mappings moved
to module level.
2026-06-14 20:14:55 +00:00
github-actions[bot] 343cb2ad0f chore: update lockfile 2026-06-14 20:04:44 +00:00
flan 02c56493f6 fix: add platform markers to GPU extras, use --extra cpu in CI, simplify lockfile workflow
- rocm/intel/gpu extras are x86_64-only; aarch64 wheels don't exist so uv
  failed to resolve them when required-environments includes aarch64
- test.yml: switch from --all-extras (broken by conflicts + missing wheels)
  to --extra cpu which is cross-platform and sufficient for unit tests
- update-lockfile.yml: drop old file-swap loop; single pyproject means a
  single uv lock run and a single uv.lock to commit
2026-06-14 20:04:25 +00:00
flan e2a1924fb0 refactor: collapse Config proxy, migrate tracker to SQLite, split reconcile module
- Config: remove _ConfigAccessor and ConfigManager; use __getattr__ for lazy
  loading on single _Config class; re-register self as _instance in __getattr__
  so reset() always clears the correct object (item 1)
- upload_tracker: replace hand-rolled JSON store with sqlite3; auto-migrates
  existing JSON on first run; remove dead record_frigate_file function;
  connection re-opens when CACHE_DIR changes for test isolation (items 2, 8)
- diversity: move ThreadPoolExecutor import to module level; inject optional
  fetch_fn parameter for testability (items 3, 6)
- pyproject: consolidate 4 variant files into extras (gpu/rocm/intel/cpu);
  update Dockerfile to use --extra flag; delete variant pyproject/lock files;
  uv.lock needs regen with `uv lock` after this change (item 4)
- jobs: extract _build_job helper to separate business logic from terminal I/O;
  auto_configure delegates dedup/selection to _build_job (item 5)
- logging: convert f-string log calls to % interpolation throughout all winnow/
  modules (item 7)
- reconcile: new module with reconcile_frigate_mappings and
  enrich_asset_with_face_data extracted from executor.py (item 9)
- scheduler: print next scheduled run time after startup and after each run;
  fix f-string logger.error call (item 10)
2026-06-14 19:59:17 +00:00
flan 321b6c66e1 Merge pull request #13 from sudolulo/feature/bump-ubuntu-26-04
build: bump amd64 rocm and cpu bases to Ubuntu 26.04
2026-06-14 15:38:00 -04:00
flan 237dd2091b build: bump amd64 gpu base to NVIDIA CUDA ubuntu24.04
Highest Ubuntu version NVIDIA currently publishes for CUDA 12.8.1.
26.04 not yet available from NVIDIA's image registry.
2026-06-14 19:17:13 +00:00
flan e7bfe00d5d build: bump amd64 rocm and cpu bases to Ubuntu 26.04
22.04 non-GPU amd64 bases were never updated when arm64 moved to 24.04.
Python 3.13 is still pulled from deadsnakes PPA (26.04 ships 3.14 natively).

intel stays on 22.04: the Intel GPU repo URL is pinned to the "jammy"
codename and cannot be bumped until Intel publishes 26.04 packages.
2026-06-14 18:55:50 +00:00
flan ad1fbd4c2a Merge pull request #12 from sudolulo/feature/document-limitations
docs+fix: annotate limitations; fix cache invalidation and version checks
2026-06-14 14:54:48 -04:00
flan 28424b0f16 fix: implement fixable limitations from annotation pass
cache.py — model fingerprint auto-invalidation:
  Replace hardcoded "buffalo_l_v1" version string with a fingerprint
  derived from buffalo_l .onnx file sizes and mtimes. EmbeddingCache now
  computes this at init time; stale embeddings from replaced or updated
  model files are automatically invalidated. Falls back to the static
  string before the model is downloaded.
  Note: existing caches built against the old key will miss on the first
  run after upgrade and recompute cleanly.

frigate_api.py — Frigate version check:
  Add get_frigate_version() (GET /api/version). Called at the start of
  upload_to_frigate(); warns if below v0.16 where the face training API
  endpoints don't exist.

immich_api.py + cli.py — Immich version check:
  Add get_immich_version() (GET /api/server/version). Called at startup
  before get_people(); warns if below v1.106 where the face data and
  merge APIs winnow depends on aren't guaranteed present.

Remaining TODO(frigate-api) annotations are left in place — they require
Frigate to expose per-file embeddings or a rebuild-complete signal before
they can be addressed.
2026-06-14 18:48:07 +00:00
flan 1d44df6e96 docs: annotate known limitations and Frigate API improvement hooks
Adds inline LIMITATION / TODO(frigate-api) comments at each specific
code site rather than a separate doc that would drift from the code.

frigate_api.py — recognize_face:
  Mean-embedding limitation: score reflects the arithmetic mean of all
  training embeddings. A bimodal set (frontals + profiles) has a mean
  between clusters, making both ends look more novel than they are.
  Fixable if Frigate exposes per-file embeddings for nearest-neighbour
  comparison.

frigate_api.py — get_all_frigate_person_files:
  "train" key exclusion is a hardcoded string. If Frigate adds other
  special top-level keys in /api/faces they'll be silently treated as
  person names. Needs a typed schema when Frigate documents the contract.

executor.py — recognize_face call site:
  Async rebuild: each deletion triggers a background model rebuild in
  Frigate. Subsequent recognize calls in the same run return None
  (rebuild in progress), degrading quality replacement for later
  candidates. Fixable with a rebuild-complete signal from Frigate.

executor.py — effective_count / manual file handling:
  Manually-added files are invisible to diversity decisions. Winnow
  observes their effect only indirectly via the Frigate score, not by
  measuring their embedding distribution. Per-file embeddings from
  Frigate would allow direct diversity measurement against the full set.

executor.py — Frigate version assumption:
  All face training endpoints are v0.16+. No version check at startup;
  failures on older versions are opaque 404s.

cache.py — MODEL_VERSIONS:
  Version string is a hardcoded constant. Manual model file replacement
  (custom weights, InsightFace update) won't invalidate cached embeddings.
  Needs file-checksum-derived versioning or a CLEAR_EMBEDDING_CACHE flag.

diversity.py — thumbnail-resolution embeddings:
  Diversity selection runs InsightFace on preview thumbnails; the actual
  training crop comes from full-resolution originals. Negligible in
  practice but degrades if Immich preview quality is low.
2026-06-14 18:41:37 +00:00
flan 4147dbec1c test: add diversity algorithm tests (60 → 93) (#11)
Tests cover the core ML pipeline algorithms in diversity.py — previously
untested. No network or model dependencies; all pure-function or
numpy-only paths:

- Face bbox and confidence extraction from Immich metadata, including
  person_id filtering and missing-data edge cases
- Face crop scaling: verifies bbox coordinates are correctly scaled when
  the thumbnail dimensions differ from the metadata image dimensions
- Near-duplicate dedup: removal below cosine threshold, quality-score
  preference between duplicates, zero-quality-score treated as zero not
  missing (falsy bug guard)
- K-Medoids: correct medoid count, distinctness, valid index range, and
  full-N edge case
- Adaptive threshold: positive output, floor at 0.05 for identical
  embeddings, single-point, scales with embedding spread
- Time-spread fallback: exact count, all-under-limit passthrough,
  auto→30 default, first/last inclusion
- Cluster-aware selection: exact limit, subset invariant, auto-stop on
  tight cluster, hard-example confidence weighting accepted
2026-06-14 14:33:38 -04:00
github-actions[bot] 4b63aafb0e chore: update lockfiles 2026-06-14 18:06:15 +00:00
flan 588a5b2af8 Merge branch 'main' of github.com:sudolulo/winnow 2026-06-14 18:05:14 +00:00
flan ea99ad10e3 Merge branch 'dev' 2026-06-14 18:05:00 +00:00
flan a237983777 chore: update uv.lock for v0.4.11 2026-06-14 18:04:21 +00:00
flan a2d0541493 chore: bump to v0.4.11 and update changelog
Documents object mode removal cleanup, dead embedding path removal,
FutureWarning fix, variant pyproject sync, and docs/wiki updates.
2026-06-14 18:04:05 +00:00
flan 1574aed7e4 docs: expand MERGE_DUPLICATE_PEOPLE docs and fix stale output dir description 2026-06-14 17:52:22 +00:00
flan 0192b6cb5b chore: sync variant pyproject files with main after object pipeline removal
Remove torch/torchvision/transformers/ultralytics from all three variant
pyproject files (rocm/cpu/intel) — no longer needed since SigLIP and YOLO
were removed in v0.4.10. Update version to 0.4.10 and description.

NOTE: uv-rocm.lock, uv-cpu.lock, uv-intel.lock are now stale and must be
regenerated in their respective platform environments before the next Docker
build of those variants.

Also fix fetch_face_data() docstring to not mention embedding (removed field).
2026-06-14 17:51:26 +00:00
flan 0ef15c5c12 chore: remove dead embedding paths and suppress insightface FutureWarning in image_processing
- immich_api.py: remove FaceData.embedding field and numpy import — Immich face
  embeddings were never consumed after being fetched; bbox/confidence is all that's used
- embeddings.py: remove immich_embedding param from get_embedding() — never passed by
  any caller; simplify docstring accordingly
- image_processing.py: wrap insightface_app.get() in warnings filter to suppress the
  scikit-image FutureWarning about estimate being deprecated (already suppressed in
  embeddings.py for the diversity path, was leaking from the crop-alignment path)
- README.md: remove two stale "object mode" / "object classification" fragments
2026-06-14 17:47:30 +00:00
flan 72dbbfa18a chore: remove object classification from package docstring 2026-06-14 17:43:17 +00:00
flan 274d50ee99 chore: remove remaining dead-mode references from executor, jobs, compose, Dockerfile
After removing the object pipeline, several dead 'mode' artifacts remained:

- executor.py: unpack `config` from job even though it was no longer read
- jobs.py: set `"mode": "face"` in both configure paths (key never consumed)
- compose.yml: TRAINING_MODE=face env, OBJECT_CLASS comment, HF_HOME, stale MAX_AUTO_IMAGES default note
- Dockerfile: "and object classification" label, /models/huggingface mkdir, HF_HOME ENV
2026-06-14 17:42:27 +00:00
flan 2d7b52470e chore: remove dead object/SigLIP references after pipeline removal
- cache.py: drop siglip MODEL_VERSIONS entry
- log_config.py: drop ultralytics/transformers/torch from noise silencers
- scheduler.py: drop HF_HOME and HuggingFace model check
- scripts/benchmark.py: drop bench_siglip and make_random_image
2026-06-14 17:31:54 +00:00
flan 835016e0e3 feat: remove object mode pipeline (YOLO, SigLIP, TRAINING_MODE, OBJECT_CLASS)
winnow is a face recognition training tool. Object mode required manual
file placement with no Frigate API, pulled in torch/torchvision/transformers/
ultralytics (~2 GB), and was architecturally misaligned with the project goal.

Removed:
- process_object_mode (YOLO inference), get_yolo_model
- SigLIP model stack (get_siglip_model, get_object_embedding, batch variant)
- entity_type branching throughout diversity, embeddings, jobs, executor
- TRAINING_MODE and OBJECT_CLASS env vars
- torch, torchvision, transformers, ultralytics dependencies
- pytorch index entries from pyproject.toml
- Object mode from README (Modes section, env var table, How It Works)
2026-06-14 17:29:25 +00:00
flan 3e030b361e Update README.md 2026-06-14 13:15:57 -04:00
flan 6ec075b9bf Update README.md 2026-06-14 13:15:20 -04:00
github-actions[bot] 3105b15beb chore: update lockfiles 2026-06-14 17:12:05 +00:00
flan d12abbc543 docs: clarify winnow's role as supplement to manual Frigate training 2026-06-14 17:11:59 +00:00
flan efce4e3443 Merge branch 'dev' of github.com:sudolulo/winnow into dev 2026-06-14 17:11:37 +00:00
flan cb670dd555 docs: clarify winnow fills the gap where manual Frigate training images are absent 2026-06-14 17:11:34 +00:00
github-actions[bot] f027f0a7d1 chore: update lockfiles 2026-06-14 17:09:33 +00:00
flan 4e989b042e Merge branch 'main' of github.com:sudolulo/winnow 2026-06-14 17:08:43 +00:00
flan d63bcfc10b chore: bump version to 0.4.10 2026-06-14 17:08:33 +00:00
flan 650629d3ed release: v0.4.10 2026-06-14 17:08:33 +00:00
flan d7dfc1446a config: lower MAX_AUTO_IMAGES default from 80 to 20 2026-06-14 17:07:52 +00:00
github-actions[bot] 20904b6b22 chore: update lockfiles 2026-06-14 17:06:47 +00:00
github-actions[bot] e47fdaadf6 chore: update lockfiles 2026-06-14 17:06:19 +00:00
flan dfaa03de47 release: v0.4.9 2026-06-14 17:05:46 +00:00
flan 9d5741f626 feat: warn on launch when unsupported advanced tuning vars are set; move ENABLE_FRIGATE_SCORES to unsupported section 2026-06-14 16:59:46 +00:00
flan d70a246a1c docs: move FRIGATE_SCORE_CEILING to unsupported advanced tuning section 2026-06-14 16:56:24 +00:00
flan 51cc7032eb docs: README accuracy audit — novelty gate, rejection scope, calibrated defaults
- How It Works step 9: "upload freely" → note novelty gate may skip below-cap candidates
- Persistence note: "Frigate rejections" → "rejected assets" (covers confidence skips too)
- RETRY_REJECTED description: explicitly covers all rejection types, not just Frigate
- Image Quality section: split into user-adjustable controls and calibrated image
  processing defaults with a support disclaimer to deter blind tuning
2026-06-14 16:54:56 +00:00
flan 105099c819 fix: mark assets rejected when faces API confidence is below threshold
Previously, assets that passed embedding-phase selection but failed the
faces API confidence check in execute_jobs were silently skipped with no
tracker entry. They appeared as valid candidates on every future run,
were re-selected, and re-skipped in an endless cycle. Now they are
marked rejected so they are excluded from future runs. RETRY_REJECTED=true
clears them if Immich later re-processes the image.
2026-06-14 16:46:09 +00:00
flan 0afc9386c6 feat: dynamic Frigate score ceiling; consolidate quality replacement branches
FRIGATE_SCORE_CEILING now defaults to dynamic mode (unset): below-cap
candidates are skipped if their pre-upload Frigate score exceeds the
most-redundant tracked file's score. This catches conditions already
covered by manually-added Frigate images that winnow cannot track —
the embedding-based diversity selection has no visibility into those.
Set FRIGATE_SCORE_CEILING=0 to disable; a positive value (e.g. 0.85)
still acts as a fixed hard ceiling. First-run safety is unchanged
(pre_run_count==0 prevents recognize_face from being called).

The two quality replacement branches (Frigate-score and blur-score)
shared identical structure and are merged into a single code path
parameterised by score source and comparison direction.

Also raises MIN_FACE_COUNT default from 0 to 3 and updates the
config test to match.
2026-06-14 16:38:25 +00:00
flan a59d05e7fd rename STRATEGY=auto to STRATEGY=adaptive; keep auto as alias
AUTO_MODE (batch/unattended) and STRATEGY=auto (diversity algorithm) shared
the same word for unrelated concepts. Renaming the strategy value to
'adaptive' eliminates the ambiguity. The old value is kept as a silent alias
so existing configs continue to work.

Interactive menu updated from "Auto (Objective Diversity)" to "Adaptive
Diversity". README AUTO_MODE description clarified to emphasise unattended
batch processing, not selection strategy.
2026-06-14 16:15:34 +00:00
flan d0cb2e17b1 config: raise MIN_FACE_COUNT default from 0 to 3
People with fewer than 3 tagged photos produce degenerate training sets
and rarely benefit from processing. Skip them by default.
2026-06-14 16:12:23 +00:00
flan 2dc26b5a3d docs: fix CUDA version, add MERGE_DUPLICATE_PEOPLE and TRACE_CROP_SIZE to README
- `:latest` image tag listed CUDA 13.3 — actual base is 12.8.1
- MERGE_DUPLICATE_PEOPLE existed in config but was absent from env var table
- TRACE_CROP_SIZE existed in CLI but was absent from env var table
- RESET_PERSON description now mentions `*` wildcard for bulk reset of all people
2026-06-14 16:11:21 +00:00
flan fe4cfac7b5 docs: add missing dedup step, fix auto-stop description in How It Works
- Step 5 (near-duplicate removal) was added in v0.4.4 but never
  documented; added between embedding and diversity selection
- Auto-stop was described as 'stops when similarity exceeds threshold'
  which is backwards; it stops when the next candidate's distance to
  already-selected images falls below the threshold (too similar, not
  enough new information)
- Renumbered steps 5-8 to 6-9
- Tightened hard-example weighting description to match the code
2026-06-14 16:06:45 +00:00
flan ec074fd279 fix: remove unused record_frigate_file import (ruff F401) 2026-06-14 04:10:15 +00:00
github-actions[bot] 6018bf7222 chore: update lockfiles 2026-06-14 04:09:08 +00:00
flan 4eb6e3169b perf: write-through in-memory cache for tracker JSON reads
All _load() calls after the first return the cached dict instead of
re-reading disk. _save() updates both disk and cache atomically.
Drops per-person tracker reads from ~90 to ~1 in the upload loop.
Keyed by resolved file path so test isolation (unique tmp_path dirs)
is preserved with no fixture changes needed.
2026-06-14 04:08:29 +00:00
github-actions[bot] f5ec9a0001 chore: update lockfiles 2026-06-14 04:06:38 +00:00
flan 7dce4a5c71 perf: eliminate O(K²) dedup allocs, vectorize kmedoids cost, batch tracker writes
- _dedup_embeddings: pre-allocated (Q,D) buffer replaces vstack-on-keep,
  dropping O(K²×D) copy overhead down to O(K×D) fill work
- _kmedoids: swap cost sum replaced with numpy fancy-index reduction,
  ~20-50x faster per swap evaluation
- _reconcile_frigate_mappings: O(L) load/save pairs collapsed to one
  batch write via record_frigate_files_batch
2026-06-14 04:03:55 +00:00
github-actions[bot] 56c802a245 chore: update lockfiles 2026-06-14 03:50:30 +00:00
flan cc092ec972 fix: cap fetch_all_assets at 5000 items; filter non-dict page entries
Fetching up to MAX_PAGES*page_size (1M) assets before the 3000-item
diversity pool cap was applied could exhaust memory on large Immich
libraries. Early-exit once 5000 items are collected — the pool cap
of 3000 makes anything beyond that wasteful. Also filter null/non-dict
items from page responses at fetch time.
2026-06-14 03:49:45 +00:00
github-actions[bot] d6e4582401 chore: update lockfiles 2026-06-14 03:42:28 +00:00
flan a85bc31da9 release: merge dev → main for 0.4.5 2026-06-14 03:41:35 +00:00
flan 3c602e2ef6 fix: code review corrections — dedup O(N²), truncated rejection check, pool warning, path guard
- _dedup_embeddings: rebuild kept_stack only on keep (was every iteration → O(N²))
- _dedup_embeddings: fix quality_score sort key to use explicit None check (falsy-zero)
- _select_by_embedding: add post-dedup pool < limit guard with warning
- executor: use full resp.text for 'face' keyword check; only truncate display snippet
- _safe_person_dir: avoid false "//" prefix when output_dir resolves to filesystem root
2026-06-14 03:40:39 +00:00
flan 96b71798ad docs: add disclaimer that winnow is not an approved Frigate training method 2026-06-14 03:25:38 +00:00
flan a7d4504db9 docs: add disclaimer that winnow is not an approved Frigate training method 2026-06-14 03:24:31 +00:00
github-actions[bot] e05363f632 chore: update lockfiles 2026-06-14 03:23:07 +00:00
github-actions[bot] b276d686f8 chore: update lockfiles 2026-06-14 03:22:57 +00:00
flan a3e54dea7a release: merge dev → main for 0.4.4 2026-06-14 03:22:38 +00:00
flan 44b717d615 release: merge dev → main for 0.4.3 2026-06-14 00:47:06 +00:00
flan 38fe4d6f0c ci: merge dev → main — drop arm64 from GPU build 2026-06-13 22:44:43 +00:00
flan 9c42da4d37 docs: merge dev → main — AI attribution disclosure 2026-06-13 22:31:26 +00:00
github-actions[bot] 68505aeb0b chore: update lockfiles 2026-06-13 22:19:49 +00:00
flanandClaude Sonnet 4.6 0f86c1054a chore: bump version to 0.4.2, update changelog
CUDA base image downgraded to 12.8.1 (driver 570 compatibility fix),
benchmark script added.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 22:19:16 +00:00
flanandClaude Sonnet 4.6 634688fc93 Fix ruff lint errors in benchmark.py
Remove unused imports, fix unsorted imports, remove bare f-strings.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 22:19:16 +00:00
flanandClaude Sonnet 4.6 9f0a78522f Downgrade GPU base image to CUDA 12.8.1; add benchmark script
CUDA 13.3 requires driver >= 575 but the host only has 570 (error 804).
CUDA 12.8.1 is the highest version supported by driver 570 and works
correctly with the NVIDIA Container Toolkit.

Add scripts/benchmark.py to measure InsightFace + SigLIP latency and
throughput across GPU and CPU modes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 22:19:16 +00:00
flanandClaude Sonnet 4.6 8d1f5da05a ci: consolidate Docker builds — eliminate duplicate builds on release
release.yml now calls docker-publish.yml via workflow_call instead of
re-running all four image builds independently. docker-publish.yml gains
workflow_call inputs (tag, version) for release context; branch trigger
is narrowed to dev only (main changes only land via tagged releases).

Each release previously built all four variants twice (~90 min) — once on
merge to main, once on tag push. Now it builds once.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 22:19:16 +00:00
github-actions[bot] a7257cf031 chore: update lockfiles 2026-06-13 19:59:04 +00:00
flanandClaude Sonnet 4.6 326fdbdf38 release: merge dev → main for 0.4.1
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 19:58:19 +00:00
flanandClaude Sonnet 4.6 91e0858aa6 refactor: merge dev → main — post-0.4.0 cleanup
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 18:55:42 +00:00
github-actions[bot] 71df0e81de chore: update lockfiles 2026-06-13 18:36:36 +00:00
flanandClaude Sonnet 4.6 9bb0727807 release: merge dev → main for 0.4.0
Frigate pre-upload scoring, quality replacement inversion, bootstrap fix,
FRIGATE_SCORE_CEILING / ENABLE_FRIGATE_SCORES, removal of post-upload gate.
See CHANGELOG.md [0.4.0] for the full list.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 18:36:02 +00:00
flanandClaude Sonnet 4.6 7c306a4423 chore: bump version to 0.3.3, update changelog and lockfile
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 15:45:06 +00:00
flanandClaude Sonnet 4.6 c22857b912 fix: raise MIN_FACE_WIDTH default from 50 to 90px (8k pixel floor)
50px crops produce ~2,500–4,225 total pixels — well below Frigate's own
camera capture range of 16k–50k px. 90px guarantees ≥8,100 total pixels
even when face margins are fully clipped by image edges.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 15:45:06 +00:00
github-actions[bot] c53172f5ff chore: update lockfiles 2026-06-13 15:33:27 +00:00
flanandClaude Sonnet 4.6 9598142997 release: v0.3.2
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 15:32:33 +00:00
40 changed files with 2717 additions and 9315 deletions
+8 -12
View File
@@ -8,19 +8,16 @@ FRIGATE_URL=http://192.168.1.10:5000
# Set AUTO_MODE=true to force auto mode even in an interactive terminal. # Set AUTO_MODE=true to force auto mode even in an interactive terminal.
# AUTO_MODE=true # AUTO_MODE=true
# VERBOSE=true # Enable DEBUG-level console output (log file is always DEBUG) # VERBOSE=true # Enable DEBUG-level console output (log file is always DEBUG)
# TRAINING_MODE: face = upload to Frigate face recognition API # STRATEGY: adaptive = embedding diversity (recommended), standard = 30 imgs, broad = 100 imgs
# object = save crops to output dir for manual Frigate placement STRATEGY=adaptive
TRAINING_MODE=face
# STRATEGY: auto = objective diversity (recommended), standard = 30 imgs, broad = 100 imgs
STRATEGY=auto
# LIMIT=50 # Custom image count; overrides STRATEGY preset # LIMIT=50 # Custom image count; overrides STRATEGY preset
# OBJECT_CLASS=dog # Object label for object mode (e.g. dog, cat, car)
# ── People Filtering ────────────────────────────────────────────────────────── # ── People Filtering ──────────────────────────────────────────────────────────
# ONLY_PEOPLE=John,Jane # Comma-separated; process only these people # ONLY_PEOPLE=John,Jane # Comma-separated; process only these people
# SKIP_PEOPLE=Unknown # Comma-separated; skip these people # SKIP_PEOPLE=Unknown # Comma-separated; skip these people
# MIN_FACE_COUNT=5 # Skip people with fewer than N assets in Immich # MIN_FACE_COUNT=3 # Skip people with fewer than N assets in Immich (default: 3)
# YEARS_FILTER=10 # Only include images from the last N years (default: 10) # YEARS_FILTER=10 # Only include images from the last N years (default: 10)
# MERGE_DUPLICATE_PEOPLE=false # Merge duplicate Immich person records permanently (default: false — warn and skip)
# ── Image Quality ───────────────────────────────────────────────────────────── # ── Image Quality ─────────────────────────────────────────────────────────────
# MIN_FACE_WIDTH=90 # Minimum face width in pixels (default: 90, guarantees ≥8,100px crop) # MIN_FACE_WIDTH=90 # Minimum face width in pixels (default: 90, guarantees ≥8,100px crop)
@@ -29,22 +26,21 @@ STRATEGY=auto
# USE_FULL_RESOLUTION=true # Use full-res images vs thumbnails (default: true) # USE_FULL_RESOLUTION=true # Use full-res images vs thumbnails (default: true)
# MIN_CONFIDENCE=0.7 # Minimum face detection confidence (default: 0.7) # MIN_CONFIDENCE=0.7 # Minimum face detection confidence (default: 0.7)
# BLUR_THRESHOLD=120.0 # Laplacian blur threshold; lower = accept more blur (default: 120.0) # BLUR_THRESHOLD=120.0 # Laplacian blur threshold; lower = accept more blur (default: 120.0)
# MAX_AUTO_IMAGES=80 # Hard cap on auto-diversity selection (default: 80) # MAX_AUTO_IMAGES=20 # Hard cap on auto-diversity selection (default: 20)
# QUALITY_REPLACEMENT=true # At cap, replace a weaker tracked image with a better candidate (default: true) # QUALITY_REPLACEMENT=true # At cap, replace a weaker tracked image with a better candidate (default: true)
# FRIGATE_SCORE_CEILING=0.0 # Skip uploads already well-covered (pre-upload score > ceiling = redundant; 0 = disabled; requires at least one prior run) # FRIGATE_SCORE_CEILING= # Below-cap novelty gate: unset = dynamic (default), 0 = disabled, e.g. 0.85 = fixed ceiling
# ENABLE_FRIGATE_SCORES=true # Call Frigate's recognize endpoint pre-upload to store diversity scores (default: true; adds ~200ms per upload) # ENABLE_FRIGATE_SCORES=true # Call Frigate's recognize endpoint pre-upload to store diversity scores (default: true; adds ~200ms per upload)
# ── Caching & Models ────────────────────────────────────────────────────────── # ── Caching & Models ──────────────────────────────────────────────────────────
# FORCE_CPU=true # Disable GPU, fall back to CPU # FORCE_CPU=true # Disable GPU, fall back to CPU
# ENABLE_CACHE=false # Disable embedding cache (default: true) # ENABLE_CACHE=false # Disable embedding cache (default: true)
CACHE_DIR=/app/.if_cache DATA_DIR=/app/data
HF_HOME=/models/huggingface
INSIGHTFACE_HOME=/models/.insightface INSIGHTFACE_HOME=/models/.insightface
# ── Tracker overrides (one-shot — remove after use) ─────────────────────────── # ── Tracker overrides (one-shot — remove after use) ───────────────────────────
# DRY_RUN=true # Preview selection without downloading/uploading # DRY_RUN=true # Preview selection without downloading/uploading
# RETRY_REJECTED=true # Re-attempt previously rejected images # RETRY_REJECTED=true # Re-attempt previously rejected images
# RESET_PERSON=John # Clear uploaded+rejected history for one person # RESET_PERSON=John # Clear uploaded+rejected history for one person (use * for all)
# ── Scheduling ──────────────────────────────────────────────────────────────── # ── Scheduling ────────────────────────────────────────────────────────────────
# CRON_SCHEDULE controls container lifetime: # CRON_SCHEDULE controls container lifetime:
+7
View File
@@ -0,0 +1,7 @@
#!/usr/bin/env bash
set -euo pipefail
if git diff --cached --name-only | grep -q "^pyproject\.toml$"; then
uv lock
git add uv.lock
fi
+21 -24
View File
@@ -12,9 +12,6 @@ on:
- ".github/workflows/lint.yml" - ".github/workflows/lint.yml"
- ".github/dependabot.yml" - ".github/dependabot.yml"
- "uv.lock" - "uv.lock"
- "uv-cpu.lock"
- "uv-rocm.lock"
- "uv-intel.lock"
workflow_call: workflow_call:
inputs: inputs:
tag: tag:
@@ -57,15 +54,15 @@ jobs:
df -h df -h
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@v6 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with: with:
ref: ${{ inputs.tag || github.ref }} ref: ${{ inputs.tag || github.ref }}
- name: Set up Docker Buildx - name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4 uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4.1.0
- name: Log in to GHCR - name: Log in to GHCR
uses: docker/login-action@v4 uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4.2.0
with: with:
registry: ${{ env.REGISTRY }} registry: ${{ env.REGISTRY }}
username: ${{ github.actor }} username: ${{ github.actor }}
@@ -82,7 +79,7 @@ jobs:
- name: Build and push by digest - name: Build and push by digest
id: build id: build
uses: docker/build-push-action@v7 uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7.2.0
with: with:
context: . context: .
file: ./Dockerfile file: ./Dockerfile
@@ -100,7 +97,7 @@ jobs:
touch "/tmp/digests/${digest#sha256:}" touch "/tmp/digests/${digest#sha256:}"
- name: Upload digest - name: Upload digest
uses: actions/upload-artifact@v4 uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with: with:
name: digest-amd64 name: digest-amd64
path: /tmp/digests/* path: /tmp/digests/*
@@ -117,17 +114,17 @@ jobs:
steps: steps:
- name: Download digests - name: Download digests
uses: actions/download-artifact@v4 uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with: with:
path: /tmp/digests path: /tmp/digests
pattern: digest-* pattern: digest-*
merge-multiple: true merge-multiple: true
- name: Set up Docker Buildx - name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4 uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4.1.0
- name: Log in to GHCR - name: Log in to GHCR
uses: docker/login-action@v4 uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4.2.0
with: with:
registry: ${{ env.REGISTRY }} registry: ${{ env.REGISTRY }}
username: ${{ github.actor }} username: ${{ github.actor }}
@@ -181,18 +178,18 @@ jobs:
df -h df -h
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@v6 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with: with:
ref: ${{ inputs.tag || github.ref }} ref: ${{ inputs.tag || github.ref }}
- name: Set up QEMU - name: Set up QEMU
uses: docker/setup-qemu-action@v4 uses: docker/setup-qemu-action@06116385d9baf250c9f4dcb4858b16962ea869c3 # v4.1.0
- name: Set up Docker Buildx - name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4 uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4.1.0
- name: Log in to GHCR - name: Log in to GHCR
uses: docker/login-action@v4 uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4.2.0
with: with:
registry: ${{ env.REGISTRY }} registry: ${{ env.REGISTRY }}
username: ${{ github.actor }} username: ${{ github.actor }}
@@ -225,7 +222,7 @@ jobs:
fi fi
- name: Build and push CPU image - name: Build and push CPU image
uses: docker/build-push-action@v7 uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7.2.0
with: with:
context: . context: .
file: ./Dockerfile file: ./Dockerfile
@@ -267,15 +264,15 @@ jobs:
df -h df -h
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@v6 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with: with:
ref: ${{ inputs.tag || github.ref }} ref: ${{ inputs.tag || github.ref }}
- name: Set up Docker Buildx - name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4 uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4.1.0
- name: Log in to GHCR - name: Log in to GHCR
uses: docker/login-action@v4 uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4.2.0
with: with:
registry: ${{ env.REGISTRY }} registry: ${{ env.REGISTRY }}
username: ${{ github.actor }} username: ${{ github.actor }}
@@ -308,7 +305,7 @@ jobs:
fi fi
- name: Build and push ROCm image - name: Build and push ROCm image
uses: docker/build-push-action@v7 uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7.2.0
with: with:
context: . context: .
file: ./Dockerfile file: ./Dockerfile
@@ -350,15 +347,15 @@ jobs:
df -h df -h
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@v6 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with: with:
ref: ${{ inputs.tag || github.ref }} ref: ${{ inputs.tag || github.ref }}
- name: Set up Docker Buildx - name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4 uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4.1.0
- name: Log in to GHCR - name: Log in to GHCR
uses: docker/login-action@v4 uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4.2.0
with: with:
registry: ${{ env.REGISTRY }} registry: ${{ env.REGISTRY }}
username: ${{ github.actor }} username: ${{ github.actor }}
@@ -391,7 +388,7 @@ jobs:
fi fi
- name: Build and push Intel image - name: Build and push Intel image
uses: docker/build-push-action@v7 uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7.2.0
with: with:
context: . context: .
file: ./Dockerfile file: ./Dockerfile
+2 -2
View File
@@ -23,10 +23,10 @@ jobs:
echo "Disk space freed." echo "Disk space freed."
- name: Checkout code - name: Checkout code
uses: actions/checkout@v6 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Run Ruff - name: Run Ruff
uses: astral-sh/ruff-action@v3 uses: astral-sh/ruff-action@0ce1b0bf8b818ef400413f810f8a11cdbda0034b # v4.0.0
with: with:
args: "check" args: "check"
+24 -23
View File
@@ -32,12 +32,12 @@ jobs:
echo "Disk space freed." echo "Disk space freed."
- name: Checkout - name: Checkout
uses: actions/checkout@v6 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with: with:
fetch-depth: 0 fetch-depth: 0
- name: Install uv - name: Install uv
uses: astral-sh/setup-uv@v7 uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # v8.2.0
- name: Set up Python - name: Set up Python
run: uv python install 3.13 run: uv python install 3.13
@@ -50,17 +50,8 @@ jobs:
exit 1 exit 1
fi fi
- name: Ensure lockfiles are current - name: Ensure lockfile is current
run: | run: uv lock
cp pyproject.toml _pyproject_orig.toml
for variant in cpu rocm intel; do
cp pyproject-${variant}.toml pyproject.toml
uv lock
cp uv.lock uv-${variant}.lock
done
cp _pyproject_orig.toml pyproject.toml
uv lock
rm _pyproject_orig.toml
- name: Resolve tag name - name: Resolve tag name
id: tag id: tag
@@ -109,7 +100,7 @@ jobs:
fi fi
- name: Create GitHub Release - name: Create GitHub Release
uses: actions/github-script@v9 uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9.0.0
env: env:
RELEASE_TAG: ${{ steps.tag.outputs.TAG }} RELEASE_TAG: ${{ steps.tag.outputs.TAG }}
RELEASE_NOTES: ${{ steps.changelog.outputs.NOTES }} RELEASE_NOTES: ${{ steps.changelog.outputs.NOTES }}
@@ -117,15 +108,25 @@ jobs:
script: | script: |
const tag = process.env.RELEASE_TAG; const tag = process.env.RELEASE_TAG;
const notes = (process.env.RELEASE_NOTES || '').trim(); const notes = (process.env.RELEASE_NOTES || '').trim();
await github.rest.repos.createRelease({ try {
owner: context.repo.owner, const existing = await github.rest.repos.getReleaseByTag({
repo: context.repo.repo, owner: context.repo.owner,
tag_name: tag, repo: context.repo.repo,
name: `Release ${tag}`, tag: tag,
body: notes || `Release ${tag}`, });
draft: false, console.log(`Release ${tag} already exists (id ${existing.data.id}), skipping creation.`);
prerelease: false, } catch (err) {
}); if (err.status !== 404) throw err;
await github.rest.repos.createRelease({
owner: context.repo.owner,
repo: context.repo.repo,
tag_name: tag,
name: `Release ${tag}`,
body: notes || `Release ${tag}`,
draft: false,
prerelease: false,
});
}
build-images: build-images:
name: Build and push Docker images name: Build and push Docker images
+6 -3
View File
@@ -13,16 +13,19 @@ jobs:
contents: read contents: read
steps: steps:
- name: Checkout code - name: Checkout code
uses: actions/checkout@v6 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Install uv - name: Install uv
uses: astral-sh/setup-uv@v7 uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # v8.2.0
- name: Set up Python - name: Set up Python
run: uv python install 3.13 run: uv python install 3.13
- name: Check lockfile is up to date
run: uv lock --check
- name: Install dependencies - name: Install dependencies
run: uv sync --all-extras run: uv sync --extra cpu
- name: Run tests - name: Run tests
run: uv run pytest run: uv run pytest
-66
View File
@@ -1,66 +0,0 @@
# .github/workflows/update-lockfile.yml
name: Update lockfiles
on:
push:
branches:
- '**'
paths:
- 'pyproject.toml'
- 'pyproject-cpu.toml'
- 'pyproject-rocm.toml'
- 'pyproject-intel.toml'
workflow_dispatch:
jobs:
update-lockfile:
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- name: Free up disk space
run: |
sudo rm -rf /usr/share/dotnet
sudo rm -rf /opt/ghc
sudo rm -rf "/usr/local/share/boost"
sudo rm -rf "$AGENT_TOOLSDIRECTORY"
echo "Disk space freed."
- name: Checkout repository
uses: actions/checkout@v6
- name: Install uv
uses: astral-sh/setup-uv@v7
- name: Set up Python
run: uv python install 3.13
- name: Regenerate all lockfiles
run: |
cp pyproject.toml _pyproject_orig.toml
for variant in cpu rocm intel; do
cp pyproject-${variant}.toml pyproject.toml
uv lock
cp uv.lock uv-${variant}.lock
done
cp _pyproject_orig.toml pyproject.toml
uv lock
rm _pyproject_orig.toml
- name: Check for changes
id: diff
run: |
if git diff --quiet uv.lock uv-cpu.lock uv-rocm.lock uv-intel.lock; then
echo "changed=false" >> "$GITHUB_OUTPUT"
else
echo "changed=true" >> "$GITHUB_OUTPUT"
fi
- name: Commit and push updated lockfiles
if: steps.diff.outputs.changed == 'true'
run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git add uv.lock uv-cpu.lock uv-rocm.lock uv-intel.lock
git commit -m "chore: update lockfiles"
git push
+248
View File
@@ -7,6 +7,254 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased] ## [Unreleased]
## [0.6.6] - 2026-06-18
### Changed
- **`MAX_AUTO_IMAGES` default lowered from 20 to 5** — existing users who have not set this variable and already have more than 5 winnow-managed images in Frigate will find themselves at cap on the next run. With `QUALITY_REPLACEMENT=true` (the default), winnow will attempt to swap weaker images rather than uploading new ones. Set `MAX_AUTO_IMAGES=20` to restore the previous behaviour.
## [0.6.5] - 2026-06-17
### Added
- **Version displayed in startup banner**: winnow now prints its installed version at launch.
### Fixed
- **GPU image: `CUDAExecutionProvider` missing due to parallel install race**: `insightface` declares `onnxruntime` (CPU) as a dependency, causing `uv sync` to install both `onnxruntime` and `onnxruntime-gpu` in parallel — both packages claim the same `pybind11_state.so` binary. On GitHub Actions the CPU binary consistently won the race, leaving the GPU build without CUDA support at runtime despite all CUDA libraries being present. Fixed by reinstalling `onnxruntime-gpu` sequentially after `uv sync` to guarantee its GPU binary is on disk.
- **GPU extra was missing three required nvidia pip packages**: `onnxruntime-gpu` 1.26.0 gates CUDA EP loading on the Python-importability of `nvidia-cuda-runtime-cu12`, `nvidia-cufft-cu12`, and `nvidia-curand-cu12`. These packages were not declared in the `gpu` extra and were absent on fresh installs, silently disabling GPU inference.
- **`_handle_duplicate_people` raises `KeyError` on id-less person records**: bare `p["id"]` subscripts in the auto-merge loop and `_smaller_duplicate_ids` raised `KeyError` when Immich returned a person dict without an `id` field (e.g. unconfirmed face clusters). Fixed by using `p.get("id")` and filtering `None` from `skip_ids`.
- **`_smaller_duplicate_ids` could include `None` in the skip set**: `p.get("id")` without a `None` guard populated `skip_ids` with `None`, causing `p.get("id") not in skip_ids` to pass for every id-less person, so unnamed face clusters were silently re-included in all return paths.
- **`_handle_duplicate_people` dead code removed**: guards `if not survivor_id` and `if not merge_ids` became unreachable after the id-gate fix; their presence suggested they still ran.
- **`_valid_people` in `jobs.py` used wrong name filter**: whitespace-only names (e.g. `" "`) passed the `p.get("name")` truthiness check and were included in the person list. Fixed using `(p.get("name") or "").strip()` consistent with the cli.py gate.
- **`interactive_configure` queued-marker check was O(N²)**: `[j for j in jobs if j["person"]["id"] == p.get("id")]` ran a full scan over jobs for every person in the display loop. Replaced with a `queued_ids` set hoisted before the loop.
- **`executor.py` slot restore did not clear `min_quality_score_for_slot`**: when a replacement upload failed all retries after a deletion, `effective_count` was restored but the stale quality-score floor from the deleted file remained, blocking the next candidate from filling the slot.
- **`get_immich_version` swallowed `KeyError` on unexpected schema**: bare `data["major"]` / `data["minor"]` / `data["patch"]` subscripts were silently caught by the surrounding `except Exception`, returning `None` without logging. Replaced with `.get()` calls that log a debug warning on unexpected schemas.
- **Face embedding selects nearest face to crop centre, not largest by area**: a 25 % margin on the crop window can pull a larger neighbouring face into the bounding box; selecting the biggest face by area then embeds the wrong person. Centre-proximity is now used instead.
- **Zero-norm face embeddings skipped before diversity selection**: InsightFace occasionally returns a zero vector for low-quality detections; zero embeddings pass deduplication with similarity 0 and score distance 1.0, causing them to be selected first as maximally diverse.
- **`executor.py` slot restore did not clear `min_quality_score_for_slot`**: stale quality floor from the deleted file blocked the next candidate from filling the restored slot in quality-replacement mode.
## [0.6.4] - 2026-06-17
### Fixed
- **Face bbox scaled to thumbnail space before quality filtering**: `assess_quality` now receives coordinates in thumbnail-pixel space rather than detection-image space. Previously, a face detected on a full-resolution image (e.g. 4000 px wide) was compared against `MIN_FACE_WIDTH` using its original pixel dimensions, causing faces that appear small on the thumbnail to pass the quality filter — and faces that appear large to be incorrectly rejected.
- **`conf_array` default restored to 1.0 for faces with missing confidence**: the default was incorrectly set to 0.5, causing images with no `score` field in the Immich faces API response to receive a 1.7× FPS diversity boost and be selected ahead of genuinely high-confidence detections. The default is now 1.0 (no boost), treating missing confidence as neutral.
- **`hard_weight` computed once outside FPS loop**: `conf_array` is constant after initialisation; moving the `np.where` call outside the `while` loop eliminates one O(n) numpy pass per selected image.
- **`has_frigate_model` snapshot prevents mid-batch `recognize_face` calls on first run**: `effective_count` is incremented inside the upload loop, so using it as the `recognize_face` gate would incorrectly trigger scoring after the first upload on a first run. A boolean snapshot is now taken before the loop.
- **`person_has_fscores` only set when tracker write succeeds**: the flag was moved outside the `try/except else` block, causing at-cap replacement to switch into Frigate-score mode even when the score was never written to the tracker — `get_most_redundant_mapped_file` then returned `None` and all replacement candidates were silently skipped. The flag is now set only in the `else` branch.
- **`STRATEGY=skip` honoured before embedding and limit checks**: the strategy was silently converted to `auto` when InsightFace was available, because two early-returns in `_resolve_strategy` ran before the `strategy_map` lookup.
- **`limit="auto"` preserved on first run**: switching to `limit = capacity` unconditionally caused the FPS adaptive early-stop to never fire on a person's first upload run. `limit="auto"` is now kept when `already_uploaded == 0`.
- **`EmbeddingCache.get` falls back gracefully on all load errors**: a `MemoryError` during `np.load` of a cached embedding was re-raised, crashing the entire diversity-selection batch for that person. Cache-read failures of any kind now return `None` so the embedding is recomputed fresh.
- **`get_people` returns `[]` when Immich sends `{"people": null}`**: `.get("people", [])` only uses the default when the key is absent, not when its value is `null`. Changed to `data.get("people") or []` so null-valued responses are handled the same as missing keys.
- **`get_people` and `fetch_all_assets` guard against non-dict responses**: a proxy or CDN returning a JSON array (or other non-dict body) previously caused an `AttributeError` from `.get()`. Both functions now check `isinstance(data, dict)` and return an empty result with an error log.
- **`filter_recent_assets` counts and logs assets with missing or unparseable timestamps** instead of silently dropping them.
- **`_suppress_output` fd cleanup restructured**: the context manager now initialises `devnull_fd`, `saved_out`, and `saved_err` to `None` before the `try` block, so the `finally` can close only the descriptors that were successfully opened. Each `os.close` is wrapped in its own `try/except OSError` so a failed close cannot prevent subsequent descriptors from being released. `OSError` from `os.dup2` restore is logged at DEBUG rather than silently swallowed.
- **`blur_score_from_image` copies the image before thumbnail resize**: `Image.thumbnail` modifies the image in-place. When the caller's image was already in RGB mode (no convert copy), the resize would have mutated the caller's object. A copy is now made when `score_img is img`.
- **`imageWidth`/`imageHeight` zero-value treated as missing** in `image_processing.py`: the old `or img_w` fallback silently set `scale = 1.0` for a zero-valued dimension (correct) but also for `None` (also correct) with no distinction. The explicit `scale = img_w / meta_w if meta_w else 1.0` form matches the pattern used in the new `_scale_bbox_to_thumbnail` helper and makes the fallback intent clear.
- **`_mark` and `update_frigate_count` copy before mutate**: both functions now create a shallow copy of the top-level tracker dict before assigning into `by_person`, so a failed `_save` cannot leave the in-memory cache ahead of the on-disk file.
- **`reset_person` flat-list guard only warns when cleanup would have run**: the `isinstance(data[flat_key], list)` check previously emitted a warning even when `person_ids` was empty (a no-op call). The warning is now gated behind `person_ids and`, matching the guard on the cleanup branch.
- **`_handle_duplicate_people` uses `p.get("id")` consistently**: all four return-path filter comprehensions and the `_smaller_duplicate_ids` set comprehension now use `.get("id")` instead of bare `p["id"]`, preventing a `KeyError` if the Immich API returns a person record without an `id` field.
- **`K-Medoids` non-medoid membership test is O(1)**: `non_medoids` now filters against `set(medoids)` instead of the list, eliminating an O(k) scan per candidate on each outer iteration.
## [0.6.3] - 2026-06-16
### Fixed
- **`record_frigate_files_batch` no longer mutates the tracker cache before write**: the function shared the same cache-corruption-on-write-failure bug that was fixed in `remove_frigate_files_batch` in v0.6.1 — `data.setdefault("by_person", {})` mutated the cached dict in-place, so a disk-full or permission error left the in-memory cache ahead of the on-disk file. Now uses the same copy-before-mutate pattern (shallow copies of the top-level dict and `by_person` sub-dict) so a failed write leaves cache and disk in sync.
- **`tracker_ok` boolean flag replaced with try/else**: the intermediate boolean was a misleading placeholder — the `True` initial value suggested success before the operation ran. The control flow is now expressed directly with a try/except/else block.
- **`LIMIT` env var guard simplified**: the two adjacent `if custom_limit is not None` checks in `_resolve_strategy` are collapsed into a single `if custom_limit is not None:` with nested branches, removing redundant evaluation.
## [0.6.2] - 2026-06-16
### Changed
- **Flat `uploaded_asset_ids` / `rejected_asset_ids` lists dropped as primary storage**: asset IDs are now derived on read from `by_person` entries, which are the single source of truth. The legacy flat lists in existing tracker files are still read (union) so no assets become re-eligible after upgrading. New writes no longer maintain the flat lists. This removes the dual-representation sync hazard and paves the way for multi-instance support (per-instance `by_person` keying in a future release).
- **Tracker writes batched per person**: `mark_uploaded` calls inside the per-person upload loop are now accumulated in memory (`begin_batch`) and flushed in a single `os.replace` write at the end of each person's loop (`flush_batch`), reducing N tracker writes per person to 1. Benefits users on slow storage (NAS, SD card, spinning disks).
- **`RESET_PERSON=*` is now O(1) disk writes**: replaced the per-person `reset_person` loop with `reset_all_people()`, which makes one Frigate API call per person for file deletion and then clears both tracker files in two writes. Previously it was O(P²) iterations and 2P writes.
- **`blur_score_from_image` inlines Laplacian computation**: replaced the `assess_quality()` call (which ran grayscale, exposure, and confidence checks whose results were discarded) with a direct `cv2.Laplacian` computation. The function is now self-contained and does not silently inherit future costs added to the full quality pipeline.
## [0.6.1] - 2026-06-16
### Fixed
- **Corrupt or truncated full-res thumbnails now marked rejected**: `OSError` (truncated file) is caught alongside `PIL.UnidentifiedImageError` in the thumbnail path so persistently bad assets are tombstoned instead of retried forever. Full-res download failures (`USE_FULL_RESOLUTION=true`) remain transient — not marked rejected — so a Immich blip doesn't permanently blacklist valid assets.
- **Quality replacement mode no longer flips mid-loop**: `person_has_fscores` was re-evaluated after each file deletion, which could switch the remaining replacements from Frigate-score mode to blur-score mode if the deleted file was the last scored one. The mode is now fixed for the duration of the upload loop.
- **`reset_person` no longer removes shared asset IDs**: the flat `uploaded_asset_ids` list is now rebuilt from all remaining `by_person` entries rather than subtracting the reset person's IDs. Previously, resetting Alice could remove an asset ID that also appeared under Bob, making it re-eligible for upload.
- **`_save` cache updated only after successful write**: the in-memory tracker cache is now updated after `os.replace` succeeds rather than before. A disk-full or permission error no longer leaves the cache permanently ahead of the on-disk file.
- **Stale Frigate file cleanup batched**: the per-file `remove_frigate_file` loop is replaced with a single `remove_frigate_files_batch` call, reducing N tracker writes to 1 when stale mappings are cleaned up.
- **`_migrate_entry` no longer mutates the cache through nested dict aliases**: all five nested dicts (`asset_ids`, `scores`, `frigate_scores`, `frigate_files`, `crop_dims`) are now individually copied so `.pop()` calls in write paths cannot reach the in-memory cache.
- **`find_by_crop_dimension` and `_pick_mapped_file` now agree on duplicate asset→file handling**: both use first-seen-wins when the same `asset_id` maps to multiple Frigate filenames, preventing inconsistent replacement decisions.
- **Non-atomic JSON write**: tracker files are written to a `.tmp` sibling then renamed with `os.replace` so a crash mid-write never leaves a truncated file.
- **`get_person_summary` uses `_migrate_entry`**: replaced three ad-hoc `isinstance` guards with a single `_migrate_entry` call, making old-format (list) entries consistent with every other read path.
- **Quality replacement floor check**: a candidate with a `None` blur score (PIL error during scoring) no longer blocks a freed slot — the `<=` floor comparison is only applied when a score is actually available.
- **`executor.py` syntax error**: the `if img is None:` block in the full-res download path was comment-only and would have raised `IndentationError` on import. Added `pass`.
- **Duplicate `if stale:` guard**: two consecutive identical guards around stale-cleanup and its log print were merged into one.
- **`_flat_key` uses constant equality** instead of substring match, removing a latent routing bug for any filename that happens to contain "uploaded".
- **`remove_frigate_file` no longer creates ghost entries**: returns early when the person is absent rather than writing an empty stub.
- **`skip_ids` extracted to helper**: the identical set comprehension in `_handle_duplicate_people` that appeared in three branches is now a single `_smaller_duplicate_ids()` inner function.
- **`blur_score_from_image` returns `None` on error** instead of `0.0`, so callers can distinguish a failed measurement from a legitimately near-zero Laplacian variance score.
## [0.6.0] - 2026-06-15
### Changed
- **Upload tracker reverted to JSON storage**: the SQLite-based tracker introduced in v0.5.0 produced 17 bug-fix releases in two days due to data-loss risks in the migration layer, schema primary key conflicts, tracker isolation races, and disk-full retry storms. The JSON backend (`frigate_uploaded_ids.json` / `frigate_rejected_ids.json` in `DATA_DIR`) is restored. It is simpler, has no migration layer, and carries no external dependency. If you ran any v0.5.x version, delete `frigate_tracker.db` from your `DATA_DIR` once you confirm the JSON files look correct. JSON files from before v0.5.0 are read automatically with no changes required.
- **`CACHE_DIR` env var accepted as `DATA_DIR` alias**: the rename introduced in v0.5.1 is preserved — `CACHE_DIR` still works with a deprecation warning. The default data path remains `data` (Docker: `/app/data`).
- **Config file now lives in `DATA_DIR`**: `.immich_config.json` resolves to `DATA_DIR/.immich_config.json` so it persists across container restarts. The legacy CWD location is still checked as a fallback for existing setups.
- **Diversity selector receives capacity as its limit directly**: instead of selecting up to `MAX_AUTO_IMAGES` and then slicing to the remaining capacity, the selector now runs with the actual remaining slot count as its budget.
### Fixed
- **Immich v2.7.5 compatibility**: `auto_configure` no longer pre-filters people by `assetCount` from `/api/people`, which Immich v2.7.5 dropped. The `MIN_FACE_COUNT` check now runs after `fetch_all_assets` using the actual fetched count.
- **`fetch_face_data` no longer falls back to a wrong person's bounding box**: when `person_id` is provided but not found in the Immich `/api/faces` response, the function now returns `None` instead of using `faces[0]`. Previously a group photo where the target person's face entry was missing would inject a different person's bounding box into the crop.
- **Corrupt thumbnail permanently rejected**: when `resp.ok=True` but `PIL.UnidentifiedImageError` is raised (Pillow cannot identify the image format), the asset is now marked rejected so it isn't re-downloaded on every future run. Transient `OSError`/truncation errors are intentionally not caught here — those are retried normally.
- **`mark_uploaded` tracker failure no longer aborts the upload loop**: a tracker write failure after a successful Frigate POST is logged and the loop continues; the asset will be re-uploaded on the next run rather than the current run dying mid-job.
- **`progress.remove_task` now in `finally` block**: the progress bar task is cleaned up even when a job exits via an exception, preventing orphaned progress rows in the terminal.
- **`SKIP_PEOPLE`/`ONLY_PEOPLE` now strip whitespace**: `"Alice, Bob".split(",")` produces `[" Bob"]`; the leading space now stripped so comma-separated values with spaces work as expected.
- **`FRIGATE_URL` with trailing slash no longer produces double-slash paths**: all Frigate API calls now use `_get_frigate_url()` for URL normalization rather than reading `FRIGATE_URL` inline.
- **Frigate version `v`-prefix now stripped**: `v0.16.0`-style version strings are correctly parsed.
- **Invalid numeric env var values warn and use defaults**: a typo such as `YEARS_FILTER=10 ` (trailing space) or `MIN_FACE_WIDTH=auto` now logs a `WARNING` and falls back to the documented default instead of raising `ValueError` at startup. Affects `YEARS_FILTER`, `MIN_FACE_WIDTH`, `MIN_FACE_COUNT`, `MAX_AUTO_IMAGES`, `BLUR_THRESHOLD`, `MIN_CONFIDENCE`, and `FACE_MARGIN`.
- **`IMMICH_URL` blank placeholder falls back to config file**: `IMMICH_URL=` (empty or blank) in `.env` is now treated as unset and falls through to `DATA_DIR/.immich_config.json`, matching pre-v0.5.0 behaviour.
- **Reconciliation checks Frigate immediately before first sleep**: the poll loop now performs an immediate check after upload, then backs off with `(1, 2, 4, 8)` s delays only if needed.
- **Dockerfile unknown `VARIANT` now fails loudly**: an unrecognised value now exits with an error instead of silently falling through to the cpu branch.
- **Embedding cache writes are now atomic**: `.npy` files are written to a `.tmp` sibling and renamed into place with `os.replace`, preventing truncated cache entries on process kill.
- **`EmbeddingCache` singleton re-creates when `DATA_DIR` changes**: prevents test runs from sharing cache state across different `DATA_DIR` values.
### Added
- **Diversity test suite** (PR #11): 33 tests covering k-medoids clustering, farthest-point sampling, adaptive threshold computation, near-duplicate deduplication, and time-spread selection. Total: 93 tests.
## [0.4.11] - 2026-06-14
### Removed
- **Object mode pipeline fully removed**: YOLO object detection, SigLIP image classification, `TRAINING_MODE`, and `OBJECT_CLASS` env vars are gone. Frigate has no training API for objects; the ~2 GB model stack (torch, torchvision, transformers, ultralytics) was dead weight.
- **Dead Immich embedding path removed**: `FaceData.embedding` field and the `immich_embedding` parameter to `get_embedding()` were never consumed by any caller. Both removed along with the NumPy import in `immich_api.py` that existed solely for that path.
- **Dead `mode` config key removed**: `"mode": "face"` was written into job config dicts in `jobs.py` but never read after object mode removal.
### Fixed
- **InsightFace `FutureWarning` suppressed in crop-alignment path**: the `insightface_app.get()` call in `image_processing.py` now wraps the same `warnings.catch_warnings()` suppressor already present in `embeddings.py`, preventing scikit-image deprecation noise in logs.
### Changed
- **Variant pyproject files synced to current state**: `pyproject-rocm.toml`, `pyproject-cpu.toml`, `pyproject-intel.toml` were at v0.2.13 and still listed torch/transformers/ultralytics. Updated to v0.4.11 and cleaned to face-only deps. Note: corresponding lockfiles (uv-rocm.lock, uv-cpu.lock, uv-intel.lock) need regeneration in their respective platform environments.
- **`MERGE_DUPLICATE_PEOPLE` documented**: README and wiki now explain the default warn-and-skip behaviour vs. setting `true` for a permanent Immich merge, with irreversibility callout.
- **Wiki fully updated**: all five wiki pages rewritten to remove object mode references, correct model size (~300 MB InsightFace vs former ~1–2 GB HuggingFace+InsightFace), fix default values (`MAX_AUTO_IMAGES` 80→20, `MIN_FACE_COUNT` 0→3), add `MERGE_DUPLICATE_PEOPLE` coverage, and update GPU verification commands for current ONNX provider API.
## [0.4.10] - 2026-06-14
### Changed
- **`MAX_AUTO_IMAGES` default lowered from `80` to `20`**: winnow is designed to fill the gap where manual Frigate training images don't exist — not to be the primary dataset. A conservative default ensures winnow-imported images remain secondary to hand-picked ones where both exist.
## [0.4.9] - 2026-06-14
### Changed
- **`FRIGATE_SCORE_CEILING` is now dynamic by default**: previously defaulted to `0` (disabled). Now unset (default) enables a self-calibrating novelty gate — below-cap candidates are skipped if their pre-upload Frigate score exceeds the most-redundant tracked file's score. This catches conditions already covered by manually-added Frigate images that winnow cannot track. Set `FRIGATE_SCORE_CEILING=0` to disable entirely; set a positive value (e.g. `0.85`) for a fixed hard ceiling.
- **Quality replacement branches consolidated**: the Frigate-score and blur-score replacement paths in the upload loop shared identical structure. Merged into a single code path parameterised by score source and comparison direction.
- **`MIN_FACE_COUNT` default raised from `0` to `3`**: people with fewer than 3 tagged photos produce degenerate training sets; skipping them by default avoids noisy runs.
- **`STRATEGY=adaptive`** is the new primary name for embedding-based diversity selection; `auto` remains a silent alias for backwards compatibility.
- **`MERGE_DUPLICATE_PEOPLE` and `TRACE_CROP_SIZE`** added to the README env var table (were in the codebase but undocumented).
- **CUDA version corrected** in the image tags table (was 13.3, actual base image is 12.8.1).
## [0.4.8] - 2026-06-14
### Changed
- **Tracker write-through cache**: `upload_tracker` now keeps an in-memory copy of each JSON file keyed by its resolved path. All reads after the first hit the cache instead of disk; writes go to both disk and cache atomically. Cuts per-person disk I/O in the upload loop from ~90 reads to ~1, with no API or behaviour changes.
## [0.4.7] - 2026-06-14
### Changed
- **`_dedup_embeddings` pre-allocated buffer**: replaced the grow-on-keep `np.vstack` pattern with a pre-allocated `(Q, D)` buffer filled row-by-row. Eliminates O(K²) copy work and the GC pressure from K intermediate heap allocations while keeping identical arithmetic for the similarity checks.
- **`_kmedoids` cost computation vectorized**: the Python-level `sum(dist_matrix[i, medoids[labels[i]]] for i in range(n))` generator (called once per swap evaluation) is replaced with `dist_matrix[np.arange(n), np.array(medoids)[labels]].sum()` — a single numpy fancy-index + reduction, ~20–50× faster in the swap loop.
- **`_reconcile_frigate_mappings` single-write batch**: previously called `record_frigate_file` once per uploaded file, each doing a full JSON load + save (O(L) disk round-trips per person). Now builds the full `{frigate_filename: asset_id}` mapping dict and writes it in one `record_frigate_files_batch` call (O(1) disk round-trip).
## [0.4.6] - 2026-06-14
### Fixed
- **OOM when Immich returns many pages per person**: `fetch_all_assets` now stops fetching once 5000 assets have been collected — the diversity selection pool is already capped at 3000 items, so fetching up to 1,000,000 was wasteful and could exhaust memory on large libraries. 5000 provides ample headroom for the pool cap while bounding per-person memory to ~2 MB.
- **Non-dict items in Immich asset pages silently skipped**: a malformed or partially-null Immich response page could include `null` or non-object items in the assets array. These are now filtered at fetch time rather than causing `AttributeError` downstream.
## [0.4.5] - 2026-06-14
### Fixed
- **Near-duplicate dedup O(N²) allocation**: `np.vstack(kept_normed)` was rebuilt on every loop iteration even for candidates that would be dropped; the stack is now rebuilt only when a new item is kept, reducing memory pressure significantly for large pools.
- **`quality_score` falsy-zero in dedup sort**: the sort key used `c.get("quality_score") or 0.0`, which treated a legitimate `quality_score=0.0` identically to a missing key. Changed to an explicit `None` check so zero is preserved as-is, and object-mode candidates (which have no `quality_score`) continue to sort stably to the back.
- **Post-dedup pool not re-checked against limit**: after near-duplicate removal the pool could silently shrink below the requested limit with no warning. A second `len < limit` guard now fires after dedup and emits the same "Only N embeddings" warning that the pre-dedup guard does.
- **`mark_rejected` could miss plain-text 400 bodies longer than 100 bytes**: `error_detail = resp.text[:100]` was being searched for the keyword `"face"` to gate `mark_rejected()`, so a response body with `"face"` after byte 100 would never mark the asset rejected and it would be retried on every future run. The `"face"` check now uses the full response body; truncation is kept only for the displayed snippet.
- **`_safe_person_dir` raised ValueError for all person names when `output_dir` resolved to `/`**: `base + os.sep` produced `"//"` when base was `"/"`, and valid paths like `/alice` don't start with `"//"`. Fixed by using `base` directly as the prefix when `base == os.sep`.
## [0.4.4] - 2026-06-14 ## [0.4.4] - 2026-06-14
### Added ### Added
+3
View File
@@ -17,8 +17,11 @@ git clone https://github.com/sudolulo/winnow.git
cd winnow cd winnow
git checkout dev git checkout dev
uv sync uv sync
git config core.hooksPath .githooks
``` ```
The last line activates the project's git hooks. The pre-commit hook automatically runs `uv lock` and stages the result whenever `pyproject.toml` is part of a commit, keeping the lockfile in sync without any extra steps.
## Running Tests and Lint ## Running Tests and Lint
```bash ```bash
+28 -22
View File
@@ -1,16 +1,17 @@
# ── Base images ─────────────────────────────────────────────────────────────── # ── Base images ───────────────────────────────────────────────────────────────
# amd64 + gpu: NVIDIA CUDA 12.8 + cuDNN (GPU acceleration via NVIDIA Container Toolkit) # amd64 + gpu: NVIDIA CUDA 12.8 + cuDNN on Ubuntu 24.04 (highest Ubuntu NVIDIA publishes)
# amd64 + rocm: Ubuntu 22.04 (AMD GPU via ROCm — pass /dev/kfd and /dev/dri) # amd64 + rocm: Ubuntu 26.04 (AMD GPU via ROCm — pass /dev/kfd and /dev/dri)
# amd64 + intel: Ubuntu 22.04 (Intel Arc / iGPU via OpenVINO — pass /dev/dri) # amd64 + intel: Ubuntu 22.04 (Intel Arc / iGPU via OpenVINO — pass /dev/dri)
# amd64 + cpu: Ubuntu 22.04 (CPU-only, ~2 GB smaller image) # Note: intel stays on 22.04 — Intel's GPU repo only publishes for jammy
# amd64 + cpu: Ubuntu 26.04 (CPU-only, ~2 GB smaller image)
# arm64: Ubuntu 24.04 (CPU-only; no CUDA/ROCm wheels on ARM) # arm64: Ubuntu 24.04 (CPU-only; no CUDA/ROCm wheels on ARM)
ARG VARIANT=gpu ARG VARIANT=gpu
FROM --platform=$BUILDPLATFORM nvidia/cuda:12.8.1-cudnn-runtime-ubuntu22.04 AS base-amd64-gpu FROM --platform=$BUILDPLATFORM nvidia/cuda:12.9.2-cudnn-runtime-ubuntu24.04 AS base-amd64-gpu
FROM ubuntu:22.04 AS base-amd64-rocm FROM ubuntu:26.04 AS base-amd64-rocm
FROM ubuntu:22.04 AS base-amd64-intel FROM ubuntu:22.04 AS base-amd64-intel
FROM ubuntu:22.04 AS base-amd64-cpu FROM ubuntu:26.04 AS base-amd64-cpu
FROM ubuntu:24.04 AS base-arm64-gpu FROM ubuntu:24.04 AS base-arm64-gpu
FROM ubuntu:24.04 AS base-arm64-rocm FROM ubuntu:24.04 AS base-arm64-rocm
FROM ubuntu:24.04 AS base-arm64-intel FROM ubuntu:24.04 AS base-arm64-intel
@@ -24,8 +25,9 @@ FROM base-${TARGETARCH}-${VARIANT} AS build
ARG VARIANT=gpu ARG VARIANT=gpu
ENV DEBIAN_FRONTEND=noninteractive ENV DEBIAN_FRONTEND=noninteractive
# Both Ubuntu 22.04 and 24.04 get Python 3.13 from the deadsnakes PPA. # All base images get Python 3.13 from the deadsnakes PPA (26.04 ships 3.14 natively;
# GNUPGHOME is isolated so gpg never contacts an agent socket under QEMU. # 3.13 is used to keep dependencies tested and aligned). GNUPGHOME is isolated
# so gpg never contacts an agent socket under QEMU.
RUN apt-get update && apt-get install -y --no-install-recommends \ RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates curl gnupg software-properties-common \ ca-certificates curl gnupg software-properties-common \
&& GNUPGHOME=$(mktemp -d) add-apt-repository ppa:deadsnakes/ppa -y \ && GNUPGHOME=$(mktemp -d) add-apt-repository ppa:deadsnakes/ppa -y \
@@ -36,23 +38,26 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
&& rm -rf /var/lib/apt/lists/* \ && rm -rf /var/lib/apt/lists/* \
&& ln -sf /usr/bin/python3.13 /usr/bin/python3 && ln -sf /usr/bin/python3.13 /usr/bin/python3
RUN curl -LsSf https://astral.sh/uv/install.sh | sh \ COPY --from=ghcr.io/astral-sh/uv:0.11.21 /uv /usr/local/bin/uv
&& cp /root/.local/bin/uv /usr/local/bin/uv
WORKDIR /app WORKDIR /app
# Swap in the variant-specific pyproject and lockfile before syncing. COPY pyproject.toml uv.lock ./
COPY pyproject.toml uv.lock pyproject-cpu.toml uv-cpu.lock \
pyproject-rocm.toml uv-rocm.lock pyproject-intel.toml uv-intel.lock ./
RUN if [ "$VARIANT" = "cpu" ]; then \ RUN if [ "$VARIANT" = "cpu" ]; then \
cp pyproject-cpu.toml pyproject.toml && cp uv-cpu.lock uv.lock; \ uv sync --frozen --no-dev --extra cpu; \
elif [ "$VARIANT" = "rocm" ]; then \ elif [ "$VARIANT" = "rocm" ]; then \
cp pyproject-rocm.toml pyproject.toml && cp uv-rocm.lock uv.lock; \ uv sync --frozen --no-dev --extra rocm; \
elif [ "$VARIANT" = "intel" ]; then \ elif [ "$VARIANT" = "intel" ]; then \
cp pyproject-intel.toml pyproject.toml && cp uv-intel.lock uv.lock; \ uv sync --frozen --no-dev --extra intel; \
elif [ "$VARIANT" = "gpu" ]; then \
uv sync --frozen --no-dev --extra gpu && \
ORT_GPU_VER=$(.venv/bin/python -c "import importlib.metadata; print(importlib.metadata.version('onnxruntime-gpu'))") && \
uv pip install --python .venv/bin/python --no-deps --reinstall "onnxruntime-gpu==$ORT_GPU_VER"; \
else \
echo "Unknown VARIANT: '$VARIANT'. Must be one of: cpu, rocm, intel, gpu" >&2; \
exit 1; \
fi && \ fi && \
uv sync --frozen --no-dev \ uv cache clean
&& uv cache clean
COPY winnow/ winnow/ COPY winnow/ winnow/
COPY entrypoint.sh scheduler.py ./ COPY entrypoint.sh scheduler.py ./
@@ -67,7 +72,7 @@ FROM base-${TARGETARCH}-${VARIANT} AS runtime
ARG VARIANT=gpu ARG VARIANT=gpu
ARG VERSION=dev ARG VERSION=dev
LABEL org.opencontainers.image.title="winnow" \ LABEL org.opencontainers.image.title="winnow" \
org.opencontainers.image.description="Selects diverse, high-quality photos from Immich as training data for Frigate face recognition and object classification." \ org.opencontainers.image.description="Selects diverse, high-quality photos from Immich as training data for Frigate face recognition." \
org.opencontainers.image.source="https://github.com/sudolulo/winnow" \ org.opencontainers.image.source="https://github.com/sudolulo/winnow" \
org.opencontainers.image.licenses="AGPL-3.0-or-later" \ org.opencontainers.image.licenses="AGPL-3.0-or-later" \
org.opencontainers.image.version="${VERSION}" org.opencontainers.image.version="${VERSION}"
@@ -116,7 +121,7 @@ https://repositories.intel.com/graphics/ubuntu jammy flex" \
fi fi
RUN groupadd -g 568 apps && useradd -u 568 -g apps -m -s /bin/bash appuser \ RUN groupadd -g 568 apps && useradd -u 568 -g apps -m -s /bin/bash appuser \
&& mkdir -p /models/.insightface /models/huggingface \ && mkdir -p /models/.insightface \
&& chown -R appuser:apps /app /models && chown -R appuser:apps /app /models
WORKDIR /app WORKDIR /app
@@ -124,7 +129,8 @@ USER appuser
# PYTHONPATH=/app makes the winnow package importable from the entry point script. # PYTHONPATH=/app makes the winnow package importable from the entry point script.
# uv sync builds the wheel before winnow/ is COPY'd, so site-packages has only # uv sync builds the wheel before winnow/ is COPY'd, so site-packages has only
# the dist-info. Explicitly adding /app lets Python find winnow/__init__.py there. # the dist-info. Explicitly adding /app lets Python find winnow/__init__.py there.
ENV HF_HOME=/models/huggingface INSIGHTFACE_HOME=/models/.insightface PYTHONPATH=/app ENV INSIGHTFACE_HOME=/models/.insightface PYTHONPATH=/app
HEALTHCHECK CMD test -f /app/entrypoint.sh || exit 1 HEALTHCHECK --interval=60s --timeout=5s --start-period=120s --retries=3 \
CMD sh -c 'if [ -f /tmp/winnow.pid ]; then kill -0 "$(cat /tmp/winnow.pid)"; fi'
ENTRYPOINT ["tini", "--", "/app/entrypoint.sh"] ENTRYPOINT ["tini", "--", "/app/entrypoint.sh"]
+59 -55
View File
@@ -2,16 +2,18 @@
[![Docker](https://github.com/sudolulo/winnow/actions/workflows/docker-publish.yml/badge.svg)](https://github.com/sudolulo/winnow/actions/workflows/docker-publish.yml) [![Test](https://github.com/sudolulo/winnow/actions/workflows/test.yml/badge.svg)](https://github.com/sudolulo/winnow/actions/workflows/test.yml) [![GitHub release](https://img.shields.io/github/v/release/sudolulo/winnow)](https://github.com/sudolulo/winnow/releases/latest) [![License: AGPL v3](https://img.shields.io/badge/License-AGPL_v3-blue.svg)](LICENSE) [![Immich](https://img.shields.io/badge/Immich-v1.106%2B-blueviolet)](https://immich.app) [![Frigate](https://img.shields.io/badge/Frigate-Ready-brightgreen)](https://frigate.video) [![Docker](https://github.com/sudolulo/winnow/actions/workflows/docker-publish.yml/badge.svg)](https://github.com/sudolulo/winnow/actions/workflows/docker-publish.yml) [![Test](https://github.com/sudolulo/winnow/actions/workflows/test.yml/badge.svg)](https://github.com/sudolulo/winnow/actions/workflows/test.yml) [![GitHub release](https://img.shields.io/github/v/release/sudolulo/winnow)](https://github.com/sudolulo/winnow/releases/latest) [![License: AGPL v3](https://img.shields.io/badge/License-AGPL_v3-blue.svg)](LICENSE) [![Immich](https://img.shields.io/badge/Immich-v1.106%2B-blueviolet)](https://immich.app) [![Frigate](https://img.shields.io/badge/Frigate-Ready-brightgreen)](https://frigate.video)
> **Note:** winnow's approach to training Frigate face recognition is not an officially documented workflow — results may vary.
> **Early Development — Use With Caution** > **Early Development — Use With Caution**
> winnow is functional but still maturing. Features that modify your Frigate training data — quality replacement, stale mapping cleanup — can remove images from your dataset and are not yet battle-tested at scale. Review the logs after each run and keep backups of your Frigate face training directory until you are confident in the results. > winnow is in an unfinished state and maturing. Features that modify your Frigate training data — quality replacement, stale mapping cleanup — can remove images from your dataset and are not yet battle-tested at scale. Review the logs after each run and keep backups of your Frigate face training directory until you are confident in the results.
**Docs:** [Setup](https://github.com/sudolulo/winnow/wiki/Setup) · [Troubleshooting](https://github.com/sudolulo/winnow/wiki/Troubleshooting) · [FAQ](https://github.com/sudolulo/winnow/wiki/FAQ) **Docs:** [Setup](https://github.com/sudolulo/winnow/wiki/Setup) · [Troubleshooting](https://github.com/sudolulo/winnow/wiki/Troubleshooting) · [FAQ](https://github.com/sudolulo/winnow/wiki/FAQ)
`winnow` pulls photos from your [Immich](https://immich.app) library, selects the most diverse and highest-quality subset using AI embeddings, and delivers them as training data for [Frigate](https://frigate.video)'s face recognition and object classification models. `winnow` pulls photos from your [Immich](https://immich.app) library, selects the most diverse and highest-quality subset using AI embeddings, and delivers them as training data for [Frigate](https://frigate.video)'s face recognition.
Frigate's face recognition is only as good as its training data — and the key quality metric is **diversity**, not volume. A hundred photos from the same week teach the model one lighting condition. What you need is a spread: different years, different angles, different lighting, different contexts. Your photo library already has that data. winnow finds and delivers the right subset automatically. The best Frigate training data is images you curate manually — photos taken specifically for recognition, in controlled conditions, uploaded directly through Frigate's UI. Winnow is meant to supplement people in your library, not replace manual training. In some cases one has people they would like to recognize that do not occur in detections often enough to train a diverse dataset. This is meant to fill that gap.
> **winnow only touches files it uploaded.** Faces added to Frigate manually through its UI are never deleted, replaced, or modified — not by quality replacement, not by `RESET_PERSON`, not by stale cleanup. If you have a curated training set you want to keep, it is safe. > **winnow only touches files it uploaded.** Faces added to Frigate manually through its UI are never deleted, replaced, or modified — not by quality replacement, not by `RESET_PERSON`, not by stale cleanup. Your manually curated images are always the primary dataset; winnow only adds to it.
--- ---
@@ -25,7 +27,7 @@ Immich library
│ │
▼ ▼
2. Filter by recency (YEARS_FILTER) and skip already-uploaded 2. Filter by recency (YEARS_FILTER) and skip already-uploaded
and rejected assets (persistent tracker in CACHE_DIR) and rejected assets (persistent tracker in DATA_DIR)
│ │
▼ ▼
3. Quality filter — download preview thumbnails and reject: 3. Quality filter — download preview thumbnails and reject:
@@ -37,48 +39,42 @@ Immich library
│ │
▼ ▼
4. Compute embeddings from the same preview thumbnails 4. Compute embeddings from the same preview thumbnails
• Faces → InsightFace (ArcFace / Buffalo_L) → 512-dim vector • InsightFace (ArcFace / Buffalo_L) → 512-dim vector
• Objects → SigLIP (Vision Transformer) → 768-dim vector
│ │
▼ ▼
5. Diversity selection 5. Near-duplicate removal — greedy cosine-distance pass drops burst shots
and near-identical photos before clustering runs; the highest-quality
image from each near-duplicate group is kept
│
▼
6. Diversity selection
• K-Medoids clustering → one representative per natural group • K-Medoids clustering → one representative per natural group
• Farthest Point Sampling → fill remaining slots with maximally spread picks • Farthest Point Sampling → fill remaining slots with maximally spread picks
• Hard example weighting — unusual angles and low-confidence detections • Hard example weighting — low-confidence detections get a distance boost
are biased toward selection, since those are where models tend to fail so unusual angles and harder looks are preferred over easy frontals
• Auto mode: stops when similarity to the existing set exceeds a threshold • Adaptive mode: stops when the next candidate is too similar to those already
(20 % of median pairwise distance for faces, 10 % for objects) selected (distance threshold = 20 % of median pairwise distance)
│ │
▼ ▼
6. Download full-resolution originals from Immich 7. Download full-resolution originals from Immich
│ │
▼ ▼
7. Crop and process 8. Crop and process — EXIF-corrected, landmark-aligned 112×112 crop (ArcFace format)
• Face mode: EXIF-corrected, landmark-aligned 112×112 crop (ArcFace format)
• Object mode: YOLOv9c detection → one crop per matched instance
│ │
▼ ▼
8. Deliver 9. Deliver — upload crops to Frigate's face registration API
• Face mode: upload crops to Frigate's face registration API ↳ below MAX_AUTO_IMAGES — upload, unless the novelty gate
↳ below MAX_AUTO_IMAGES — upload freely (FRIGATE_SCORE_CEILING) determines the candidate is already
↳ at cap + QUALITY_REPLACEMENT=true — with Frigate scoring active, covered by the current training set
swap the most redundant tracked image (highest pre-upload recognize ↳ at cap + QUALITY_REPLACEMENT=true — with Frigate scoring active,
score) if the candidate is more novel (lower score); falling back to swap the most redundant tracked image (highest pre-upload recognize
blur-score comparison when no Frigate scores are available; manually score) if the candidate is more novel (lower score); falling back to
added files are never touched blur-score comparison when no Frigate scores are available; manually
↳ at cap + QUALITY_REPLACEMENT=false — skip this person added files are never touched
• Object mode: save crops to disk → place into your Frigate data directory ↳ at cap + QUALITY_REPLACEMENT=false — skip this person
``` ```
Uploaded and rejected asset IDs are persisted across runs. The same image is never processed twice; Frigate rejections are permanently skipped unless `RETRY_REJECTED=true`. Uploaded and rejected asset IDs are persisted across runs in two JSON files (`frigate_uploaded_ids.json` and `frigate_rejected_ids.json` in `DATA_DIR`). The same image is never processed twice; rejected assets are permanently skipped unless `RETRY_REJECTED=true`.
---
## Modes
**Face mode** (default) — extracts face crops using Immich's bounding box metadata, applies EXIF orientation correction, and aligns them to ArcFace's standard 112×112 format using 5-point facial landmarks. Crops are uploaded directly to Frigate's face registration API.
**Object mode** — runs each full-resolution image through YOLOv9c to detect instances of a target class (dog, cat, car, etc.), crops each detection, and saves it to the output directory. Frigate has no API for uploading object training data; place the crops into your Frigate data directory manually.
--- ---
@@ -88,7 +84,7 @@ Uploaded and rejected asset IDs are persisted across runs. The same image is nev
| Tag | Arch | Acceleration | | Tag | Arch | Acceleration |
| :-- | :-- | :-- | | :-- | :-- | :-- |
| `:latest` | amd64 + arm64 | NVIDIA CUDA 13.3 (amd64) · requires [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html) | | `:latest` | amd64 | NVIDIA CUDA 12.8 · requires [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html) |
| `:rocm` | amd64 | AMD ROCm · pass `/dev/kfd` + `/dev/dri` | | `:rocm` | amd64 | AMD ROCm · pass `/dev/kfd` + `/dev/dri` |
| `:intel` | amd64 | Intel Arc / iGPU via OpenVINO · pass `/dev/dri`, set `OPENVINO_DEVICE=GPU` | | `:intel` | amd64 | Intel Arc / iGPU via OpenVINO · pass `/dev/dri`, set `OPENVINO_DEVICE=GPU` |
| `:cpu` | amd64 + arm64 | CPU only · ~2 GB smaller · no GPU required | | `:cpu` | amd64 + arm64 | CPU only · ~2 GB smaller · no GPU required |
@@ -106,8 +102,8 @@ services:
- FRIGATE_URL=http://192.168.1.10:5000 - FRIGATE_URL=http://192.168.1.10:5000
- CRON_SCHEDULE=0 3 * * 0 - CRON_SCHEDULE=0 3 * * 0
volumes: volumes:
- /path/to/models:/models - /path/to/models:/models # INSIGHTFACE_HOME — persists Buffalo_L model (~300 MB)
- /path/to/cache:/app/.if_cache - /path/to/data:/app/data
- /path/to/output:/app/frigate_train - /path/to/output:/app/frigate_train
deploy: deploy:
resources: resources:
@@ -152,7 +148,7 @@ See [compose.yml](compose.yml) for the full annotated example with all options.
| *(empty string)* | Stay alive, run nothing — trigger manually with `docker exec -it winnow winnow` | | *(empty string)* | Stay alive, run nothing — trigger manually with `docker exec -it winnow winnow` |
| Cron expression | Run on startup, then repeat on schedule | | Cron expression | Run on startup, then repeat on schedule |
In scheduled mode the process (and loaded models) stays resident between runs. The first run after a fresh install downloads the embedding models (~1–2 GB); subsequent runs use the cached models from the mounted volume. In scheduled mode the process (and loaded models) stays resident between runs. The first run after a fresh install downloads InsightFace Buffalo_L (~300 MB); subsequent runs use the cached model from the mounted volume.
--- ---
@@ -170,11 +166,9 @@ In scheduled mode the process (and loaded models) stays resident between runs. T
| Variable | Default | Description | | Variable | Default | Description |
| :--- | :--- | :--- | | :--- | :--- | :--- |
| `TRAINING_MODE` | `face` | `face` — upload crops to Frigate; `object` — save crops to disk | | `STRATEGY` | `adaptive` | `adaptive` — embedding-based diversity selection, stops when candidates become redundant; `standard` — fixed 30 images; `broad` — fixed 100 images |
| `STRATEGY` | `auto` | `auto` (embedding-based adaptive), `standard` (30 images), `broad` (100 images) |
| `LIMIT` | *(unset)* | Exact image count — overrides `STRATEGY` | | `LIMIT` | *(unset)* | Exact image count — overrides `STRATEGY` |
| `OBJECT_CLASS` | `dog` | Target class for object mode (any YOLO class: `dog`, `cat`, `car`, etc.) | | `AUTO_MODE` | *(auto)* | Skip interactive prompts and process all people unattended — auto-detected when no TTY is present (Docker, cron); set `true` to force in a terminal |
| `AUTO_MODE` | *(auto)* | Force non-interactive mode in a terminal; auto-detected otherwise |
| `VERBOSE` | `false` | Enable DEBUG-level console output (log file is always DEBUG) | | `VERBOSE` | `false` | Enable DEBUG-level console output (log file is always DEBUG) |
### People Filtering ### People Filtering
@@ -183,23 +177,33 @@ In scheduled mode the process (and loaded models) stays resident between runs. T
| :--- | :--- | :--- | | :--- | :--- | :--- |
| `ONLY_PEOPLE` | *(unset)* | Comma-separated whitelist — process only these people | | `ONLY_PEOPLE` | *(unset)* | Comma-separated whitelist — process only these people |
| `SKIP_PEOPLE` | *(unset)* | Comma-separated list — skip these people | | `SKIP_PEOPLE` | *(unset)* | Comma-separated list — skip these people |
| `MIN_FACE_COUNT` | `0` | Skip people with fewer than N tagged assets in Immich | | `MIN_FACE_COUNT` | `3` | Skip people with fewer than N tagged assets in Immich |
| `MERGE_DUPLICATE_PEOPLE` | `false` | When Immich has duplicate entries for the same person (same face split across multiple names), merge their asset pools before processing. Without this, each duplicate group emits a warning and is skipped |
| `YEARS_FILTER` | `10` | Ignore images older than N years | | `YEARS_FILTER` | `10` | Ignore images older than N years |
> **Duplicate people detection** — winnow warns at startup if the same name appears on multiple Immich person records (a common side-effect of Immich's face clustering creating separate pools for the same individual). By default (`false`) it logs the duplicates, keeps only the person with the most assets, and skips the rest — no data is changed. Set `MERGE_DUPLICATE_PEOPLE=true` to permanently merge each duplicate group inside Immich (the person with the most assets absorbs the others). **This modifies Immich and cannot be undone.** Only enable it once you've verified the duplicates are actually the same person.
### Image Quality ### Image Quality
| Variable | Default | Description | | Variable | Default | Description |
| :--- | :--- | :--- | | :--- | :--- | :--- |
| `MAX_AUTO_IMAGES` | `5` | Maximum training images per person in Frigate |
| `QUALITY_REPLACEMENT` | `true` | When at cap, swap a weaker tracked image for a better candidate. With Frigate scoring active, targets the most redundant image (highest pre-upload recognize score); otherwise uses blur score. Never touches manually added Frigate files. Set `false` to skip people at cap |
#### Advanced Tuning *(calibrated — do not adjust)*
These defaults are tuned for Frigate's ArcFace requirements. winnow will warn on launch if any are set. Image quality issues caused by non-default values will not be investigated.
| Variable | Default | Description |
| :--- | :--- | :--- |
| `ENABLE_FRIGATE_SCORES` | `true` | Call Frigate's recognize endpoint pre-upload to store diversity scores used for quality replacement. Adds ~200 ms per upload. Disabling also disables the below-cap novelty gate |
| `FRIGATE_SCORE_CEILING` | *(unset)* | Below-cap novelty gate. Unset: dynamic — skips candidates whose Frigate score exceeds the most-redundant tracked file's score, auto-calibrates each run. `0`: disable entirely. Positive value (e.g. `0.85`): fixed hard ceiling |
| `MIN_FACE_WIDTH` | `90` | Minimum face crop width in pixels | | `MIN_FACE_WIDTH` | `90` | Minimum face crop width in pixels |
| `FACE_MARGIN` | `0.15` | Padding around bounding box crop (fraction of face size) | | `FACE_MARGIN` | `0.15` | Padding around bounding box crop (fraction of face size) |
| `ENABLE_FACE_ALIGNMENT` | `true` | Align to ArcFace 112×112 format using facial landmarks | | `ENABLE_FACE_ALIGNMENT` | `true` | Align to ArcFace 112×112 format using facial landmarks |
| `USE_FULL_RESOLUTION` | `true` | Download full-resolution originals rather than preview thumbnails | | `USE_FULL_RESOLUTION` | `true` | Download full-resolution originals rather than preview thumbnails |
| `MIN_CONFIDENCE` | `0.7` | Minimum Immich face detection confidence | | `MIN_CONFIDENCE` | `0.7` | Minimum Immich face detection confidence |
| `BLUR_THRESHOLD` | `120.0` | Laplacian variance threshold — lower accepts more blur | | `BLUR_THRESHOLD` | `120.0` | Laplacian variance threshold — lower accepts more blur |
| `MAX_AUTO_IMAGES` | `80` | Maximum training images per person in Frigate |
| `QUALITY_REPLACEMENT` | `true` | When at cap, swap a weaker tracked image for a better candidate. With Frigate scoring active, targets the most redundant image (highest pre-upload recognize score); otherwise uses blur score. Never touches manually added Frigate files. Set `false` to skip people at cap |
| `FRIGATE_SCORE_CEILING` | `0.0` | Skip uploads whose pre-upload Frigate recognize score exceeds this value — they are already well-covered. `0` disables; requires at least one prior run to have scores |
| `ENABLE_FRIGATE_SCORES` | `true` | Call Frigate's recognize endpoint pre-upload to store diversity scores used for quality replacement. Adds ~200 ms per upload. Disable to use blur-score replacement only |
### GPU & Models ### GPU & Models
@@ -208,23 +212,23 @@ In scheduled mode the process (and loaded models) stays resident between runs. T
| `FORCE_CPU` | `false` | Disable GPU — fall back to CPU for all inference | | `FORCE_CPU` | `false` | Disable GPU — fall back to CPU for all inference |
| `OPENVINO_DEVICE` | `CPU` | Intel variant only: set `GPU` to use Arc or iGPU; default runs on CPU | | `OPENVINO_DEVICE` | `CPU` | Intel variant only: set `GPU` to use Arc or iGPU; default runs on CPU |
| `ENABLE_CACHE` | `true` | Cache computed embeddings to disk (speeds up re-runs on the same library) | | `ENABLE_CACHE` | `true` | Cache computed embeddings to disk (speeds up re-runs on the same library) |
| `CACHE_DIR` | `.if_cache` | Path for embedding cache and upload tracker files | | `DATA_DIR` | `data` | Path for embedding cache and upload tracker JSON files |
| `HF_HOME` | *(system)* | HuggingFace model cache path (SigLIP) |
| `INSIGHTFACE_HOME` | *(system)* | InsightFace model cache path (Buffalo_L) | | `INSIGHTFACE_HOME` | *(system)* | InsightFace model cache path (Buffalo_L) |
### Output ### Output
| Variable | Default | Description | | Variable | Default | Description |
| :--- | :--- | :--- | | :--- | :--- | :--- |
| `OUTPUT_DIR` | `./frigate_train` | Directory for object-mode crops and the `winnow.log` file. In Docker, set this via the volume mount instead. | | `OUTPUT_DIR` | `./frigate_train` | Directory where face crops are staged before upload and where `winnow.log` is written. In Docker, set this via the volume mount instead. |
### Tracker Overrides *(one-shot — remove after use)* ### Tracker Overrides *(one-shot — remove after use)*
| Variable | Default | Description | | Variable | Default | Description |
| :--- | :--- | :--- | | :--- | :--- | :--- |
| `DRY_RUN` | `false` | Preview selection without downloading or uploading | | `DRY_RUN` | `false` | Preview selection without downloading or uploading |
| `RETRY_REJECTED` | `false` | Re-attempt assets previously rejected by Frigate | | `RETRY_REJECTED` | `false` | Re-attempt all previously rejected assets (low-confidence skips, Frigate rejections, and other permanent exclusions) |
| `RESET_PERSON` | *(unset)* | Clear upload history for one person and delete their winnow-managed Frigate training files so the next run starts fresh. Manually added Frigate files are never touched | | `RESET_PERSON` | *(unset)* | Set to a person's name to clear their upload history and delete their winnow-managed Frigate training files so the next run starts fresh. Set to `*` to reset all tracked people at once. Manually added Frigate files are never touched |
| `TRACE_CROP_SIZE` | *(unset)* | Debug: print all tracked crops whose width or height matches this pixel value, then exit |
### Scheduling ### Scheduling
@@ -245,14 +249,14 @@ uv run winnow
Requires Python 3.13+ and [uv](https://astral.sh/uv). An NVIDIA, AMD, or Intel GPU is recommended — CPU mode works but embedding computation is slower. Requires Python 3.13+ and [uv](https://astral.sh/uv). An NVIDIA, AMD, or Intel GPU is recommended — CPU mode works but embedding computation is slower.
When run with a terminal attached, winnow starts an interactive session: select which people to process and choose a strategy (auto, standard, broad, or a custom count) per person. Without a TTY — Docker, cron, or `AUTO_MODE=true` — it processes all people automatically using the configured defaults. When run with a terminal attached, winnow starts an interactive session: select which people to process and choose a strategy (adaptive, standard, broad, or a custom count) per person. Without a TTY — Docker, cron, or `AUTO_MODE=true` — it processes all people unattended using the configured defaults.
--- ---
## Requirements ## Requirements
- **Immich** v1.106+ - **Immich** v1.106+
- **Frigate** v0.16+ (face mode only — object mode has no Frigate dependency) - **Frigate** v0.16+
- **GPU** recommended: NVIDIA (CUDA), AMD (ROCm), or Intel (Arc / iGPU via OpenVINO) - **GPU** recommended: NVIDIA (CUDA), AMD (ROCm), or Intel (Arc / iGPU via OpenVINO)
- **Python** 3.13+ - **Python** 3.13+
+3 -8
View File
@@ -13,13 +13,9 @@ services:
# Set AUTO_MODE=true to force auto mode in an interactive terminal. # Set AUTO_MODE=true to force auto mode in an interactive terminal.
# To run interactively: docker exec -it winnow winnow # To run interactively: docker exec -it winnow winnow
# - VERBOSE=true # Enable DEBUG-level console output # - VERBOSE=true # Enable DEBUG-level console output
# TRAINING_MODE: face = upload to Frigate face recognition API
# object = save crops to output dir for manual Frigate placement
- TRAINING_MODE=face
# STRATEGY: auto = objective diversity (recommended), standard = 30 imgs, broad = 100 imgs # STRATEGY: auto = objective diversity (recommended), standard = 30 imgs, broad = 100 imgs
- STRATEGY=auto - STRATEGY=auto
# - LIMIT=50 # Custom image count; overrides STRATEGY preset # - LIMIT=50 # Custom image count; overrides STRATEGY preset
# - OBJECT_CLASS=dog # Object label for object mode (e.g. dog, cat, car)
# ── People Filtering ────────────────────────────────────────────────── # ── People Filtering ──────────────────────────────────────────────────
# - ONLY_PEOPLE=John,Jane # Comma-separated; process only these people # - ONLY_PEOPLE=John,Jane # Comma-separated; process only these people
@@ -35,14 +31,13 @@ services:
# - USE_FULL_RESOLUTION=true # Use full-res images vs thumbnails (default: true) # - USE_FULL_RESOLUTION=true # Use full-res images vs thumbnails (default: true)
# - MIN_CONFIDENCE=0.7 # Minimum face detection confidence (default: 0.7) # - MIN_CONFIDENCE=0.7 # Minimum face detection confidence (default: 0.7)
# - BLUR_THRESHOLD=100.0 # Laplacian blur threshold; lower = accept more blur (default: 100.0) # - BLUR_THRESHOLD=100.0 # Laplacian blur threshold; lower = accept more blur (default: 100.0)
# - MAX_AUTO_IMAGES=80 # Hard cap on auto-diversity selection (default: 80) # - MAX_AUTO_IMAGES=80 # Hard cap on auto-diversity selection (default: 5)
# ── Caching & Models ────────────────────────────────────────────────── # ── Caching & Models ──────────────────────────────────────────────────
# - FORCE_CPU=true # Disable GPU, fall back to CPU # - FORCE_CPU=true # Disable GPU, fall back to CPU
# - OPENVINO_DEVICE=GPU # Intel variant only: use Arc/iGPU instead of CPU (default: CPU) # - OPENVINO_DEVICE=GPU # Intel variant only: use Arc/iGPU instead of CPU (default: CPU)
# - ENABLE_CACHE=false # Disable embedding cache (default: true) # - ENABLE_CACHE=false # Disable embedding cache (default: true)
- CACHE_DIR=/app/.if_cache - DATA_DIR=/app/data
- HF_HOME=/models/huggingface
- INSIGHTFACE_HOME=/models/.insightface - INSIGHTFACE_HOME=/models/.insightface
# ── Tracker overrides (one-shot, remove after use) ──────────────────── # ── Tracker overrides (one-shot, remove after use) ────────────────────
@@ -62,7 +57,7 @@ services:
volumes: volumes:
# Replace with absolute paths on your host, e.g. /opt/winnow/models # Replace with absolute paths on your host, e.g. /opt/winnow/models
- /path/to/winnow/models:/models - /path/to/winnow/models:/models
- /path/to/winnow/cache:/app/.if_cache - /path/to/winnow/data:/app/data
- /path/to/winnow/output:/app/frigate_train - /path/to/winnow/output:/app/frigate_train
restart: unless-stopped restart: unless-stopped
-93
View File
@@ -1,93 +0,0 @@
[project]
name = "winnow"
version = "0.2.13"
description = "Immich to Frigate training sets"
license = "AGPL-3.0-or-later"
requires-python = ">=3.13"
authors = [{ name = "Holden Salomon", email = "holden@arch.fyi" }]
keywords = ["immich", "frigate", "face-recognition", "training-data", "arcface", "insightface"]
classifiers = [
"Development Status :: 3 - Alpha",
"Intended Audience :: Developers",
"License :: OSI Approved :: GNU Affero General Public License v3 or later (AGPLv3+)",
"Programming Language :: Python :: 3.13",
"Topic :: Scientific/Engineering :: Image Recognition",
]
dependencies = [
"croniter>=5.0.2",
"insightface>=0.7.3",
"numpy>=2.2.6",
"onnxruntime>=1.23.2",
"opencv-python-headless>=4.12.0.88",
"pillow>=12.1.0",
"python-dotenv>=1.2.1",
"requests>=2.32.5",
"rich>=14.2.0",
"torch>=2.12.0",
"torchvision>=0.27.0",
"transformers>=5.12.0",
"ultralytics>=8.4.66",
]
[project.scripts]
winnow = "winnow.cli:main"
[project.urls]
Repository = "https://github.com/sudolulo/winnow"
[tool.uv]
required-environments = [
"sys_platform == 'linux' and platform_machine == 'x86_64'",
]
[tool.uv.sources]
torch = [
{ index = "pytorch-cpu", marker = "sys_platform == 'linux' and platform_machine == 'x86_64'" },
]
torchvision = [
{ index = "pytorch-cpu", marker = "sys_platform == 'linux' and platform_machine == 'x86_64'" },
]
[[tool.uv.index]]
name = "pytorch-cpu"
url = "https://download.pytorch.org/whl/cpu"
explicit = true
[dependency-groups]
dev = [
"pytest>=8.0",
"ruff>=0.15.17",
]
[tool.hatch.build.targets.wheel]
packages = ["winnow"]
[tool.ruff]
line-length = 120
target-version = "py313"
[tool.ruff.lint]
select = ["E", "F", "I"]
[tool.deptry]
pep621_dev_dependency_groups = ["dev"]
[tool.deptry.package_module_name_map]
pillow = "PIL"
opencv-python-headless = "cv2"
python-dotenv = "dotenv"
insightface = "insightface"
numpy = "numpy"
onnxruntime = "onnxruntime"
requests = "requests"
rich = "rich"
torch = "torch"
transformers = "transformers"
ultralytics = "ultralytics"
[tool.pytest.ini_options]
testpaths = ["tests"]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
-103
View File
@@ -1,103 +0,0 @@
[project]
name = "winnow"
version = "0.2.13"
description = "Immich to Frigate training sets"
license = "AGPL-3.0-or-later"
requires-python = ">=3.13"
authors = [{ name = "Holden Salomon", email = "holden@arch.fyi" }]
keywords = ["immich", "frigate", "face-recognition", "training-data", "arcface", "insightface"]
classifiers = [
"Development Status :: 3 - Alpha",
"Intended Audience :: Developers",
"License :: OSI Approved :: GNU Affero General Public License v3 or later (AGPLv3+)",
"Programming Language :: Python :: 3.13",
"Topic :: Scientific/Engineering :: Image Recognition",
]
dependencies = [
"croniter>=5.0.2",
"insightface>=0.7.3",
"numpy>=2.2.6",
"onnxruntime-openvino>=1.20.0",
"opencv-python-headless>=4.12.0.88",
"pillow>=12.1.0",
"python-dotenv>=1.2.1",
"requests>=2.32.5",
"rich>=14.2.0",
"torch>=2.12.0",
"torchvision>=0.27.0",
"transformers>=5.12.0",
"ultralytics>=8.4.66",
]
[project.scripts]
winnow = "winnow.cli:main"
[project.urls]
Repository = "https://github.com/sudolulo/winnow"
[tool.uv]
conflicts = [
[
{ package = "onnxruntime" },
{ package = "onnxruntime-gpu" },
{ package = "onnxruntime-openvino" },
],
]
required-environments = [
"sys_platform == 'linux' and platform_machine == 'x86_64'",
]
[tool.uv.sources]
torch = [
{ index = "pytorch-cpu", marker = "sys_platform == 'linux' and platform_machine == 'x86_64'" },
]
torchvision = [
{ index = "pytorch-cpu", marker = "sys_platform == 'linux' and platform_machine == 'x86_64'" },
]
[[tool.uv.index]]
name = "pytorch-cpu"
url = "https://download.pytorch.org/whl/cpu"
explicit = true
[dependency-groups]
dev = [
"pytest>=8.0",
"ruff>=0.15.17",
]
[tool.hatch.build.targets.wheel]
packages = ["winnow"]
[tool.ruff]
line-length = 120
target-version = "py313"
[tool.ruff.lint]
select = ["E", "F", "I"]
[tool.deptry]
pep621_dev_dependency_groups = ["dev"]
[tool.deptry.package_module_name_map]
pillow = "PIL"
opencv-python-headless = "cv2"
python-dotenv = "dotenv"
insightface = "insightface"
numpy = "numpy"
onnxruntime-openvino = "onnxruntime"
requests = "requests"
rich = "rich"
torch = "torch"
transformers = "transformers"
ultralytics = "ultralytics"
[tool.deptry.per_rule_ignores]
DEP002 = ["onnxruntime-openvino"]
[tool.pytest.ini_options]
testpaths = ["tests"]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
-103
View File
@@ -1,103 +0,0 @@
[project]
name = "winnow"
version = "0.2.13"
description = "Immich to Frigate training sets"
license = "AGPL-3.0-or-later"
requires-python = ">=3.13"
authors = [{ name = "Holden Salomon", email = "holden@arch.fyi" }]
keywords = ["immich", "frigate", "face-recognition", "training-data", "arcface", "insightface"]
classifiers = [
"Development Status :: 3 - Alpha",
"Intended Audience :: Developers",
"License :: OSI Approved :: GNU Affero General Public License v3 or later (AGPLv3+)",
"Programming Language :: Python :: 3.13",
"Topic :: Scientific/Engineering :: Image Recognition",
]
dependencies = [
"croniter>=5.0.2",
"insightface>=0.7.3",
"numpy>=2.2.6",
"onnxruntime-rocm>=1.16.0",
"opencv-python-headless>=4.12.0.88",
"pillow>=12.1.0",
"python-dotenv>=1.2.1",
"requests>=2.32.5",
"rich>=14.2.0",
"torch>=2.5.0",
"torchvision>=0.20.0",
"transformers>=5.12.0",
"ultralytics>=8.4.66",
]
[project.scripts]
winnow = "winnow.cli:main"
[project.urls]
Repository = "https://github.com/sudolulo/winnow"
[tool.uv]
index-strategy = "unsafe-best-match"
conflicts = [
[
{ package = "onnxruntime" },
{ package = "onnxruntime-gpu" },
{ package = "onnxruntime-rocm" },
],
]
required-environments = [
"sys_platform == 'linux' and platform_machine == 'x86_64'",
]
[tool.uv.sources]
torch = [
{ index = "pytorch-rocm63", marker = "sys_platform == 'linux' and platform_machine == 'x86_64'" },
]
torchvision = [
{ index = "pytorch-rocm63", marker = "sys_platform == 'linux' and platform_machine == 'x86_64'" },
]
[[tool.uv.index]]
name = "pytorch-rocm63"
url = "https://download.pytorch.org/whl/rocm6.3"
[dependency-groups]
dev = [
"pytest>=8.0",
"ruff>=0.15.17",
]
[tool.hatch.build.targets.wheel]
packages = ["winnow"]
[tool.ruff]
line-length = 120
target-version = "py313"
[tool.ruff.lint]
select = ["E", "F", "I"]
[tool.deptry]
pep621_dev_dependency_groups = ["dev"]
[tool.deptry.package_module_name_map]
pillow = "PIL"
opencv-python-headless = "cv2"
python-dotenv = "dotenv"
insightface = "insightface"
numpy = "numpy"
onnxruntime-rocm = "onnxruntime"
requests = "requests"
rich = "rich"
torch = "torch"
transformers = "transformers"
ultralytics = "ultralytics"
[tool.deptry.per_rule_ignores]
DEP002 = ["onnxruntime-rocm"]
[tool.pytest.ini_options]
testpaths = ["tests"]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
+29 -44
View File
@@ -1,7 +1,7 @@
[project] [project]
name = "winnow" name = "winnow"
version = "0.4.4" version = "0.6.6"
description = "Selects diverse, high-quality photos from Immich as training data for Frigate face recognition and object classification." description = "Selects diverse, high-quality photos from Immich as training data for Frigate face recognition."
license = "AGPL-3.0-or-later" license = "AGPL-3.0-or-later"
requires-python = ">=3.13" requires-python = ">=3.13"
authors = [{ name = "Holden Salomon", email = "holden@arch.fyi" }] authors = [{ name = "Holden Salomon", email = "holden@arch.fyi" }]
@@ -15,24 +15,28 @@ classifiers = [
"Topic :: Scientific/Engineering :: Image Recognition", "Topic :: Scientific/Engineering :: Image Recognition",
] ]
dependencies = [ dependencies = [
"croniter>=5.0.2", "croniter>=6.2.3",
"insightface>=0.7.3", "insightface>=0.7.3",
"nvidia-cudnn-cu12>=9.0.0", "numpy>=2.5.1",
"numpy>=2.2.6", "opencv-python-headless>=5.0.0.93",
"onnxruntime-gpu>=1.23.2; sys_platform == 'linux' and platform_machine == 'x86_64'", "pillow>=12.3.0",
"onnxruntime>=1.23.2; sys_platform == 'linux' and platform_machine != 'x86_64'",
"onnxruntime>=1.23.2; sys_platform != 'linux'",
"opencv-python-headless>=4.12.0.88",
"pillow>=12.1.0",
"python-dotenv>=1.2.1", "python-dotenv>=1.2.1",
"requests>=2.32.5", "requests>=2.32.5",
"rich>=14.2.0", "rich>=14.2.0",
"torch>=2.12.0",
"torchvision>=0.27.0",
"transformers>=5.12.0",
"ultralytics>=8.4.66",
] ]
[project.optional-dependencies]
gpu = [
"onnxruntime-gpu>=1.27.0; sys_platform == 'linux' and platform_machine == 'x86_64'",
"nvidia-cudnn-cu12>=9.24.0.43; sys_platform == 'linux' and platform_machine == 'x86_64'",
"nvidia-cuda-runtime-cu12>=12.0; sys_platform == 'linux' and platform_machine == 'x86_64'",
"nvidia-cufft-cu12>=11.0; sys_platform == 'linux' and platform_machine == 'x86_64'",
"nvidia-curand-cu12>=10.0; sys_platform == 'linux' and platform_machine == 'x86_64'",
]
rocm = ["onnxruntime-rocm>=1.16.0; sys_platform == 'linux' and platform_machine == 'x86_64'"]
intel = ["onnxruntime-openvino>=1.20.0; sys_platform == 'linux' and platform_machine == 'x86_64'"]
cpu = ["onnxruntime>=1.27.0"]
[project.scripts] [project.scripts]
winnow = "winnow.cli:main" winnow = "winnow.cli:main"
@@ -42,10 +46,13 @@ Changelog = "https://github.com/sudolulo/winnow/blob/main/CHANGELOG.md"
Documentation = "https://github.com/sudolulo/winnow/wiki" Documentation = "https://github.com/sudolulo/winnow/wiki"
[tool.uv] [tool.uv]
index-strategy = "unsafe-best-match"
conflicts = [ conflicts = [
[ [
{ package = "onnxruntime" }, { extra = "gpu" },
{ package = "onnxruntime-gpu" }, { extra = "rocm" },
{ extra = "intel" },
{ extra = "cpu" },
], ],
] ]
required-environments = [ required-environments = [
@@ -53,32 +60,11 @@ required-environments = [
"sys_platform == 'linux' and platform_machine == 'aarch64'", "sys_platform == 'linux' and platform_machine == 'aarch64'",
] ]
[tool.uv.sources]
torch = [
{ index = "pytorch-cu126", marker = "sys_platform == 'linux' and platform_machine == 'x86_64'" },
{ index = "pytorch-cpu", marker = "sys_platform == 'linux' and platform_machine == 'aarch64'" },
{ index = "pytorch-cpu", marker = "sys_platform != 'linux'" },
]
torchvision = [
{ index = "pytorch-cu126", marker = "sys_platform == 'linux' and platform_machine == 'x86_64'" },
{ index = "pytorch-cpu", marker = "sys_platform == 'linux' and platform_machine == 'aarch64'" },
{ index = "pytorch-cpu", marker = "sys_platform != 'linux'" },
]
[[tool.uv.index]]
name = "pytorch-cu126"
url = "https://download.pytorch.org/whl/cu126"
explicit = true
[[tool.uv.index]]
name = "pytorch-cpu"
url = "https://download.pytorch.org/whl/cpu"
explicit = true
[dependency-groups] [dependency-groups]
dev = [ dev = [
"pytest>=8.0", "pytest>=9.1.1",
"ruff>=0.15.17", "ruff>=0.15.20",
] ]
[tool.hatch.build.targets.wheel] [tool.hatch.build.targets.wheel]
@@ -102,14 +88,14 @@ insightface = "insightface"
numpy = "numpy" numpy = "numpy"
nvidia-cudnn-cu12 = "nvidia.cudnn" nvidia-cudnn-cu12 = "nvidia.cudnn"
onnxruntime-gpu = "onnxruntime" onnxruntime-gpu = "onnxruntime"
onnxruntime-rocm = "onnxruntime"
onnxruntime-openvino = "onnxruntime"
onnxruntime = "onnxruntime"
requests = "requests" requests = "requests"
rich = "rich" rich = "rich"
torch = "torch"
transformers = "transformers"
ultralytics = "ultralytics"
[tool.deptry.per_rule_ignores] [tool.deptry.per_rule_ignores]
DEP002 = ["onnxruntime-gpu", "nvidia-cudnn-cu12"] DEP002 = ["onnxruntime-gpu", "nvidia-cudnn-cu12", "onnxruntime-rocm", "onnxruntime-openvino", "onnxruntime"]
[tool.pytest.ini_options] [tool.pytest.ini_options]
testpaths = ["tests"] testpaths = ["tests"]
@@ -117,4 +103,3 @@ testpaths = ["tests"]
[build-system] [build-system]
requires = ["hatchling"] requires = ["hatchling"]
build-backend = "hatchling.build" build-backend = "hatchling.build"
+38 -26
View File
@@ -15,38 +15,50 @@ except ImportError:
# resident in memory across all subsequent scheduled runs. # resident in memory across all subsequent scheduled runs.
from winnow.cli import main from winnow.cli import main
SCHEDULE = os.environ["CRON_SCHEDULE"]
MODELS_DIR = os.environ.get("HF_HOME", "/models/huggingface")
INSIGHTFACE_HOME = os.environ.get("INSIGHTFACE_HOME", "/models/.insightface")
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
def check_models() -> None: def _check_models() -> None:
buffalo = Path(INSIGHTFACE_HOME) / "models" / "buffalo_l" insightface_home = os.environ.get("INSIGHTFACE_HOME", "/models/.insightface")
hf_hub = Path(MODELS_DIR) / "hub" buffalo = Path(insightface_home) / "models" / "buffalo_l"
if not buffalo.exists(): if not buffalo.exists():
print(" InsightFace Buffalo_L not found — will download on first run", flush=True) print(" InsightFace Buffalo_L not found — will download on first run", flush=True)
if not (hf_hub.exists() and any(hf_hub.iterdir())):
print(" HuggingFace models not found — will download on first run", flush=True)
NOW = time.time() def _run_scheduler() -> None:
cron = croniter(SCHEDULE, NOW) schedule = os.environ.get("CRON_SCHEDULE")
next_run = cron.get_next(float) if not schedule:
print("Error: CRON_SCHEDULE environment variable is required.", flush=True)
sys.exit(1)
try:
Path("/tmp/winnow.pid").write_text(str(os.getpid()))
except OSError as e:
print(f"Warning: could not write PID file: {e}", flush=True)
while True:
now = time.time() now = time.time()
if now >= next_run: cron = croniter(schedule, now)
print(f"\n[{time.strftime('%Y-%m-%d %H:%M:%S')}] Starting winnow run...", flush=True) next_run = cron.get_next(float)
check_models() print(f"Next run: {time.strftime('%Y-%m-%d %H:%M:%S', time.localtime(next_run))}", flush=True)
try:
main() while True:
print("winnow run complete", flush=True) now = time.time()
except KeyboardInterrupt: if now >= next_run:
raise print(f"\n[{time.strftime('%Y-%m-%d %H:%M:%S')}] Starting winnow run...", flush=True)
except Exception as e: _check_models()
logger.error(f"winnow run failed: {e}", exc_info=True) try:
print(f"winnow run failed: {e}", flush=True) main()
next_run = cron.get_next(float) print("winnow run complete", flush=True)
time.sleep(max(1, next_run - time.time())) except (KeyboardInterrupt, SystemExit):
raise
except Exception as e:
logger.error("winnow run failed: %s", e, exc_info=True)
print(f"winnow run failed: {e}", flush=True)
cron = croniter(schedule, time.time())
next_run = cron.get_next(float)
print(f"Next run: {time.strftime('%Y-%m-%d %H:%M:%S', time.localtime(next_run))}", flush=True)
time.sleep(min(60, max(1, next_run - time.time())))
if __name__ == "__main__":
_run_scheduler()
+4 -70
View File
@@ -2,8 +2,8 @@
""" """
winnow inference benchmark: GPU vs CPU throughput. winnow inference benchmark: GPU vs CPU throughput.
Measures InsightFace (face mode) and SigLIP (object mode) latency and Measures InsightFace (ArcFace) latency and throughput.
throughput. Run with FORCE_CPU=true for CPU-only baseline. Run with FORCE_CPU=true for CPU-only baseline.
Usage inside container: Usage inside container:
# GPU mode: # GPU mode:
@@ -13,7 +13,6 @@ Usage inside container:
docker exec -e FORCE_CPU=true winnow python /app/scripts/benchmark.py docker exec -e FORCE_CPU=true winnow python /app/scripts/benchmark.py
""" """
import os
import sys import sys
import time import time
@@ -22,9 +21,8 @@ from PIL import Image, ImageDraw
def _mode_label() -> str: def _mode_label() -> str:
if os.getenv("FORCE_CPU", "").lower() in ("true", "1", "yes"): from winnow.config import _getenv_bool
return "CPU (FORCE_CPU=true)" return "CPU (FORCE_CPU=true)" if _getenv_bool("FORCE_CPU", False) else "GPU (auto)"
return "GPU (auto)"
def make_face_image(size: int = 640) -> Image.Image: def make_face_image(size: int = 640) -> Image.Image:
@@ -47,11 +45,6 @@ def make_face_image(size: int = 640) -> Image.Image:
return img return img
def make_random_image(width: int = 224, height: int = 224) -> Image.Image:
rng = np.random.default_rng(42)
return Image.fromarray(rng.integers(0, 256, (height, width, 3), dtype=np.uint8), "RGB")
def _stats(times_s: list[float]) -> dict: def _stats(times_s: list[float]) -> dict:
arr = np.array(times_s) * 1000 # ms arr = np.array(times_s) * 1000 # ms
return { return {
@@ -119,61 +112,6 @@ def bench_insightface(n_warmup: int = 5, n_runs: int = 30) -> None:
print(f" 320×320 median : {s2['median_ms']:.1f} ms ({s2['ips']:.1f} img/s)") print(f" 320×320 median : {s2['median_ms']:.1f} ms ({s2['ips']:.1f} img/s)")
def bench_siglip(
n_warmup: int = 3,
n_runs: int = 20,
batch_sizes: tuple = (1, 4, 8, 16, 32),
) -> None:
import torch
import winnow.embeddings as emb_mod
emb_mod._siglip_model = None
emb_mod._siglip_processor = None
emb_mod._siglip_loaded = False
print(" Loading model...")
t_load = time.perf_counter()
model, processor = emb_mod.get_siglip_model()
load_s = time.perf_counter() - t_load
if model is None:
print(" SKIP: SigLIP failed to load")
return
device = next(model.parameters()).device
print(f" Model load time : {load_s:.2f} s (device: {device})")
print(f" {'Batch':>5} {'ms/batch':>10} {'ms/img':>8} {'img/s':>8} {'p95/img':>9}")
for bs in batch_sizes:
imgs = [make_random_image(224, 224) for _ in range(bs)]
inputs = processor(images=imgs, return_tensors="pt")
inputs = {k: v.to(device) for k, v in inputs.items()}
# Warmup
for _ in range(n_warmup):
with torch.no_grad():
model(**inputs)
if str(device) != "cpu":
torch.cuda.synchronize()
times: list[float] = []
for _ in range(n_runs):
if str(device) != "cpu":
torch.cuda.synchronize()
t0 = time.perf_counter()
with torch.no_grad():
model(**inputs)
if str(device) != "cpu":
torch.cuda.synchronize()
times.append(time.perf_counter() - t0)
s = _stats(times)
print(
f" {bs:>5} {s['median_ms']:>10.1f} {s['median_ms']/bs:>8.2f}"
f" {bs * 1000 / s['median_ms']:>8.1f} {s['p95_ms']/bs:>9.2f}"
)
def main() -> None: def main() -> None:
print("=" * 56) print("=" * 56)
print(" winnow inference benchmark") print(" winnow inference benchmark")
@@ -185,10 +123,6 @@ def main() -> None:
bench_insightface() bench_insightface()
print() print()
print("── SigLIP google/siglip-base-patch16-224 (objects) ───")
bench_siglip()
print()
if __name__ == "__main__": if __name__ == "__main__":
# Add winnow to path when run directly inside container # Add winnow to path when run directly inside container
+2 -2
View File
@@ -18,10 +18,10 @@ def test_config_loads_defaults(monkeypatch):
assert cfg.OUTPUT_DIR == "./frigate_train" assert cfg.OUTPUT_DIR == "./frigate_train"
assert cfg.YEARS_FILTER == 10 assert cfg.YEARS_FILTER == 10
assert cfg.MIN_FACE_WIDTH == 90 assert cfg.MIN_FACE_WIDTH == 90
assert cfg.MIN_FACE_COUNT == 0 assert cfg.MIN_FACE_COUNT == 3
assert cfg.BLUR_THRESHOLD == 120.0 assert cfg.BLUR_THRESHOLD == 120.0
assert cfg.MIN_CONFIDENCE == 0.7 assert cfg.MIN_CONFIDENCE == 0.7
assert cfg.MAX_AUTO_IMAGES == 80 assert cfg.MAX_AUTO_IMAGES == 5
assert cfg.QUALITY_REPLACEMENT is True assert cfg.QUALITY_REPLACEMENT is True
assert cfg.FACE_MARGIN == 0.15 assert cfg.FACE_MARGIN == 0.15
assert cfg.USE_FULL_RESOLUTION is True assert cfg.USE_FULL_RESOLUTION is True
+336
View File
@@ -0,0 +1,336 @@
"""Tests for core diversity selection algorithms (pure functions, no network)."""
import numpy as np
import pytest
from PIL import Image
def _unit_embeddings(n: int, d: int = 512, seed: int = 42) -> list:
"""Return n normalised random embeddings — well-separated in high-dim space."""
rng = np.random.default_rng(seed)
embs = rng.standard_normal((n, d)).astype(np.float32)
embs /= np.linalg.norm(embs, axis=1, keepdims=True)
return list(embs)
def _asset_with_face(
asset_id="a1", person_id="p1",
x1=10, y1=10, x2=60, y2=60,
img_w=100, img_h=100, score=0.9,
):
return {
"id": asset_id,
"people": [{
"id": person_id,
"faces": [{
"boundingBoxX1": x1, "boundingBoxY1": y1,
"boundingBoxX2": x2, "boundingBoxY2": y2,
"imageWidth": img_w, "imageHeight": img_h,
"score": score,
}],
}],
}
# ── _get_face_bbox ─────────────────────────────────────────────────────────────
def test_get_face_bbox_returns_coords():
from winnow.diversity import _get_face_bbox
asset = _asset_with_face(x1=5, y1=10, x2=55, y2=70)
assert _get_face_bbox(asset) == (5, 10, 55, 70)
def test_get_face_bbox_filters_by_person_id():
from winnow.diversity import _get_face_bbox
asset = _asset_with_face(person_id="p1")
assert _get_face_bbox(asset, person_id="p999") is None
def test_get_face_bbox_returns_none_empty_faces():
from winnow.diversity import _get_face_bbox
asset = {"id": "a1", "people": [{"id": "p1", "faces": []}]}
assert _get_face_bbox(asset) is None
def test_get_face_bbox_returns_none_no_people():
from winnow.diversity import _get_face_bbox
assert _get_face_bbox({"id": "a1"}) is None
# ── _get_face_confidence ───────────────────────────────────────────────────────
def test_get_face_confidence_returns_score():
from winnow.diversity import _get_face_confidence
asset = _asset_with_face(score=0.92)
assert _get_face_confidence(asset) == pytest.approx(0.92)
def test_get_face_confidence_filters_by_person_id():
from winnow.diversity import _get_face_confidence
asset = _asset_with_face(person_id="p1", score=0.9)
assert _get_face_confidence(asset, person_id="p999") is None
def test_get_face_confidence_returns_none_empty_faces():
from winnow.diversity import _get_face_confidence
asset = {"id": "a1", "people": [{"id": "p1", "faces": []}]}
assert _get_face_confidence(asset) is None
# ── _crop_face_from_thumbnail ──────────────────────────────────────────────────
def test_crop_face_returns_image():
from winnow.diversity import _crop_face_from_thumbnail
img = Image.new("RGB", (200, 200), color=(128, 64, 32))
asset = _asset_with_face(x1=50, y1=50, x2=150, y2=150, img_w=200, img_h=200)
crop = _crop_face_from_thumbnail(img, asset)
assert crop is not None
assert crop.width > 0 and crop.height > 0
def test_crop_face_scales_bbox_to_thumbnail():
"""When thumbnail is half the metadata dimensions, bbox is scaled accordingly."""
from winnow.diversity import _crop_face_from_thumbnail
img = Image.new("RGB", (200, 200))
# Metadata says 400×400; bbox covers the centre quarter
asset = _asset_with_face(x1=100, y1=100, x2=300, y2=300, img_w=400, img_h=400)
crop = _crop_face_from_thumbnail(img, asset)
assert crop is not None
assert crop.width <= 200 and crop.height <= 200
def test_crop_face_returns_none_no_metadata():
from winnow.diversity import _crop_face_from_thumbnail
img = Image.new("RGB", (100, 100))
assert _crop_face_from_thumbnail(img, {"id": "a1"}) is None
def test_crop_face_returns_none_for_sub_30px_bbox():
"""A 1×1 bbox produces a crop too small to embed — should be rejected."""
from winnow.diversity import _crop_face_from_thumbnail
img = Image.new("RGB", (100, 100))
asset = _asset_with_face(x1=50, y1=50, x2=51, y2=51, img_w=100, img_h=100)
assert _crop_face_from_thumbnail(img, asset) is None
def test_crop_face_respects_person_id_filter():
from winnow.diversity import _crop_face_from_thumbnail
img = Image.new("RGB", (200, 200))
asset = _asset_with_face(person_id="p1", x1=50, y1=50, x2=150, y2=150)
assert _crop_face_from_thumbnail(img, asset, person_id="p999") is None
# ── _dedup_embeddings ──────────────────────────────────────────────────────────
def test_dedup_keeps_all_diverse_embeddings():
from winnow.diversity import _dedup_embeddings
embs = _unit_embeddings(20)
candidates = [{"id": str(i)} for i in range(20)]
_, out_cands, _ = _dedup_embeddings(embs, candidates, [None] * 20)
# Random 512-dim unit vectors are far apart — all should survive
assert len(out_cands) == 20
def test_dedup_removes_near_duplicate():
from winnow.diversity import _dedup_embeddings
base = np.zeros(512, dtype=np.float32)
base[0] = 1.0
# Cosine distance ≈ 0.01 — well within the 0.20 dedup threshold
near_dup = base.copy()
near_dup[1] = 0.014
near_dup /= np.linalg.norm(near_dup)
embs = [base, near_dup]
candidates = [{"id": "base", "quality_score": 0.9}, {"id": "dup", "quality_score": 0.5}]
_, out_cands, _ = _dedup_embeddings(embs, candidates, [None, None])
assert len(out_cands) == 1
assert out_cands[0]["id"] == "base"
def test_dedup_keeps_higher_quality_from_duplicate_pair():
from winnow.diversity import _dedup_embeddings
base = np.zeros(512, dtype=np.float32)
base[0] = 1.0
near_dup = base.copy()
near_dup[1] = 0.014
near_dup /= np.linalg.norm(near_dup)
# Reversed quality: near_dup is sharper
embs = [base, near_dup]
candidates = [{"id": "base", "quality_score": 0.3}, {"id": "dup", "quality_score": 0.95}]
_, out_cands, _ = _dedup_embeddings(embs, candidates, [None, None])
assert len(out_cands) == 1
assert out_cands[0]["id"] == "dup"
def test_dedup_single_embedding_passes_through():
from winnow.diversity import _dedup_embeddings
embs = _unit_embeddings(1)
out_embs, out_cands, _ = _dedup_embeddings(embs, [{"id": "only"}], [None])
assert len(out_cands) == 1
def test_dedup_treats_zero_quality_score_as_zero_not_missing():
"""quality_score=0.0 is a valid score — should not be treated as absent."""
from winnow.diversity import _dedup_embeddings
base = np.zeros(512, dtype=np.float32)
base[0] = 1.0
near_dup = base.copy()
near_dup[1] = 0.014
near_dup /= np.linalg.norm(near_dup)
embs = [base, near_dup]
# base has explicit 0.0; near_dup has 0.5 — near_dup should win
candidates = [{"id": "base", "quality_score": 0.0}, {"id": "dup", "quality_score": 0.5}]
_, out_cands, _ = _dedup_embeddings(embs, candidates, [None, None])
assert out_cands[0]["id"] == "dup"
# ── _kmedoids ──────────────────────────────────────────────────────────────────
def _dist_matrix(embs):
m = np.vstack(embs)
m /= np.linalg.norm(m, axis=1, keepdims=True)
return 1 - m @ m.T
def test_kmedoids_returns_k_distinct_medoids():
from winnow.diversity import _kmedoids
dist = _dist_matrix(_unit_embeddings(30))
medoids, _ = _kmedoids(dist, k=5)
assert len(medoids) == 5
assert len(set(medoids)) == 5
def test_kmedoids_labels_cover_all_points():
from winnow.diversity import _kmedoids
dist = _dist_matrix(_unit_embeddings(20))
medoids, labels = _kmedoids(dist, k=4)
assert len(labels) == 20
assert set(labels).issubset(set(range(4)))
def test_kmedoids_medoids_are_valid_indices():
from winnow.diversity import _kmedoids
n = 15
dist = _dist_matrix(_unit_embeddings(n))
medoids, _ = _kmedoids(dist, k=3)
assert all(0 <= m < n for m in medoids)
def test_kmedoids_k_equals_n_selects_all():
from winnow.diversity import _kmedoids
n = 5
dist = _dist_matrix(_unit_embeddings(n))
medoids, _ = _kmedoids(dist, k=n)
assert len(medoids) == n
# ── _compute_adaptive_threshold ────────────────────────────────────────────────
def test_adaptive_threshold_positive():
from winnow.diversity import _compute_adaptive_threshold
embs = np.array(_unit_embeddings(50))
assert _compute_adaptive_threshold(embs) > 0
def test_adaptive_threshold_floor_for_identical_embeddings():
"""All-identical embeddings → median pairwise distance = 0 → floor at 0.05."""
from winnow.diversity import _compute_adaptive_threshold
base = np.zeros((10, 512), dtype=np.float32)
base[:, 0] = 1.0
assert _compute_adaptive_threshold(base) == pytest.approx(0.05)
def test_adaptive_threshold_single_point_returns_floor():
from winnow.diversity import _compute_adaptive_threshold
single = np.ones((1, 512), dtype=np.float32)
single /= np.linalg.norm(single)
assert _compute_adaptive_threshold(single) == pytest.approx(0.05)
def test_adaptive_threshold_scales_with_spread():
"""A more spread-out embedding set should produce a higher threshold."""
from winnow.diversity import _compute_adaptive_threshold
tight = np.array(_unit_embeddings(30, seed=0)) * 0.001 + np.array([1.0] + [0.0] * 511)
tight /= np.linalg.norm(tight, axis=1, keepdims=True)
diverse = np.array(_unit_embeddings(30, seed=1))
assert _compute_adaptive_threshold(diverse) > _compute_adaptive_threshold(tight)
# ── _select_time_spread ────────────────────────────────────────────────────────
def test_time_spread_returns_exact_n():
from winnow.diversity import _select_time_spread
assets = [{"id": str(i)} for i in range(100)]
assert len(_select_time_spread(assets, limit=10)) == 10
def test_time_spread_returns_all_when_under_limit():
from winnow.diversity import _select_time_spread
assets = [{"id": str(i)} for i in range(5)]
assert len(_select_time_spread(assets, limit=20)) == 5
def test_time_spread_auto_defaults_to_30():
from winnow.diversity import _select_time_spread
assets = [{"id": str(i)} for i in range(200)]
assert len(_select_time_spread(assets, limit="auto")) == 30
def test_time_spread_includes_first_and_last():
from winnow.diversity import _select_time_spread
assets = [{"id": str(i)} for i in range(100)]
result = _select_time_spread(assets, limit=5)
ids = [int(a["id"]) for a in result]
assert ids[0] == 0
assert ids[-1] == 99
# ── _cluster_aware_selection ───────────────────────────────────────────────────
def test_cluster_selection_returns_exact_limit(monkeypatch):
from winnow.diversity import _cluster_aware_selection
monkeypatch.setattr("winnow.diversity.Config.MAX_AUTO_IMAGES", 20)
embs = _unit_embeddings(50)
candidates = [{"id": str(i)} for i in range(50)]
result = _cluster_aware_selection(embs, candidates, limit=10)
assert len(result) == 10
def test_cluster_selection_output_is_subset_of_input(monkeypatch):
from winnow.diversity import _cluster_aware_selection
monkeypatch.setattr("winnow.diversity.Config.MAX_AUTO_IMAGES", 20)
embs = _unit_embeddings(30)
candidates = [{"id": str(i)} for i in range(30)]
result = _cluster_aware_selection(embs, candidates, limit=10)
result_ids = {a["id"] for a in result}
assert result_ids.issubset({a["id"] for a in candidates})
def test_cluster_selection_auto_stops_early_on_tight_cluster(monkeypatch):
"""When all embeddings are nearly identical auto mode should stop early."""
from winnow.diversity import _cluster_aware_selection
monkeypatch.setattr("winnow.diversity.Config.MAX_AUTO_IMAGES", 20)
rng = np.random.default_rng(0)
base = np.zeros(512, dtype=np.float32)
base[0] = 1.0
embs = []
for _ in range(50):
v = base + rng.standard_normal(512).astype(np.float32) * 0.001
v /= np.linalg.norm(v)
embs.append(v)
candidates = [{"id": str(i)} for i in range(50)]
result = _cluster_aware_selection(list(embs), candidates, limit="auto")
assert len(result) < 20
def test_cluster_selection_hard_example_weighting_accepted(monkeypatch):
"""Confidence scores are accepted without error."""
from winnow.diversity import _cluster_aware_selection
monkeypatch.setattr("winnow.diversity.Config.MAX_AUTO_IMAGES", 20)
embs = _unit_embeddings(20)
candidates = [{"id": str(i)} for i in range(20)]
conf = [0.7 if i % 2 == 0 else 0.95 for i in range(20)]
result = _cluster_aware_selection(embs, candidates, limit=5, confidence_scores=conf)
assert len(result) == 5
+1 -1
View File
@@ -7,7 +7,7 @@ import pytest
@pytest.fixture(autouse=True) @pytest.fixture(autouse=True)
def isolated_cache(monkeypatch, tmp_path): def isolated_cache(monkeypatch, tmp_path):
"""Point tracker at a temp directory so tests don't touch real cache files.""" """Point tracker at a temp directory so tests don't touch real cache files."""
monkeypatch.setenv("CACHE_DIR", str(tmp_path)) monkeypatch.setenv("DATA_DIR", str(tmp_path))
from winnow.config import _Config from winnow.config import _Config
_Config.reset() _Config.reset()
yield tmp_path yield tmp_path
-1884
View File
File diff suppressed because it is too large Load Diff
-1941
View File
File diff suppressed because it is too large Load Diff
-1933
View File
File diff suppressed because it is too large Load Diff
Generated
+263 -1712
View File
File diff suppressed because it is too large Load Diff
+1 -2
View File
@@ -1,8 +1,7 @@
"""Immich to Frigate training set curator. """Immich to Frigate training set curator.
AI-powered tool to extract high-quality, diverse training images from your AI-powered tool to extract high-quality, diverse training images from your
Immich library for Frigate's Face Recognition (ArcFace) and Object/State Immich library for Frigate's face recognition (ArcFace/Buffalo_L).
Classification models.
""" """
from importlib.metadata import PackageNotFoundError, version from importlib.metadata import PackageNotFoundError, version
+55 -19
View File
@@ -7,38 +7,62 @@ recomputing on reruns. Uses numpy binary format for fast I/O.
import hashlib import hashlib
import logging import logging
import os import os
from pathlib import Path
import numpy as np import numpy as np
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# Model versions — bump these when the upstream model changes
MODEL_VERSIONS = { MODEL_VERSIONS = {
"insightface": "buffalo_l_v1",
"siglip": "siglip-base-patch16-224_v1",
"immich": "immich_buffalo_l_v1", "immich": "immich_buffalo_l_v1",
} }
def _insightface_model_fingerprint() -> str:
"""Derive a version string from buffalo_l .onnx file sizes and mtimes.
Changes automatically when model files are replaced or updated, preventing
stale embeddings from a previous model being served from cache.
Falls back to a static string before the model is downloaded (first run).
"""
insightface_home = os.environ.get("INSIGHTFACE_HOME", os.path.expanduser("~/.insightface"))
model_dir = Path(insightface_home) / "models" / "buffalo_l"
if not model_dir.exists():
return "buffalo_l_v1"
onnx_files = sorted(model_dir.glob("*.onnx"))
if not onnx_files:
return "buffalo_l_v1"
fingerprint = "|".join(
f"{f.name}:{f.stat().st_size}:{int(f.stat().st_mtime)}"
for f in onnx_files
)
return hashlib.sha256(fingerprint.encode()).hexdigest()[:12]
class EmbeddingCache: class EmbeddingCache:
"""Simple disk-based embedding cache. """Simple disk-based embedding cache.
Embeddings are stored as .npy files in a flat directory, Embeddings are stored as .npy files in a flat directory,
keyed by a hash of (asset_id, model_version). keyed by a hash of (asset_id, model_version). The InsightFace version
is derived from buffalo_l model file metadata so the cache auto-invalidates
when model files are replaced or updated.
""" """
def __init__(self, cache_dir: str = ".if_cache") -> None: def __init__(self, cache_dir: str = ".if_cache") -> None:
self.cache_dir = cache_dir self.cache_dir = cache_dir
self._ensured = False self._ensured = False
self._model_versions = {
**MODEL_VERSIONS,
"insightface": _insightface_model_fingerprint(),
}
def _ensure_dir(self) -> None: def _ensure_dir(self) -> None:
if not self._ensured: if not self._ensured:
os.makedirs(self.cache_dir, exist_ok=True) os.makedirs(self.cache_dir, exist_ok=True)
self._ensured = True self._ensured = True
@staticmethod def _key(self, asset_id: str, model: str) -> str:
def _key(asset_id: str, model: str) -> str: version = self._model_versions.get(model, model)
version = MODEL_VERSIONS.get(model, model)
raw = f"{asset_id}:{version}" raw = f"{asset_id}:{version}"
return hashlib.sha256(raw.encode()).hexdigest()[:16] return hashlib.sha256(raw.encode()).hexdigest()[:16]
@@ -58,10 +82,19 @@ class EmbeddingCache:
def put(self, asset_id: str, embedding: np.ndarray, model: str = "insightface") -> None: def put(self, asset_id: str, embedding: np.ndarray, model: str = "insightface") -> None:
"""Store an embedding in the cache.""" """Store an embedding in the cache."""
self._ensure_dir() self._ensure_dir()
final = self._path(asset_id, model)
# Insert .tmp before .npy so np.save doesn't auto-append another .npy extension
# (np.save appends .npy to paths that don't already end in .npy).
tmp = final.removesuffix(".npy") + ".tmp.npy"
try: try:
np.save(self._path(asset_id, model), embedding) np.save(tmp, embedding)
os.replace(tmp, final)
except Exception as e: except Exception as e:
logger.debug(f"Cache write failed for {asset_id}: {e}") logger.warning("Cache write failed for %s: %s", asset_id, e)
try:
os.remove(tmp)
except OSError:
pass
def clear(self) -> None: def clear(self) -> None:
"""Delete all cached embeddings.""" """Delete all cached embeddings."""
@@ -70,25 +103,28 @@ class EmbeddingCache:
count = 0 count = 0
for f in os.listdir(self.cache_dir): for f in os.listdir(self.cache_dir):
if f.endswith(".npy"): if f.endswith(".npy"):
os.remove(os.path.join(self.cache_dir, f)) try:
count += 1 os.remove(os.path.join(self.cache_dir, f))
logger.info(f"Cleared {count} cached embeddings.") count += 1
except OSError:
pass
logger.info("Cleared %s cached embeddings.", count)
# Singleton instance # Singleton instance
_cache: EmbeddingCache | None = None _cache: EmbeddingCache | None = None
_cache_dir: str | None = None
def get_cache(cache_dir: str = ".if_cache") -> EmbeddingCache: def get_cache(cache_dir: str = ".if_cache") -> EmbeddingCache:
"""Get or create the singleton cache instance. """Get or create the singleton cache instance.
Note: The ``cache_dir`` parameter is only used when creating the Re-creates the instance when ``cache_dir`` changes so that test
singleton for the first time. Subsequent calls return the existing isolation (which resets Config.DATA_DIR via _Config.reset()) always
instance regardless of ``cache_dir``. If you need a cache with a writes to the correct directory rather than a stale one.
different directory, instantiate ``EmbeddingCache`` directly.
""" """
global _cache global _cache, _cache_dir
if _cache is None: if _cache is None or _cache_dir != cache_dir:
_cache = EmbeddingCache(cache_dir) _cache = EmbeddingCache(cache_dir)
_cache_dir = cache_dir
return _cache return _cache
+87 -24
View File
@@ -7,12 +7,13 @@ import sys
from rich import print as rprint from rich import print as rprint
from rich.prompt import Confirm from rich.prompt import Confirm
from .config import Config, ConfigManager from . import __version__
from .config import Config, _getenv_bool
from .executor import execute_jobs, upload_to_frigate from .executor import execute_jobs, upload_to_frigate
from .immich_api import get_people, merge_people from .immich_api import get_immich_version, get_people, merge_people
from .jobs import _show_preview, auto_configure, interactive_configure from .jobs import _show_preview, auto_configure, interactive_configure
from .log_config import console, setup_logging from .log_config import console, setup_logging
from .upload_tracker import find_by_crop_dimension, get_person_summary, reset_person from .upload_tracker import find_by_crop_dimension, get_person_summary, reset_all_people, reset_person
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -71,19 +72,33 @@ def _handle_duplicate_people(people: list[dict]) -> list[dict]:
by_name: dict[str, list[dict]] = defaultdict(list) by_name: dict[str, list[dict]] = defaultdict(list)
for p in people: for p in people:
name = (p.get("name") or "").strip() name = (p.get("name") or "").strip()
if name: if name and p.get("id"):
by_name[name].append(p) by_name[name].append(p)
duplicates = {name: ps for name, ps in by_name.items() if len(ps) > 1} duplicates = {name: ps for name, ps in by_name.items() if len(ps) > 1}
if not duplicates: if not duplicates:
return people return people
def _smaller_duplicate_ids(groups: dict) -> set[str]:
"""IDs of all but the largest person in each duplicate group."""
return {
pid
for ps in groups.values()
for p in sorted(ps, key=lambda x: x.get("assetCount", 0), reverse=True)[1:]
if (pid := p.get("id"))
}
skip_ids = _smaller_duplicate_ids(duplicates)
def _excl(lst: list[dict]) -> list[dict]:
return [p for p in lst if p.get("id") not in skip_ids]
if not Config.MERGE_DUPLICATE_PEOPLE: if not Config.MERGE_DUPLICATE_PEOPLE:
rprint("\n[bold yellow]⚠ Duplicate person names detected in Immich:[/bold yellow]") rprint("\n[bold yellow]⚠ Duplicate person names detected in Immich:[/bold yellow]")
for name, ps in sorted(duplicates.items()): for name, ps in sorted(duplicates.items()):
ordered = sorted(ps, key=lambda x: x.get("assetCount", 0), reverse=True) ordered = sorted(ps, key=lambda x: x.get("assetCount", 0), reverse=True)
entries = ", ".join( entries = ", ".join(
f"[dim]{p['id'][:8]}…[/dim] ({p.get('assetCount', 0)} assets)" f"[dim]{(p.get('id') or '?')[:8]}…[/dim] ({p.get('assetCount', 0)} assets)"
for p in ordered for p in ordered
) )
rprint(f" [yellow]{name}[/yellow] → {len(ps)} people: {entries}") rprint(f" [yellow]{name}[/yellow] → {len(ps)} people: {entries}")
@@ -99,25 +114,21 @@ def _handle_duplicate_people(people: list[dict]) -> list[dict]:
) )
# Return deduplicated list — keep only the largest per name so that # Return deduplicated list — keep only the largest per name so that
# downstream job creation never runs two jobs for the same Frigate folder. # downstream job creation never runs two jobs for the same Frigate folder.
skip_ids = { return _excl(people)
p["id"]
for ps in duplicates.values()
for p in sorted(ps, key=lambda x: x.get("assetCount", 0), reverse=True)[1:]
}
return [p for p in people if p["id"] not in skip_ids]
# Auto-merge: survivor = largest asset count, rest merge into it inside Immich # Auto-merge: survivor = largest asset count, rest merge into it inside Immich
merged_any = False merged_any = False
for name, ps in sorted(duplicates.items()): for name, ps in sorted(duplicates.items()):
ordered = sorted(ps, key=lambda x: x.get("assetCount", 0), reverse=True) ordered = sorted(ps, key=lambda x: x.get("assetCount", 0), reverse=True)
survivor = ordered[0] survivor = ordered[0]
merge_ids = [p["id"] for p in ordered[1:]] survivor_id = survivor.get("id")
merge_ids = [pid for p in ordered[1:] if (pid := p.get("id")) is not None]
rprint( rprint(
f" [cyan]Merging {name!r} inside Immich:[/cyan] keeping " f" [cyan]Merging {name!r} inside Immich:[/cyan] keeping "
f"[dim]{survivor['id'][:8]}…[/dim] ({survivor.get('assetCount', 0)} assets), " f"[dim]{survivor_id[:8]}…[/dim] ({survivor.get('assetCount', 0)} assets), "
f"absorbing {len(merge_ids)} smaller duplicate(s)..." f"absorbing {len(merge_ids)} smaller duplicate(s)..."
) )
if merge_people(survivor["id"], merge_ids): if merge_people(survivor_id, merge_ids):
rprint(f" [green]✓ Merged {name!r}[/green]") rprint(f" [green]✓ Merged {name!r}[/green]")
merged_any = True merged_any = True
else: else:
@@ -125,27 +136,73 @@ def _handle_duplicate_people(people: list[dict]) -> list[dict]:
if merged_any: if merged_any:
rprint(" [dim]Re-fetching people after merge...[/dim]") rprint(" [dim]Re-fetching people after merge...[/dim]")
return get_people() fresh = get_people()
if not fresh:
# Retry once: get_people() returns [] for both transient failures and
# auth errors (401); a second empty result strongly suggests a real failure.
fresh = get_people()
if not fresh:
logger.warning(
"Re-fetch after merge returned no people (tried twice)"
" — possible transient error or expired API key;"
" proceeding with pre-merge list. Check IMMICH_API_KEY if this recurs."
)
return _excl(people)
# Filter out the smaller duplicate from any group whose merge failed — those
# IDs still exist in Immich and would produce two jobs for the same folder.
# IDs from groups that merged successfully are already gone from Immich, so
# this filter is a no-op for them.
return _excl(fresh)
return people # All merges failed — fall back to local deduplication (keep largest per name) so
# downstream job creation never runs two jobs for the same Frigate folder.
rprint(
" [yellow]All merges failed — applying local deduplication"
" to avoid overwriting output.[/yellow]"
)
return _excl(people)
_UNSUPPORTED_VARS = [
"ENABLE_FRIGATE_SCORES",
"FRIGATE_SCORE_CEILING",
"MIN_FACE_WIDTH",
"FACE_MARGIN",
"ENABLE_FACE_ALIGNMENT",
"USE_FULL_RESOLUTION",
"MIN_CONFIDENCE",
"BLUR_THRESHOLD",
]
def main() -> None: def main() -> None:
"""Entry point for winnow CLI.""" """Entry point for winnow CLI."""
try: try:
verbose = os.environ.get("VERBOSE", "").lower() in ("true", "1", "yes") verbose = _getenv_bool("VERBOSE", False)
setup_logging(verbose=verbose) setup_logging(verbose=verbose)
trace_size = os.environ.get("TRACE_CROP_SIZE", "").strip() trace_size = os.environ.get("TRACE_CROP_SIZE", "").strip()
if trace_size: if trace_size:
_handle_trace_crop(trace_size) _handle_trace_crop(trace_size)
console.print(r""" console.print(f"""
[bold blue]winnow[/bold blue] [bold blue]winnow[/bold blue] [dim]v{__version__}[/dim]
[dim]Immich -> Frigate Training Data Curator[/dim] [dim]Immich -> Frigate Training Data Curator[/dim]
""") """)
ConfigManager.get().interactive_setup() _FALSY = {"", "false", "0", "no", "off"}
set_unsupported = [v for v in _UNSUPPORTED_VARS if os.environ.get(v, "").strip().lower() not in _FALSY]
if set_unsupported:
console.print(
f"[bold yellow]⚠ Advanced tuning vars set: "
f"{', '.join(set_unsupported)}[/bold yellow]"
)
console.print(
"[dim] These defaults are calibrated for Frigate's ArcFace requirements. "
"Image quality issues caused by non-default values will not be investigated.[/dim]\n"
)
Config.interactive_setup()
try: try:
Config.validate() Config.validate()
@@ -169,8 +226,7 @@ def main() -> None:
"and will be reset along with everyone else.[/yellow]" "and will be reset along with everyone else.[/yellow]"
) )
if names: if names:
for name in names: reset_all_people()
reset_person(name)
rprint(f"[bold yellow]Reset tracking data for all {len(names)} people.[/bold yellow]") rprint(f"[bold yellow]Reset tracking data for all {len(names)} people.[/bold yellow]")
else: else:
rprint("[dim]No tracking data to reset.[/dim]") rprint("[dim]No tracking data to reset.[/dim]")
@@ -193,6 +249,13 @@ def main() -> None:
f" {counts['rejected']} rejected{frigate_part}[/dim]" f" {counts['rejected']} rejected{frigate_part}[/dim]"
) )
_immich_version = get_immich_version()
if _immich_version is not None and _immich_version < (1, 106, 0):
rprint(
f" [yellow]⚠ Immich {'.'.join(str(x) for x in _immich_version)} detected — "
"winnow requires v1.106+. Some features may not work.[/yellow]"
)
people = get_people() people = get_people()
if not people: if not people:
rprint("[bold red]Could not fetch people from Immich. Check URL/Key.[/bold red]") rprint("[bold red]Could not fetch people from Immich. Check URL/Key.[/bold red]")
@@ -202,8 +265,8 @@ def main() -> None:
# Auto mode when no TTY (Docker, cron, pipes) — the primary use case. # Auto mode when no TTY (Docker, cron, pipes) — the primary use case.
# A TTY means local interactive use; AUTO_MODE=true overrides that for scripting. # A TTY means local interactive use; AUTO_MODE=true overrides that for scripting.
auto_mode = not sys.stdin.isatty() or os.environ.get("AUTO_MODE", "").lower() in ("true", "1", "yes") auto_mode = not sys.stdin.isatty() or _getenv_bool("AUTO_MODE", False)
dry_run = os.environ.get("DRY_RUN", "false").lower() in ("true", "1", "yes") dry_run = _getenv_bool("DRY_RUN", False)
if dry_run: if dry_run:
rprint("[bold yellow]DRY RUN — no images will be downloaded or uploaded[/bold yellow]") rprint("[bold yellow]DRY RUN — no images will be downloaded or uploaded[/bold yellow]")
+151 -90
View File
@@ -9,84 +9,183 @@ from typing import ClassVar
from dotenv import load_dotenv from dotenv import load_dotenv
from rich.prompt import Prompt from rich.prompt import Prompt
load_dotenv() _LEGACY_CONFIG_FILE = Path(".immich_config.json") # pre-v0.6: lived in process CWD, not on a volume
CONFIG_FILE = Path(".immich_config.json")
def _getenv_num(name: str, default, cast):
raw = os.getenv(name)
if raw is None:
return default
raw = raw.strip()
if not raw:
return default
try:
return cast(raw)
except ValueError:
logging.warning("%s=%r is not a valid %s — using default %s", name, raw, cast.__name__, default)
return default
def _getenv_int(name: str, default: int) -> int:
return _getenv_num(name, default, int)
def _getenv_float(name: str, default: float) -> float:
return _getenv_num(name, default, float)
def _getenv_optional_float(name: str) -> float | None:
"""Return float value of env var, or None if unset/empty. Warns and returns None on invalid."""
return _getenv_num(name, None, float)
def _getenv_optional_int(name: str) -> int | None:
"""Return int value of env var, or None if unset/empty. Warns and returns None on invalid."""
return _getenv_num(name, None, int)
def _getenv_bool(name: str, default: bool) -> bool:
raw = os.getenv(name)
if raw is None:
return default
raw = raw.strip()
if not raw:
return default
return raw.lower() in ("true", "1", "yes")
class _Config: class _Config:
"""Singleton configuration with uppercase attribute access for backward compatibility.""" """Singleton configuration with lazy loading via __getattr__.
Class-level attributes are annotations only (no defaults), so attribute
access on an un-loaded instance falls through to __getattr__, which
triggers _load() exactly once.
"""
_instance: ClassVar["_Config | None"] = None _instance: ClassVar["_Config | None"] = None
# Configuration values # Annotations only — no class-level defaults so __getattr__ fires on first access
IMMICH_URL: str | None = None IMMICH_URL: str | None
API_KEY: str | None = None API_KEY: str | None
OUTPUT_DIR: str = "./frigate_train" OUTPUT_DIR: str
YEARS_FILTER: int = 10 YEARS_FILTER: int
# Quality filtering # Quality filtering
MIN_FACE_WIDTH: int = 90 MIN_FACE_WIDTH: int
BLUR_THRESHOLD: float = 120.0 BLUR_THRESHOLD: float
MIN_CONFIDENCE: float = 0.7 MIN_CONFIDENCE: float
MAX_AUTO_IMAGES: int = 80 MAX_AUTO_IMAGES: int
QUALITY_REPLACEMENT: bool = True QUALITY_REPLACEMENT: bool
FRIGATE_SCORE_CEILING: float = 0.0 FRIGATE_SCORE_CEILING: float | None
ENABLE_FRIGATE_SCORES: bool = True ENABLE_FRIGATE_SCORES: bool
# People filtering # People filtering
MIN_FACE_COUNT: int = 0 MIN_FACE_COUNT: int
MERGE_DUPLICATE_PEOPLE: bool = False MERGE_DUPLICATE_PEOPLE: bool
# Output quality # Output quality
FACE_MARGIN: float = 0.15 FACE_MARGIN: float
USE_FULL_RESOLUTION: bool = True USE_FULL_RESOLUTION: bool
ENABLE_FACE_ALIGNMENT: bool = True ENABLE_FACE_ALIGNMENT: bool
ENABLE_CACHE: bool = True ENABLE_CACHE: bool
CACHE_DIR: str = ".if_cache" DATA_DIR: str
def __new__(cls) -> "_Config": def __new__(cls) -> "_Config":
if cls._instance is None: if cls._instance is None:
cls._instance = super().__new__(cls) cls._instance = super().__new__(cls)
cls._instance._load() # Do NOT call _load() here — keep __new__ I/O-free so that import
# time does not trigger env/file reads.
return cls._instance return cls._instance
def __getattr__(self, name: str):
"""Called only when the attribute is not found on the instance.
On first access to any config attribute, load all values from env/file
and return the requested one. Re-registers self as _instance so that
a subsequent reset() correctly finds and clears this object's attrs.
"""
if name.startswith("_"):
raise AttributeError(name)
self._load()
# Re-register self as the singleton so reset() can clear our __dict__.
# This handles the case where __getattr__ is called on the module-level
# Config object after a reset() set _instance to None.
_Config._instance = self
# _load() sets the attribute as an instance attr; retrieve it directly
# to avoid infinite recursion through __getattr__.
try:
return self.__dict__[name]
except KeyError:
raise AttributeError(f"_Config has no attribute {name!r}")
def _load(self) -> None: def _load(self) -> None:
"""Load configuration from environment and config file.""" """Load configuration from environment and config file."""
load_dotenv()
# Load from environment (highest priority) # Load from environment (highest priority)
self.IMMICH_URL = os.getenv("IMMICH_URL") self.IMMICH_URL = os.getenv("IMMICH_URL")
self.API_KEY = os.getenv("API_KEY") self.API_KEY = os.getenv("API_KEY")
self.OUTPUT_DIR = os.getenv("OUTPUT_DIR", "./frigate_train") self.OUTPUT_DIR = os.getenv("OUTPUT_DIR", "./frigate_train")
self.YEARS_FILTER = int(os.getenv("YEARS_FILTER", "10")) self.YEARS_FILTER = _getenv_int("YEARS_FILTER", 10)
self.MIN_FACE_WIDTH = int(os.getenv("MIN_FACE_WIDTH", "90")) if self.YEARS_FILTER < 0:
self.MIN_FACE_COUNT = int(os.getenv("MIN_FACE_COUNT", "0")) logging.warning("YEARS_FILTER=%s is negative — using default 10", self.YEARS_FILTER)
self.MERGE_DUPLICATE_PEOPLE = os.getenv("MERGE_DUPLICATE_PEOPLE", "false").lower() in ("true", "1", "yes") self.YEARS_FILTER = 10
self.BLUR_THRESHOLD = float(os.getenv("BLUR_THRESHOLD", "120.0")) self.MIN_FACE_WIDTH = _getenv_int("MIN_FACE_WIDTH", 90)
self.MIN_CONFIDENCE = float(os.getenv("MIN_CONFIDENCE", "0.7")) self.MIN_FACE_COUNT = _getenv_int("MIN_FACE_COUNT", 3)
self.MAX_AUTO_IMAGES = int(os.getenv("MAX_AUTO_IMAGES", "80")) self.MERGE_DUPLICATE_PEOPLE = _getenv_bool("MERGE_DUPLICATE_PEOPLE", False)
self.QUALITY_REPLACEMENT = os.getenv("QUALITY_REPLACEMENT", "true").lower() in ("true", "1", "yes") self.BLUR_THRESHOLD = _getenv_float("BLUR_THRESHOLD", 120.0)
self.FRIGATE_SCORE_CEILING = float(os.getenv("FRIGATE_SCORE_CEILING", "0.0")) self.MIN_CONFIDENCE = _getenv_float("MIN_CONFIDENCE", 0.7)
self.ENABLE_FRIGATE_SCORES = os.getenv("ENABLE_FRIGATE_SCORES", "true").lower() in ("true", "1", "yes") self.MAX_AUTO_IMAGES = _getenv_int("MAX_AUTO_IMAGES", 5)
self.FACE_MARGIN = float(os.getenv("FACE_MARGIN", "0.15")) self.QUALITY_REPLACEMENT = _getenv_bool("QUALITY_REPLACEMENT", True)
self.USE_FULL_RESOLUTION = os.getenv("USE_FULL_RESOLUTION", "true").lower() in ("true", "1", "yes") self.FRIGATE_SCORE_CEILING = _getenv_optional_float("FRIGATE_SCORE_CEILING")
self.ENABLE_FACE_ALIGNMENT = os.getenv("ENABLE_FACE_ALIGNMENT", "true").lower() in ("true", "1", "yes") self.ENABLE_FRIGATE_SCORES = _getenv_bool("ENABLE_FRIGATE_SCORES", True)
self.ENABLE_CACHE = os.getenv("ENABLE_CACHE", "true").lower() in ("true", "1", "yes") self.FACE_MARGIN = _getenv_float("FACE_MARGIN", 0.15)
self.CACHE_DIR = os.getenv("CACHE_DIR", ".if_cache") self.USE_FULL_RESOLUTION = _getenv_bool("USE_FULL_RESOLUTION", True)
self.ENABLE_FACE_ALIGNMENT = _getenv_bool("ENABLE_FACE_ALIGNMENT", True)
self.ENABLE_CACHE = _getenv_bool("ENABLE_CACHE", True)
_data_dir = os.getenv("DATA_DIR")
_cache_dir_legacy = os.getenv("CACHE_DIR")
if _data_dir:
self.DATA_DIR = _data_dir
elif _cache_dir_legacy:
logging.warning(
"CACHE_DIR is deprecated — rename it to DATA_DIR in your .env or compose.yml"
)
self.DATA_DIR = _cache_dir_legacy
else:
self.DATA_DIR = "data"
# Fall back to config file for non-sensitive values (API_KEY not stored here) # Fall back to config file when the env var is absent or blank — a blank
if CONFIG_FILE.exists(): # IMMICH_URL= placeholder in .env should not override the config file.
# Prefer DATA_DIR/.immich_config.json (volume-safe in Docker) and fall back
# to the legacy CWD path so existing installations continue to work.
_data_cfg = Path(self.DATA_DIR) / ".immich_config.json"
_data_cfg_exists = _data_cfg.exists()
if _data_cfg_exists and _LEGACY_CONFIG_FILE.exists():
logging.warning(
"Two config files found: %s and %s — using %s. Remove the legacy file to silence this.",
_data_cfg,
_LEGACY_CONFIG_FILE,
_data_cfg,
)
config_file = _data_cfg if _data_cfg_exists else _LEGACY_CONFIG_FILE
# _data_cfg_exists already confirmed the primary path — avoid re-stat.
# The short-circuit means the legacy path is stat'd at most once here.
if _data_cfg_exists or config_file.exists():
try: try:
data = json.loads(CONFIG_FILE.read_text()) data = json.loads(config_file.read_text())
self.IMMICH_URL = self.IMMICH_URL or data.get("IMMICH_URL") if not self.IMMICH_URL:
self.IMMICH_URL = data.get("IMMICH_URL")
if not os.getenv("OUTPUT_DIR"): if not os.getenv("OUTPUT_DIR"):
self.OUTPUT_DIR = data.get("OUTPUT_DIR", self.OUTPUT_DIR) self.OUTPUT_DIR = data.get("OUTPUT_DIR", self.OUTPUT_DIR)
except (json.JSONDecodeError, OSError) as e: except (json.JSONDecodeError, OSError) as e:
logging.warning(f"Failed to load config file: {e}") logging.warning("Failed to load config file: %s", e)
@classmethod @classmethod
def reset(cls) -> None: def reset(cls) -> None:
"""Reset the singleton — mainly useful for testing or delayed env setup.""" """Reset the singleton — mainly useful for testing or delayed env setup."""
if cls._instance is not None:
cls._instance.__dict__.clear()
cls._instance = None cls._instance = None
def save(self) -> None: def save(self) -> None:
@@ -94,9 +193,13 @@ class _Config:
API_KEY is intentionally excluded — store it in .env or as an API_KEY is intentionally excluded — store it in .env or as an
environment variable instead of a plain-text config file. environment variable instead of a plain-text config file.
Writes to DATA_DIR/.immich_config.json so the file survives container
restarts when DATA_DIR is a mounted volume.
""" """
config_file = Path(self.DATA_DIR) / ".immich_config.json"
try: try:
CONFIG_FILE.write_text( Path(self.DATA_DIR).mkdir(parents=True, exist_ok=True)
config_file.write_text(
json.dumps( json.dumps(
{ {
"IMMICH_URL": self.IMMICH_URL, "IMMICH_URL": self.IMMICH_URL,
@@ -105,9 +208,9 @@ class _Config:
indent=2, indent=2,
) )
) )
logging.info(f"Configuration saved to {CONFIG_FILE}") logging.info("Configuration saved to %s", config_file)
except OSError as e: except OSError as e:
logging.error(f"Failed to save config: {e}") logging.error("Failed to save config: %s", e)
def interactive_setup(self) -> None: def interactive_setup(self) -> None:
"""Prompt user for missing configuration.""" """Prompt user for missing configuration."""
@@ -131,52 +234,10 @@ class _Config:
raise ValueError("Missing Immich URL or API Key.") raise ValueError("Missing Immich URL or API Key.")
# Singleton instance — use a lazy property pattern to avoid import-time side effects # Module-level singleton — lazy: no I/O until first attribute access.
# when env vars aren't yet set. Call Config.instance() or just access attributes on Config = _Config()
# the module-level `Config` (which delegates to the singleton).
class _ConfigAccessor:
"""Lazy accessor that defers singleton creation until first attribute access.
This avoids reading .env and config files at import time, so environment
variables set after importing the module are properly picked up.
"""
def __getattr__(self, name: str):
return getattr(_Config(), name)
def __setattr__(self, name: str, value):
if name.startswith("_"):
super().__setattr__(name, value)
else:
setattr(_Config(), name, value)
def reset(self) -> None:
"""Reset the underlying singleton."""
_Config.reset()
def interactive_setup(self) -> None:
"""Delegate to the singleton."""
_Config().interactive_setup()
def validate(self) -> None:
"""Delegate to the singleton."""
_Config().validate()
def save(self) -> None:
"""Delegate to the singleton."""
_Config().save()
Config = _ConfigAccessor()
class ConfigManager:
@staticmethod
def get() -> _Config:
return _Config()
def get_headers() -> dict[str, str]: def get_headers() -> dict[str, str]:
"""Return HTTP headers for Immich API requests.""" """Return HTTP headers for Immich API requests."""
return {"x-api-key": Config.API_KEY or "", "Accept": "application/json"} return {"x-api-key": Config.API_KEY or "", "Accept": "application/json"}
+143 -92
View File
@@ -5,11 +5,12 @@ Selection pipeline:
1. Concurrent thumbnail download 1. Concurrent thumbnail download
2. Quality filtering (blur, IR, exposure, confidence, face size) 2. Quality filtering (blur, IR, exposure, confidence, face size)
3. Face crop extraction (embed person's face, not full image) 3. Face crop extraction (embed person's face, not full image)
4. Embedding computation (InsightFace or SigLIP) 4. Embedding computation (InsightFace)
5. Cluster-aware selection (K-Medoids + FPS with hard example weighting) 5. Cluster-aware selection (K-Medoids + FPS with hard example weighting)
""" """
import logging import logging
from concurrent.futures import ThreadPoolExecutor, as_completed
from io import BytesIO from io import BytesIO
import numpy as np import numpy as np
@@ -22,15 +23,24 @@ from .quality import assess_quality
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# Candidate pool: cap at _POOL_CAP assets, but take at least _POOL_SCALE × the
# requested limit so small limits don't artificially narrow the search space.
_POOL_CAP = 3000
_POOL_SCALE = 20
# Embedding batch size: bounds decoded thumbnails in memory.
# At ~3-8 MB each, 32 images ≈ 100–250 MB peak — safe in a 4 GB container.
_EMBEDDING_BATCH_SIZE = 32
def select_diverse_assets( def select_diverse_assets(
assets: list, assets: list,
limit: int | str, limit: int | str,
entity_name: str, entity_name: str,
selection_mode: str = "smart", selection_mode: str = "smart",
entity_type: str = "face",
person_id: str | None = None, person_id: str | None = None,
progress_callback=None, progress_callback=None,
fetch_fn=None,
) -> list: ) -> list:
""" """
Select diverse assets using cluster-aware FPS or time spread. Select diverse assets using cluster-aware FPS or time spread.
@@ -38,31 +48,31 @@ def select_diverse_assets(
Args: Args:
assets: List of asset dicts from Immich API assets: List of asset dicts from Immich API
limit: Number to select, or "auto" for dynamic selection limit: Number to select, or "auto" for dynamic selection
entity_name: Name of the person/object for logging entity_name: Name of the person for logging
selection_mode: 'smart' (embedding-based) or 'time' (time spread) selection_mode: 'smart' (embedding-based) or 'time' (time spread)
entity_type: 'face' or 'object' - determines embedding model
progress_callback: Optional callback(current, total) for progress progress_callback: Optional callback(current, total) for progress
fetch_fn: Optional callable(asset_id) -> Image | None; defaults to
_fetch_thumbnail. Injected for testability.
Returns: Returns:
List of selected assets List of selected assets
""" """
# Fast path: fewer assets than limit # Fast path: fewer assets than limit — sort for consistent ordering with other paths
if limit != "auto" and len(assets) <= limit: if limit != "auto" and len(assets) <= limit:
return assets return sorted(assets, key=lambda x: x.get("fileCreatedAt", ""))
# Sort by creation time # Sort by creation time
assets = sorted(assets, key=lambda x: x.get("fileCreatedAt", "")) assets = sorted(assets, key=lambda x: x.get("fileCreatedAt", ""))
if selection_mode != "smart" or not is_embedding_available(entity_type): if selection_mode != "smart" or not is_embedding_available():
if selection_mode == "smart": if selection_mode == "smart":
model_name = "InsightFace" if entity_type == "face" else "SigLIP" logger.warning("InsightFace unavailable. Falling back to time spread.")
logger.warning(f"{model_name} unavailable. Falling back to time spread.")
return _select_time_spread(assets, limit) return _select_time_spread(assets, limit)
try: try:
return _select_by_embedding(assets, limit, entity_type, person_id, progress_callback) return _select_by_embedding(assets, limit, person_id, progress_callback, fetch_fn=fetch_fn)
except Exception as e: except Exception as e:
logger.error(f"Smart Diversity failed: {e}. Falling back to time spread.") logger.error("Smart Diversity failed: %s. Falling back to time spread.", e)
return _select_time_spread(assets, limit) return _select_time_spread(assets, limit)
@@ -171,6 +181,28 @@ def _crop_face_from_thumbnail(
return crop return crop
def _scale_bbox_to_thumbnail(
bbox: tuple[float, float, float, float],
img: Image.Image,
asset: dict,
person_id: str | None = None,
) -> tuple[float, float, float, float]:
"""Scale a face bbox from original detection-image space to thumbnail-pixel space."""
x1, y1, x2, y2 = bbox
img_w, img_h = img.size
for person in asset.get("people", []):
if person_id and person.get("id") != person_id:
continue
faces = person.get("faces", [])
if faces:
meta_w = faces[0].get("imageWidth") or 0
meta_h = faces[0].get("imageHeight") or 0
scale_x = img_w / meta_w if meta_w else 1.0
scale_y = img_h / meta_h if meta_h else 1.0
return (x1 * scale_x, y1 * scale_y, x2 * scale_x, y2 * scale_y)
return bbox
# ============================================================================= # =============================================================================
# Embedding Collection # Embedding Collection
# ============================================================================= # =============================================================================
@@ -179,22 +211,21 @@ def _crop_face_from_thumbnail(
def _select_by_embedding( def _select_by_embedding(
assets: list, assets: list,
limit: int | str, limit: int | str,
entity_type: str,
person_id: str | None = None, person_id: str | None = None,
progress_callback=None, progress_callback=None,
fetch_fn=None,
) -> list: ) -> list:
"""Select assets using embedding-based cluster-aware FPS. """Select assets using embedding-based cluster-aware FPS.
Pipeline: Pipeline:
1. Concurrent thumbnail download 1. Concurrent thumbnail download
2. Quality filtering 2. Quality filtering
3. Face crop extraction (face mode only) 3. Face crop extraction
4. Embedding computation 4. Embedding computation
5. Cluster-aware selection with hard example weighting 5. Cluster-aware selection with hard example weighting
""" """
# Determine candidate pool (cap at 3000 for performance)
effective_limit = 30 if limit == "auto" else limit effective_limit = 30 if limit == "auto" else limit
pool_size = min(3000, max(effective_limit * 20, len(assets))) pool_size = min(_POOL_CAP, max(effective_limit * _POOL_SCALE, len(assets)))
# Subsample if needed (evenly distributed in time) # Subsample if needed (evenly distributed in time)
if len(assets) > pool_size: if len(assets) > pool_size:
@@ -207,20 +238,25 @@ def _select_by_embedding(
# Process in bounded batches so at most _BATCH decoded images live in RAM # Process in bounded batches so at most _BATCH decoded images live in RAM
# at once. With 472 candidates each thumbnail is ~3-8 MB decoded; loading # at once. With 472 candidates each thumbnail is ~3-8 MB decoded; loading
# all at once easily exhausts a 4 GB container limit on CPU. # all at once easily exhausts a 4 GB container limit on CPU.
from concurrent.futures import ThreadPoolExecutor, as_completed # LIMITATION — thumbnail-resolution embeddings drive full-res crop selection:
# diversity selection runs InsightFace on Immich preview thumbnails (~720p)
_BATCH = 32 # to avoid downloading full-res for every candidate, but the training crop
# comes from the full-resolution original. Embeddings from thumbnails are
# representative in practice, but heavy JPEG compression on a preview could
# produce a subtly different embedding than the full-res version. For most
# libraries this is negligible; it matters if Immich preview quality is low.
_fetch = fetch_fn or _fetch_thumbnail
embeddings, valid_candidates, confidence_scores = [], [], [] embeddings, valid_candidates, confidence_scores = [], [], []
quality_filtered = 0 quality_filtered = 0
processed = 0 processed = 0
for batch_start in range(0, len(candidates), _BATCH): for batch_start in range(0, len(candidates), _EMBEDDING_BATCH_SIZE):
batch = candidates[batch_start : batch_start + _BATCH] batch = candidates[batch_start : batch_start + _EMBEDDING_BATCH_SIZE]
# Download this batch concurrently # Download this batch concurrently
batch_images: dict[str, Image.Image] = {} batch_images: dict[str, Image.Image] = {}
with ThreadPoolExecutor(max_workers=min(8, len(batch))) as pool: with ThreadPoolExecutor(max_workers=min(8, len(batch))) as pool:
futures = {pool.submit(_fetch_thumbnail, a["id"]): a for a in batch} futures = {pool.submit(_fetch, a["id"]): a for a in batch}
for future in as_completed(futures): for future in as_completed(futures):
asset = futures[future] asset = futures[future]
try: try:
@@ -228,7 +264,7 @@ def _select_by_embedding(
if img is not None: if img is not None:
batch_images[asset["id"]] = img batch_images[asset["id"]] = img
except Exception as e: except Exception as e:
logger.debug(f"Failed to fetch thumbnail for {asset['id']}: {e}") logger.debug("Failed to fetch thumbnail for %s: %s", asset["id"], e)
continue continue
# Process each image; batch_images goes out of scope after this loop, # Process each image; batch_images goes out of scope after this loop,
@@ -243,42 +279,46 @@ def _select_by_embedding(
confidence = _get_face_confidence(asset, person_id=person_id) confidence = _get_face_confidence(asset, person_id=person_id)
if entity_type == "face": face_bbox = _get_face_bbox(asset, person_id=person_id)
face_bbox = _get_face_bbox(asset, person_id=person_id) thumbnail_bbox = (
quality = assess_quality( _scale_bbox_to_thumbnail(face_bbox, img, asset, person_id)
img, if face_bbox is not None else None
face_bbox=face_bbox, )
confidence=confidence, quality = assess_quality(
blur_threshold=Config.BLUR_THRESHOLD, img,
min_face_px=Config.MIN_FACE_WIDTH, face_bbox=thumbnail_bbox,
min_confidence=Config.MIN_CONFIDENCE, confidence=confidence,
) blur_threshold=Config.BLUR_THRESHOLD,
if not quality.passed: min_face_px=Config.MIN_FACE_WIDTH,
quality_filtered += 1 min_confidence=Config.MIN_CONFIDENCE,
logger.debug(f"Quality filtered {asset['id']}: {quality.reason}") )
continue if not quality.passed:
quality_filtered += 1
logger.debug("Quality filtered %s: %s", asset["id"], quality.reason)
continue
asset["quality_score"] = quality.blur_score asset["quality_score"] = quality.blur_score
face_crop = _crop_face_from_thumbnail(img, asset, person_id=person_id) face_crop = _crop_face_from_thumbnail(img, asset, person_id=person_id)
embed_img = face_crop if face_crop is not None else img embed_img = face_crop if face_crop is not None else img
else:
embed_img = img
emb = get_embedding(embed_img, entity_type, asset_id=asset["id"]) emb = get_embedding(embed_img, asset_id=asset["id"])
if emb is not None: if emb is not None:
if np.linalg.norm(emb) < 1e-6:
logger.debug("Zero-norm embedding for asset %s, skipping", asset["id"])
continue
embeddings.append(emb) embeddings.append(emb)
valid_candidates.append(asset) valid_candidates.append(asset)
confidence_scores.append(confidence) confidence_scores.append(confidence)
if quality_filtered > 0: if quality_filtered > 0:
logger.info(f"Quality filtering removed {quality_filtered} images.") logger.info("Quality filtering removed %s images.", quality_filtered)
if not embeddings: if not embeddings:
logger.warning("No valid embeddings found. Falling back to time spread.") logger.warning("No valid embeddings found. Falling back to time spread.")
return _select_time_spread(assets, limit) return _select_time_spread(assets, limit)
if limit != "auto" and len(valid_candidates) < limit: if limit != "auto" and len(valid_candidates) < limit:
logger.warning(f"Only {len(valid_candidates)} valid embeddings. Returning all.") logger.warning("Only %s valid embeddings. Returning all.", len(valid_candidates))
return valid_candidates return valid_candidates
# --- Phase 5: Near-duplicate removal --- # --- Phase 5: Near-duplicate removal ---
@@ -290,12 +330,16 @@ def _select_by_embedding(
embeddings, valid_candidates, confidence_scores embeddings, valid_candidates, confidence_scores
) )
# Re-check after dedup: pool may have shrunk below limit
if limit != "auto" and len(valid_candidates) < limit:
logger.warning("Only %s embeddings after near-duplicate removal. Returning all.", len(valid_candidates))
return valid_candidates
# --- Phase 6: Cluster-aware selection --- # --- Phase 6: Cluster-aware selection ---
return _cluster_aware_selection( return _cluster_aware_selection(
embeddings, embeddings,
valid_candidates, valid_candidates,
limit, limit,
entity_type=entity_type,
confidence_scores=confidence_scores, confidence_scores=confidence_scores,
) )
@@ -326,25 +370,29 @@ def _dedup_embeddings(
norms = np.linalg.norm(emb_matrix, axis=1, keepdims=True) norms = np.linalg.norm(emb_matrix, axis=1, keepdims=True)
emb_normed = emb_matrix / np.maximum(norms, 1e-8) emb_normed = emb_matrix / np.maximum(norms, 1e-8)
# Sort by quality descending so the best image in each near-duplicate group wins # Sort by quality descending so the best image in each near-duplicate group wins.
quality_scores = [c.get("quality_score") or 0.0 for c in candidates] # Use explicit None check so a legitimate quality_score=0.0 isn't treated as missing.
quality_scores = [qs if (qs := c.get("quality_score")) is not None else 0.0 for c in candidates]
order = sorted(range(len(candidates)), key=lambda i: quality_scores[i], reverse=True) order = sorted(range(len(candidates)), key=lambda i: quality_scores[i], reverse=True)
kept_indices = [] kept_indices = []
kept_normed = [] # Pre-allocate a max-size buffer and fill row-by-row — eliminates the O(K²)
# copy overhead from vstack-on-keep while keeping identical arithmetic.
kept_buf = np.empty((len(order), emb_normed.shape[1]), dtype=emb_normed.dtype)
n_kept = 0
for i in order: for i in order:
if kept_normed: if n_kept > 0:
kept_stack = np.vstack(kept_normed) sims = emb_normed[i] @ kept_buf[:n_kept].T
sims = emb_normed[i] @ kept_stack.T
if np.any(sims > 1 - _DEDUP_THRESHOLD): if np.any(sims > 1 - _DEDUP_THRESHOLD):
continue continue
kept_buf[n_kept] = emb_normed[i]
n_kept += 1
kept_indices.append(i) kept_indices.append(i)
kept_normed.append(emb_normed[i])
dropped = len(embeddings) - len(kept_indices) dropped = len(embeddings) - len(kept_indices)
if dropped: if dropped:
logger.info(f"Near-duplicate removal dropped {dropped} images (threshold {_DEDUP_THRESHOLD}).") logger.info("Near-duplicate removal dropped %s images (threshold %s).", dropped, _DEDUP_THRESHOLD)
return ( return (
[embeddings[i] for i in kept_indices], [embeddings[i] for i in kept_indices],
@@ -384,12 +432,13 @@ def _kmedoids(dist_matrix: np.ndarray, k: int, max_iter: int = 50) -> tuple[list
# Iterative swap step # Iterative swap step
medoids = list(medoids) medoids = list(medoids)
labels = np.argmin(dist_matrix[:, medoids], axis=1) labels = np.argmin(dist_matrix[:, medoids], axis=1)
cost = sum(dist_matrix[i, medoids[labels[i]]] for i in range(n)) cost = dist_matrix[np.arange(n), np.array(medoids)[labels]].sum()
for _ in range(max_iter): for _ in range(max_iter):
improved = False improved = False
# Try swapping each medoid with a random non-medoid # Try swapping each medoid with a random non-medoid
non_medoids = [i for i in range(n) if i not in medoids] medoid_set = set(medoids)
non_medoids = [i for i in range(n) if i not in medoid_set]
if not non_medoids: if not non_medoids:
break break
@@ -399,7 +448,7 @@ def _kmedoids(dist_matrix: np.ndarray, k: int, max_iter: int = 50) -> tuple[list
new_medoids = medoids.copy() new_medoids = medoids.copy()
new_medoids[m_idx] = cand new_medoids[m_idx] = cand
new_labels = np.argmin(dist_matrix[:, new_medoids], axis=1) new_labels = np.argmin(dist_matrix[:, new_medoids], axis=1)
new_cost = sum(dist_matrix[i, new_medoids[new_labels[i]]] for i in range(n)) new_cost = dist_matrix[np.arange(n), np.array(new_medoids)[new_labels]].sum()
if new_cost < cost: if new_cost < cost:
medoids = new_medoids medoids = new_medoids
labels = new_labels labels = new_labels
@@ -420,11 +469,11 @@ def _kmedoids(dist_matrix: np.ndarray, k: int, max_iter: int = 50) -> tuple[list
# ============================================================================= # =============================================================================
def _compute_adaptive_threshold(emb_normed: np.ndarray, entity_type: str) -> float: def _compute_adaptive_threshold(emb_normed: np.ndarray) -> float:
"""Compute adaptive FPS stop threshold based on actual embedding distribution. """Compute adaptive FPS stop threshold based on actual embedding distribution.
Instead of a hardcoded threshold, samples pairwise distances and sets Instead of a hardcoded threshold, samples pairwise distances and sets
the threshold as a fraction of the median pairwise distance. the threshold as 20% of the median pairwise distance.
""" """
n = len(emb_normed) n = len(emb_normed)
sample_size = min(200, n) sample_size = min(200, n)
@@ -432,22 +481,14 @@ def _compute_adaptive_threshold(emb_normed: np.ndarray, entity_type: str) -> flo
indices = rng.choice(n, sample_size, replace=False) if n > sample_size else np.arange(n) indices = rng.choice(n, sample_size, replace=False) if n > sample_size else np.arange(n)
sample = emb_normed[indices] sample = emb_normed[indices]
# Compute pairwise cosine distances for the sample
pairwise = 1 - sample @ sample.T pairwise = 1 - sample @ sample.T
upper_tri = pairwise[np.triu_indices(len(sample), k=1)] upper_tri = pairwise[np.triu_indices(len(sample), k=1)]
if len(upper_tri) == 0: if len(upper_tri) == 0:
return 0.05 return 0.05
median_dist = float(np.median(upper_tri)) median_dist = float(np.median(upper_tri))
threshold = max(0.05, median_dist * 0.20)
# Faces: 20% of median (tighter — want fewer, more distinct images) logger.debug("Adaptive threshold: %.4f (median_dist=%.4f)", threshold, median_dist)
# Objects: 10% of median (wider — want more diversity)
fraction = 0.20 if entity_type == "face" else 0.10
threshold = max(0.05, median_dist * fraction)
logger.debug(
f"Adaptive threshold: {threshold:.4f} "
f"(median_dist={median_dist:.4f}, fraction={fraction}, type={entity_type})"
)
return threshold return threshold
@@ -455,7 +496,6 @@ def _cluster_aware_selection(
embeddings: list, embeddings: list,
candidates: list, candidates: list,
limit: int | str, limit: int | str,
entity_type: str = "face",
confidence_scores: list | None = None, confidence_scores: list | None = None,
) -> list: ) -> list:
"""Two-stage selection: K-Medoids clustering → FPS with hard example weighting. """Two-stage selection: K-Medoids clustering → FPS with hard example weighting.
@@ -473,29 +513,36 @@ def _cluster_aware_selection(
norms = np.linalg.norm(emb_matrix, axis=1, keepdims=True) norms = np.linalg.norm(emb_matrix, axis=1, keepdims=True)
emb_normed = emb_matrix / np.maximum(norms, 1e-8) emb_normed = emb_matrix / np.maximum(norms, 1e-8)
# Build confidence weight array for hard example boosting # Build confidence weight array for hard example boosting.
# Default to 1.0 for faces with no confidence score: treat as high-confidence
# (no boost) rather than hard-example territory. A missing score field should
# not cause these images to beat genuinely high-confidence detections in FPS.
conf_array = np.ones(n) conf_array = np.ones(n)
if confidence_scores and entity_type == "face": if confidence_scores:
for i, c in enumerate(confidence_scores): for i, c in enumerate(confidence_scores):
if c is not None: if c is not None:
conf_array[i] = c conf_array[i] = c
# Compute adaptive threshold for auto mode # Compute adaptive threshold for auto mode
auto_threshold = _compute_adaptive_threshold(emb_normed, entity_type) if limit == "auto" else 0.0 auto_threshold = _compute_adaptive_threshold(emb_normed) if limit == "auto" else 0.0
target = Config.MAX_AUTO_IMAGES if limit == "auto" else limit target = Config.MAX_AUTO_IMAGES if limit == "auto" else limit
# Short-circuit: nothing to select
if limit != "auto" and target <= 0:
return []
# --- Stage 1: K-Medoids clustering --- # --- Stage 1: K-Medoids clustering ---
k = min(max(5, target // 4), max(1, n // 3), n) # e.g., 1-20 clusters # Cap k at target so we never seed more cluster representatives than requested.
logger.debug(f"Clustering {n} embeddings into {k} groups (K-Medoids)...") k = min(max(5, target // 4), max(1, n // 3), n, target) # e.g., 1-20 clusters
logger.debug("Clustering %s embeddings into %s groups (K-Medoids)...", n, k)
# Compute full cosine distance matrix # Compute full cosine distance matrix
dist_matrix = 1 - emb_normed @ emb_normed.T dist_matrix = 1 - emb_normed @ emb_normed.T
medoid_indices, cluster_labels = _kmedoids(dist_matrix, k) medoid_indices, cluster_labels = _kmedoids(dist_matrix, k)
selected = list(medoid_indices) selected = list(medoid_indices)
selected_set = set(selected)
logger.debug(f"Selected {len(selected)} cluster medoids as initial picks.") logger.debug("Selected %s cluster medoids as initial picks.", len(selected))
# --- Stage 2: FPS with hard example weighting --- # --- Stage 2: FPS with hard example weighting ---
min_dists = np.full(n, np.inf) min_dists = np.full(n, np.inf)
@@ -507,10 +554,11 @@ def _cluster_aware_selection(
for idx in selected: for idx in selected:
min_dists[idx] = -np.inf min_dists[idx] = -np.inf
# Hard example weighting: boost distance for low-confidence candidates.
# conf_array is constant after this point, so compute once outside the loop.
hard_weight = np.where(conf_array < 0.85, 1.0 + (0.85 - conf_array) * 2.0, 1.0)
while len(selected) < target: while len(selected) < target:
# Hard example weighting: boost distance for low-confidence candidates
# Confidence < 0.85 gets up to 1.5× distance boost
hard_weight = np.where(conf_array < 0.85, 1.0 + (0.85 - conf_array) * 2.0, 1.0)
weighted_dists = min_dists * hard_weight weighted_dists = min_dists * hard_weight
best_idx = int(np.argmax(weighted_dists)) best_idx = int(np.argmax(weighted_dists))
@@ -526,24 +574,27 @@ def _cluster_aware_selection(
break break
selected.append(best_idx) selected.append(best_idx)
selected_set.add(best_idx)
# Update min distances # Update min distances
dists_to_new = dist_matrix[best_idx] dists_to_new = dist_matrix[best_idx]
min_dists = np.minimum(min_dists, dists_to_new) min_dists = np.minimum(min_dists, dists_to_new)
min_dists[best_idx] = -np.inf min_dists[best_idx] = -np.inf
# Log hard example stats hard_count = sum(
if entity_type == "face": 1 for i in selected
selected_conf = [conf_array[i] for i in selected if conf_array[i] < 1.0] if confidence_scores
hard_count = sum(1 for c in selected_conf if c < 0.85) and i < len(confidence_scores)
logger.info( and confidence_scores[i] is not None
f"Selection complete: {len(selected)} images " f"({hard_count} hard examples with confidence < 0.85)." and confidence_scores[i] < 0.85
) )
else: logger.info("Selection complete: %s images (%s hard examples with confidence < 0.85).", len(selected), hard_count)
logger.info(f"Selection complete: {len(selected)} diverse images.")
return [candidates[i] for i in selected] # Slice to target: the while loop enforces this for non-auto mode, but
# guard here too in case the medoid seed already exceeded target (small target).
result = [candidates[i] for i in selected]
if limit != "auto":
result = result[:target]
return result
# ============================================================================= # =============================================================================
@@ -556,10 +607,10 @@ def _select_time_spread(assets: list, limit: int | str) -> list:
if limit == "auto": if limit == "auto":
limit = 30 limit = 30
logger.info(f"Selecting {limit} images using time spread.") logger.info("Selecting %s images using time spread.", limit)
if len(assets) <= limit: if len(assets) <= limit:
return assets return assets
indices = np.linspace(0, len(assets) - 1, limit, dtype=int) indices = np.linspace(0, len(assets) - 1, limit, dtype=int)
return [assets[i] for i in np.unique(indices)] return [assets[i] for i in indices]
+82 -204
View File
@@ -1,8 +1,7 @@
""" """
Unified embedding interface for faces and objects. Embedding interface for face diversity selection.
- Faces: InsightFace (ArcFace/Buffalo_L) — or reuse from Immich - Faces: InsightFace (ArcFace/Buffalo_L) — or reuse from Immich
- Objects: SigLIP (Vision Transformer via transformers)
- Caching: Disk-based cache avoids recomputation on reruns - Caching: Disk-based cache avoids recomputation on reruns
""" """
@@ -19,6 +18,7 @@ import numpy as np
from PIL import Image from PIL import Image
from .cache import get_cache from .cache import get_cache
from .config import _getenv_bool
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -26,33 +26,57 @@ logger = logging.getLogger(__name__)
@contextmanager @contextmanager
def _suppress_output(): def _suppress_output():
"""Suppress stdout/stderr at the file-descriptor level, silencing C extension noise.""" """Suppress stdout/stderr at the file-descriptor level, silencing C extension noise."""
devnull_fd = os.open(os.devnull, os.O_WRONLY) devnull_fd = None
saved_out, saved_err = os.dup(1), os.dup(2) saved_out = None
saved_err = None
try: try:
devnull_fd = os.open(os.devnull, os.O_WRONLY)
saved_out = os.dup(1)
saved_err = os.dup(2)
os.dup2(devnull_fd, 1) os.dup2(devnull_fd, 1)
os.dup2(devnull_fd, 2) os.dup2(devnull_fd, 2)
yield yield
finally: finally:
try: # Each block is a separate sequential statement. A BaseException (e.g.
os.dup2(saved_out, 1) # KeyboardInterrupt) raised inside block N would propagate past blocks N+1
finally: # and N+2, leaving saved_err or devnull_fd unclosed. In CPython, KI is
os.dup2(saved_err, 2) # delivered between bytecodes, not mid-syscall; os.dup2 is a single C call
os.close(devnull_fd) # and completes atomically, so this race is not realistically triggerable.
os.close(saved_out) if saved_out is not None:
os.close(saved_err) try:
os.dup2(saved_out, 1)
except OSError as e:
logger.debug("_suppress_output: failed to restore stdout fd: %s", e)
finally:
try:
os.close(saved_out)
except OSError:
pass
if saved_err is not None:
try:
os.dup2(saved_err, 2)
except OSError as e:
logger.debug("_suppress_output: failed to restore stderr fd: %s", e)
finally:
try:
os.close(saved_err)
except OSError:
pass
if devnull_fd is not None:
try:
os.close(devnull_fd)
except OSError:
pass
# Lazy-loaded singletons # Lazy-loaded singleton
_insightface_app = None _insightface_app = None
_insightface_loaded = False _insightface_loaded = False
_siglip_model = None
_siglip_processor = None
_siglip_loaded = False
def _is_force_cpu() -> bool: def _is_force_cpu() -> bool:
"""Check if CPU mode is forced via environment variable.""" """Check if CPU mode is forced via environment variable."""
return os.getenv("FORCE_CPU", "").lower() in ("true", "1", "yes") return _getenv_bool("FORCE_CPU", False)
def _preload_cuda_libs() -> None: def _preload_cuda_libs() -> None:
@@ -70,7 +94,7 @@ def _preload_cuda_libs() -> None:
else: else:
logger.debug("onnxruntime.preload_dlls() not available (ORT < 1.21)") logger.debug("onnxruntime.preload_dlls() not available (ORT < 1.21)")
except Exception as e: except Exception as e:
logger.warning(f"Failed to preload CUDA/cuDNN DLLs: {e}") logger.warning("Failed to preload CUDA/cuDNN DLLs: %s", e)
# ============================================================================= # =============================================================================
@@ -83,7 +107,6 @@ def get_insightface_app():
global _insightface_app, _insightface_loaded global _insightface_app, _insightface_loaded
if _insightface_loaded: if _insightface_loaded:
return _insightface_app return _insightface_app
_insightface_loaded = True
ctx_id = -1 ctx_id = -1
insightface_home = os.environ.get("INSIGHTFACE_HOME", os.path.expanduser("~/.insightface")) insightface_home = os.environ.get("INSIGHTFACE_HOME", os.path.expanduser("~/.insightface"))
@@ -104,7 +127,7 @@ def get_insightface_app():
# Get providers, excluding TensorRT to avoid noisy errors # Get providers, excluding TensorRT to avoid noisy errors
providers = [p for p in ort.get_available_providers() if p != "TensorrtExecutionProvider"] providers = [p for p in ort.get_available_providers() if p != "TensorrtExecutionProvider"]
logger.debug(f"ONNX providers available: {providers}") logger.debug("ONNX providers available: %s", providers)
gpu_providers = { gpu_providers = {
"CUDAExecutionProvider", "CUDAExecutionProvider",
@@ -125,7 +148,7 @@ def get_insightface_app():
if p == "OpenVINOExecutionProvider" else p if p == "OpenVINOExecutionProvider" else p
for p in providers for p in providers
] ]
logger.debug(f"OpenVINO EP: device_type={openvino_device}") logger.debug("OpenVINO EP: device_type=%s", openvino_device)
if not has_gpu_provider and not _is_force_cpu(): if not has_gpu_provider and not _is_force_cpu():
logger.warning( logger.warning(
@@ -140,21 +163,23 @@ def get_insightface_app():
device_str = f"OpenVINO ({os.getenv('OPENVINO_DEVICE', 'CPU')})" device_str = f"OpenVINO ({os.getenv('OPENVINO_DEVICE', 'CPU')})"
else: else:
device_str = "GPU" device_str = "GPU"
logger.info(f"InsightFace Buffalo_L: loading into memory on {device_str}...") logger.info("InsightFace Buffalo_L: loading into memory on %s...", device_str)
t0 = time.time() t0 = time.time()
with _suppress_output(): with _suppress_output():
_insightface_app = FaceAnalysis(name="buffalo_l", root=insightface_home, providers=providers) _insightface_app = FaceAnalysis(name="buffalo_l", root=insightface_home, providers=providers)
_insightface_app.prepare(ctx_id=ctx_id, det_size=(640, 640)) _insightface_app.prepare(ctx_id=ctx_id, det_size=(640, 640))
logger.info(f"InsightFace Buffalo_L: ready on {device_str} ({time.time() - t0:.1f}s)") logger.info("InsightFace Buffalo_L: ready on %s (%.1fs)", device_str, time.time() - t0)
_insightface_loaded = True
return _insightface_app return _insightface_app
except ImportError: except ImportError:
logger.error("InsightFace not installed!") logger.error("InsightFace not installed!")
_insightface_loaded = True
return None return None
except Exception as e: except Exception as e:
logger.error(f"Failed to load InsightFace: {e}") logger.error("Failed to load InsightFace: %s", e)
if ctx_id == 0: if ctx_id == 0:
logger.warning("InsightFace GPU load failed — retrying on CPU...") logger.warning("InsightFace GPU load failed — retrying on CPU...")
try: try:
@@ -168,10 +193,12 @@ def get_insightface_app():
providers=["CPUExecutionProvider"], providers=["CPUExecutionProvider"],
) )
_insightface_app.prepare(ctx_id=-1, det_size=(640, 640)) _insightface_app.prepare(ctx_id=-1, det_size=(640, 640))
logger.info(f"InsightFace Buffalo_L: ready on CPU (fallback, {time.time() - t0:.1f}s)") logger.info("InsightFace Buffalo_L: ready on CPU (fallback, %.1fs)", time.time() - t0)
_insightface_loaded = True
return _insightface_app return _insightface_app
except Exception as ex: except Exception as ex:
logger.error(f"InsightFace CPU fallback failed: {ex}") logger.error("InsightFace CPU fallback failed: %s", ex)
_insightface_loaded = True
return None return None
@@ -182,8 +209,9 @@ def get_face_embedding(img_pil: Image.Image) -> np.ndarray | None:
return None return None
try: try:
# InsightFace expects BGR cv2 image # InsightFace expects BGR cv2 image; normalise mode first so RGBA/grayscale don't
img_bgr = cv2.cvtColor(np.asarray(img_pil), cv2.COLOR_RGB2BGR) # raise a channel-count error inside cvtColor.
img_bgr = cv2.cvtColor(np.asarray(img_pil.convert("RGB")), cv2.COLOR_RGB2BGR)
# Suppress scikit-image FutureWarning from InsightFace's face_align.py # Suppress scikit-image FutureWarning from InsightFace's face_align.py
with warnings.catch_warnings(): with warnings.catch_warnings():
@@ -193,178 +221,47 @@ def get_face_embedding(img_pil: Image.Image) -> np.ndarray | None:
if not faces: if not faces:
return None return None
# Return embedding of largest face # Return embedding of the face nearest the crop centre; a large margin can pull
largest = max(faces, key=lambda f: (f.bbox[2] - f.bbox[0]) * (f.bbox[3] - f.bbox[1])) # a bigger neighbouring face into frame, and max-by-area would pick the wrong person.
return largest.embedding cx, cy = img_pil.width / 2, img_pil.height / 2
nearest = min(
faces,
key=lambda f: ((f.bbox[0] + f.bbox[2]) / 2 - cx) ** 2 + ((f.bbox[1] + f.bbox[3]) / 2 - cy) ** 2,
)
return nearest.embedding
except Exception as e: except Exception as e:
logger.error(f"Error getting face embedding: {e}") logger.error("Error getting face embedding: %s", e)
return None return None
# ============================================================================= # =============================================================================
# SigLIP (Objects) # Embedding Interface with Caching
# =============================================================================
def get_siglip_model():
"""Singleton for SigLIP model and processor with GPU auto-detection."""
global _siglip_model, _siglip_processor, _siglip_loaded
if _siglip_loaded:
return _siglip_model, _siglip_processor
_siglip_loaded = True
try:
import warnings
import torch
from transformers import AutoImageProcessor, SiglipVisionModel
model_name = "google/siglip-base-patch16-224"
# Disk cache check — path derived from model_name using HuggingFace's slug convention
hf_home = os.environ.get("HF_HOME", os.path.join(os.path.expanduser("~"), ".cache", "huggingface"))
cache_slug = "models--" + model_name.replace("/", "--")
model_cache = Path(hf_home) / "hub" / cache_slug
if model_cache.exists() and any(model_cache.iterdir()):
logger.info(f"SigLIP {model_name}: found in model cache")
else:
logger.info(f"SigLIP {model_name}: not cached — downloading now (~380 MB)")
logger.info(f"SigLIP {model_name}: loading into memory...")
t0 = time.time()
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=FutureWarning)
warnings.filterwarnings("ignore", message=".*use_fast.*")
_siglip_processor = AutoImageProcessor.from_pretrained(model_name, use_fast=True)
_siglip_model = SiglipVisionModel.from_pretrained(model_name)
_siglip_model.eval()
# Move to GPU if available (ROCm builds expose torch.cuda.is_available() == True)
if not _is_force_cpu():
if torch.cuda.is_available():
_siglip_model = _siglip_model.cuda()
device_name = "CUDA GPU"
elif hasattr(torch, "xpu") and torch.xpu.is_available():
_siglip_model = _siglip_model.to("xpu")
device_name = "Intel XPU"
elif hasattr(torch.backends, "mps") and torch.backends.mps.is_available():
_siglip_model = _siglip_model.to("mps")
device_name = "Apple MPS"
else:
device_name = "CPU"
else:
device_name = "CPU (FORCE_CPU)"
logger.info(f"SigLIP {model_name}: ready on {device_name} ({time.time() - t0:.1f}s)")
return _siglip_model, _siglip_processor
except ImportError as e:
logger.error(f"transformers/torch not installed: {e}")
return None, None
except Exception as e:
logger.error(f"Failed to load SigLIP: {e}")
return None, None
def get_object_embedding(img_pil: Image.Image) -> np.ndarray | None:
"""Get 768-dim SigLIP embedding for an image."""
model, processor = get_siglip_model()
if model is None:
return None
try:
import torch
inputs = processor(images=img_pil, return_tensors="pt")
device = next(model.parameters()).device
inputs = {k: v.to(device) for k, v in inputs.items()}
with torch.no_grad():
outputs = model(**inputs)
return outputs.pooler_output.squeeze().cpu().numpy()
except Exception as e:
logger.error(f"Error getting object embedding: {e}")
return None
def get_object_embeddings_batch(images: list[Image.Image]) -> list[np.ndarray | None]:
"""Get SigLIP embeddings for a batch of images (GPU-efficient)."""
model, processor = get_siglip_model()
if model is None:
return [None] * len(images)
try:
import torch
inputs = processor(images=images, return_tensors="pt", padding=True)
device = next(model.parameters()).device
inputs = {k: v.to(device) for k, v in inputs.items()}
with torch.no_grad():
outputs = model(**inputs)
embeddings = outputs.pooler_output.cpu().numpy()
return [embeddings[i] for i in range(len(embeddings))]
except Exception as e:
logger.error(f"Error in batch embedding: {e}")
# Fall back to individual computation
return [get_object_embedding(img) for img in images]
# =============================================================================
# Unified Interface with Caching
# ============================================================================= # =============================================================================
def get_embedding( def get_embedding(
img_pil: Image.Image, img_pil: Image.Image,
entity_type: str = "face",
asset_id: str | None = None, asset_id: str | None = None,
immich_embedding: np.ndarray | None = None,
) -> np.ndarray | None: ) -> np.ndarray | None:
"""Get embedding for an image based on entity type. """Get embedding for a face image.
Priority: Checks disk cache first (if enabled and asset_id provided),
1. Pre-fetched Immich embedding (if provided) then falls back to local InsightFace computation.
2. Disk cache (if enabled and asset_id provided)
3. Local model computation (InsightFace or SigLIP)
Args:
img_pil: The image to embed
entity_type: 'face' or 'object'
asset_id: Optional asset ID for cache lookup
immich_embedding: Optional pre-fetched embedding from Immich API
""" """
from .config import Config from .config import Config
use_cache = Config.ENABLE_CACHE and asset_id is not None use_cache = Config.ENABLE_CACHE and asset_id is not None
cache = get_cache(Config.CACHE_DIR) if use_cache else None cache = get_cache(Config.DATA_DIR) if use_cache else None
# Use a single consistent cache key per model so lookups and stores always match.
# "immich" was previously used as the face key on the lookup path but "insightface"
# on the store path — meaning the cache was never hit for locally-computed embeddings.
cache_key = "insightface" if entity_type == "face" else "siglip"
# 1. Use Immich embedding if provided
if immich_embedding is not None:
if cache:
cache.put(asset_id, immich_embedding, cache_key)
return immich_embedding
# 2. Check disk cache
if cache: if cache:
cached = cache.get(asset_id, cache_key) cached = cache.get(asset_id, "insightface")
if cached is not None: if cached is not None:
return cached return cached
# 3. Compute locally emb = get_face_embedding(img_pil)
if entity_type == "face":
emb = get_face_embedding(img_pil)
else:
emb = get_object_embedding(img_pil)
if emb is not None and cache: if emb is not None and cache:
cache.put(asset_id, emb, cache_key) cache.put(asset_id, emb, "insightface")
return emb return emb
@@ -372,42 +269,23 @@ def get_embedding(
def _is_module_available(module_name: str) -> bool: def _is_module_available(module_name: str) -> bool:
"""Check if a Python module is importable without importing it fully.""" """Check if a Python module is importable without importing it fully."""
try: try:
importlib.util.find_spec(module_name) return importlib.util.find_spec(module_name) is not None
return True
except (ModuleNotFoundError, ValueError): except (ModuleNotFoundError, ValueError):
return False return False
def is_embedding_available(entity_type: str = "face", *, load: bool = False) -> bool: def is_embedding_available(*, load: bool = False) -> bool:
"""Check if embedding model is available for the given entity type. """Check if InsightFace is available.
By default this performs a lightweight import-check only (no model loading). By default this performs a lightweight import-check only (no model loading).
Pass ``load=True`` to actually load the model (expensive, hundreds of MB). Pass ``load=True`` to actually load the model (expensive, ~300 MB).
Args:
entity_type: 'face' or 'object'
load: If True, fully load the model to verify. If False (default),
only check that the required packages are importable.
""" """
if load: if load:
if entity_type == "face":
return get_insightface_app() is not None
model, _ = get_siglip_model()
return model is not None
# Lightweight check: just verify the packages are importable
if entity_type == "face":
return _is_module_available("insightface") and _is_module_available("onnxruntime")
return _is_module_available("transformers") and _is_module_available("torch")
def load_embedding_model(entity_type: str = "face") -> bool:
"""Explicitly load the embedding model for the given entity type.
Returns True if the model loaded successfully.
"""
if entity_type == "face":
return get_insightface_app() is not None return get_insightface_app() is not None
model, _ = get_siglip_model() return _is_module_available("insightface") and _is_module_available("onnxruntime")
return model is not None
def load_embedding_model() -> bool:
"""Explicitly load InsightFace. Returns True if the model loaded successfully."""
return get_insightface_app() is not None
+403 -411
View File
File diff suppressed because it is too large Load Diff
+62 -18
View File
@@ -8,9 +8,32 @@ import requests
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
def _get_frigate_url() -> str:
"""Return normalized FRIGATE_URL with whitespace and trailing slash stripped, or '' if unset."""
return os.environ.get("FRIGATE_URL", "").strip().rstrip("/")
def get_frigate_version() -> str | None:
"""Fetch Frigate's version string from GET /api/version.
Returns the version string (e.g. "0.16.0-beta4") or None if FRIGATE_URL
is unset, the endpoint is unreachable, or the response is not parseable.
"""
frigate_url = _get_frigate_url()
if not frigate_url:
return None
try:
resp = requests.get(f"{frigate_url}/api/version", timeout=5)
if resp.ok:
return resp.text.strip().strip('"')
return None
except Exception:
return None
def _get_faces_data() -> dict | None: def _get_faces_data() -> dict | None:
"""Fetch raw GET /api/faces response. Returns None if unavailable.""" """Fetch raw GET /api/faces response. Returns None if unavailable."""
frigate_url = os.environ.get("FRIGATE_URL", "").rstrip("/") frigate_url = _get_frigate_url()
if not frigate_url: if not frigate_url:
return None return None
try: try:
@@ -18,7 +41,7 @@ def _get_faces_data() -> dict | None:
resp.raise_for_status() resp.raise_for_status()
return resp.json() return resp.json()
except Exception as e: except Exception as e:
logger.warning(f"Could not query Frigate faces API: {e}") logger.warning("Could not query Frigate faces API: %s", e)
return None return None
@@ -33,11 +56,17 @@ def get_all_frigate_person_files() -> dict[str, list[str]] | None:
return None return None
# Response: {person_name: [file, ...], "train": [...], ...} # Response: {person_name: [file, ...], "train": [...], ...}
# "train" is a flat pending list, not a person — skip it. # "train" is a flat pending list, not a person — skip it.
return { # TODO(frigate-api): "train" is the only known special key as of Frigate v0.16.
name: files # Log unexpected non-list values so future Frigate schema additions are visible.
for name, files in data.items() result = {}
if name != "train" and isinstance(files, list) for name, files in data.items():
} if name == "train":
continue
if isinstance(files, list):
result[name] = files
else:
logger.debug("Frigate API: skipping unexpected key %r (got %s, not list)", name, type(files).__name__)
return result
def get_frigate_face_counts() -> dict[str, int] | None: def get_frigate_face_counts() -> dict[str, int] | None:
@@ -62,7 +91,10 @@ def get_frigate_person_files(person_name: str) -> list[str] | None:
if data is None: if data is None:
return None return None
files = data.get(person_name) files = data.get(person_name)
return files if isinstance(files, list) else [] if files is not None and not isinstance(files, list):
logger.debug("Frigate API: unexpected type for %r — got %s, not list", person_name, type(files).__name__)
return []
return files if files is not None else []
def recognize_face(file_path: str) -> tuple[str | None, float] | None: def recognize_face(file_path: str) -> tuple[str | None, float] | None:
@@ -75,8 +107,18 @@ def recognize_face(file_path: str) -> tuple[str | None, float] | None:
Returns None if FRIGATE_URL is unset, the API is unreachable, no face is Returns None if FRIGATE_URL is unset, the API is unreachable, no face is
detected, or face recognition is not enabled in Frigate. detected, or face recognition is not enabled in Frigate.
LIMITATION — mean embedding comparison: the score reflects similarity to
the arithmetic mean of all training embeddings, not to individual ones.
A bimodal training set (e.g. frontals + profiles) has a mean that sits
between both clusters, making candidates from either cluster look more
novel than they are. Winnow could add redundant frontals while the score
suggests novelty, because the mean is pulled toward profiles.
TODO(frigate-api): if Frigate exposes per-file embeddings via the API,
replace mean-comparison with nearest-neighbour distance across individual
training embeddings for accurate coverage detection.
""" """
frigate_url = os.environ.get("FRIGATE_URL", "").rstrip("/") frigate_url = _get_frigate_url()
if not frigate_url: if not frigate_url:
return None return None
try: try:
@@ -93,7 +135,7 @@ def recognize_face(file_path: str) -> tuple[str | None, float] | None:
return (data.get("face_name"), round(float(data["score"]), 4)) return (data.get("face_name"), round(float(data["score"]), 4))
return None return None
except Exception as e: except Exception as e:
logger.debug(f"Frigate recognize failed for {file_path}: {e}") logger.debug("Frigate recognize failed for %s: %s", file_path, e)
return None return None
@@ -103,27 +145,29 @@ def delete_frigate_person_files(person_name: str, filenames: list[str]) -> bool:
Uses POST /api/faces/{name}/delete with body {"ids": [filename, ...]}. Uses POST /api/faces/{name}/delete with body {"ids": [filename, ...]}.
Returns True on success, False if unreachable or the request fails. Returns True on success, False if unreachable or the request fails.
""" """
frigate_url = os.environ.get("FRIGATE_URL", "").rstrip("/") frigate_url = _get_frigate_url()
if not frigate_url or not filenames: if not frigate_url:
return False return False
if not filenames:
return True
from urllib.parse import quote from urllib.parse import quote
encoded = quote(person_name, safe="") encoded_name = quote(person_name, safe="")
try: try:
resp = requests.post( resp = requests.post(
f"{frigate_url}/api/faces/{encoded}/delete", f"{frigate_url}/api/faces/{encoded_name}/delete",
json={"ids": filenames}, json={"ids": filenames},
timeout=10, timeout=10,
) )
if resp.ok: if resp.ok:
logger.debug(f"Deleted {len(filenames)} Frigate file(s) for {person_name}") logger.debug("Deleted %s Frigate file(s) for %s", len(filenames), person_name)
return True return True
if resp.status_code == 404: if resp.status_code == 404:
# File already absent — stale tracker entry. Return True so the caller # File already absent — stale tracker entry. Return True so the caller
# removes it from the tracker and frees the slot cleanly. # removes it from the tracker and frees the slot cleanly.
logger.warning(f"Frigate file(s) not found for {person_name} (stale tracker entry?): {filenames}") logger.warning("Frigate file(s) not found for %s (stale tracker entry?): %s", person_name, filenames)
return True return True
logger.warning(f"Frigate delete returned {resp.status_code} for {person_name}") logger.warning("Frigate delete returned %s for %s", resp.status_code, person_name)
return False return False
except Exception as e: except Exception as e:
logger.warning(f"Failed to delete Frigate files for {person_name}: {e}") logger.warning("Failed to delete Frigate files for %s: %s", person_name, e)
return False return False
+35 -76
View File
@@ -1,7 +1,8 @@
"""Image processing functions for cropping faces and objects.""" """Image processing functions for cropping faces."""
import logging import logging
import os import os
import warnings
import numpy as np import numpy as np
from PIL import Image from PIL import Image
@@ -10,25 +11,20 @@ from .config import Config
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# Lazy singleton
_yolo_model = None
def _save_jpeg(img: Image.Image, path: str) -> None: def _save_jpeg(img: Image.Image, path: str) -> None:
if img.mode != "RGB": if img.mode != "RGB":
img = img.convert("RGB") img = img.convert("RGB")
img.save(path, format="JPEG") tmp = path + ".tmp"
try:
img.save(tmp, format="JPEG")
def get_yolo_model(): os.replace(tmp, path)
"""Singleton for YOLO model.""" except Exception:
global _yolo_model try:
if _yolo_model is None: os.remove(tmp)
from ultralytics import YOLO except OSError:
pass
logger.info("Loading YOLOv9c model...") raise
_yolo_model = YOLO("yolov9c.pt")
return _yolo_model
def align_face(img: Image.Image, landmarks: list[list[float]] | np.ndarray) -> Image.Image | None: def align_face(img: Image.Image, landmarks: list[list[float]] | np.ndarray) -> Image.Image | None:
@@ -50,15 +46,17 @@ def align_face(img: Image.Image, landmarks: list[list[float]] | np.ndarray) -> I
img_np = np.asarray(img) img_np = np.asarray(img)
lm = np.array(landmarks, dtype=np.float32) lm = np.array(landmarks, dtype=np.float32)
if lm.shape != (5, 2): if lm.shape != (5, 2):
logger.debug(f"Invalid landmark shape: {lm.shape}, expected (5, 2)") logger.debug("Invalid landmark shape: %s, expected (5, 2)", lm.shape)
return None return None
aligned = norm_crop(img_np, lm) with warnings.catch_warnings():
warnings.filterwarnings("ignore", message=".*estimate.*is deprecated", category=FutureWarning)
aligned = norm_crop(img_np, lm)
return Image.fromarray(aligned) return Image.fromarray(aligned)
except ImportError: except ImportError:
logger.debug("InsightFace not available for face alignment") logger.debug("InsightFace not available for face alignment")
return None return None
except Exception as e: except Exception as e:
logger.debug(f"Face alignment failed: {e}") logger.debug("Face alignment failed: %s", e)
return None return None
@@ -70,10 +68,11 @@ def process_face_mode(
count: int, count: int,
min_width: int | None = None, min_width: int | None = None,
insightface_app=None, insightface_app=None,
) -> tuple[int, int] | None: ) -> tuple[int, int] | str:
"""Crop face based on Immich metadata and save to output directory. """Crop face based on Immich metadata and save to output directory.
Returns (width, height) of the saved crop, or None if no crop was saved. Returns (width, height) of the saved crop, or a skip-reason string if the
face was filtered out.
When insightface_app is provided and ENABLE_FACE_ALIGNMENT is True, When insightface_app is provided and ENABLE_FACE_ALIGNMENT is True,
re-detects the face in the Immich bbox region using InsightFace to get re-detects the face in the Immich bbox region using InsightFace to get
precise landmarks for a proper 112x112 aligned crop. Falls back to precise landmarks for a proper 112x112 aligned crop. Falls back to
@@ -92,15 +91,18 @@ def process_face_mode(
break break
if not face_info: if not face_info:
logger.debug(f"No face info for {person.get('name')} in asset {asset.get('id')}") logger.debug("No face info for %s in asset %s", person.get("name"), asset.get("id"))
return None return "no face metadata"
img_w, img_h = img.size img_w, img_h = img.size
meta_w = face_info.get("imageWidth") or img_w meta_w = face_info.get("imageWidth") or 0
meta_h = face_info.get("imageHeight") or img_h meta_h = face_info.get("imageHeight") or 0
# Scale bounding box to actual image dimensions # Scale bounding box from detection-image space to actual image dimensions.
scale_x, scale_y = img_w / meta_w, img_h / meta_h # Fall back to 1.0 if Immich omits the field — bbox is assumed to already
# be in image space (correct for thumbnails, wrong for full-res).
scale_x = img_w / meta_w if meta_w else 1.0
scale_y = img_h / meta_h if meta_h else 1.0
x1 = face_info["boundingBoxX1"] * scale_x x1 = face_info["boundingBoxX1"] * scale_x
y1 = face_info["boundingBoxY1"] * scale_y y1 = face_info["boundingBoxY1"] * scale_y
x2 = face_info["boundingBoxX2"] * scale_x x2 = face_info["boundingBoxX2"] * scale_x
@@ -108,8 +110,8 @@ def process_face_mode(
face_w, face_h = x2 - x1, y2 - y1 face_w, face_h = x2 - x1, y2 - y1
if face_w < min_width or face_h < min_width: if face_w < min_width or face_h < min_width:
logger.debug(f"Face too small ({face_w:.1f}x{face_h:.1f})") logger.debug("Face too small (%.1fx%.1f)", face_w, face_h)
return None return f"face too small ({face_w:.0f}x{face_h:.0f}px, min {min_width}px)"
# Re-detect face with InsightFace for landmark-based alignment. # Re-detect face with InsightFace for landmark-based alignment.
# Immich's /api/faces endpoint does not include landmarks, so the # Immich's /api/faces endpoint does not include landmarks, so the
@@ -127,7 +129,9 @@ def process_face_mode(
min(img_h, y2 + pad_y), min(img_h, y2 + pad_y),
) )
search_crop = img.crop(search_box) search_crop = img.crop(search_box)
detected = insightface_app.get(np.asarray(search_crop)) with warnings.catch_warnings():
warnings.filterwarnings("ignore", message=".*estimate.*is deprecated", category=FutureWarning)
detected = insightface_app.get(np.asarray(search_crop))
if detected: if detected:
cx, cy = search_crop.width / 2, search_crop.height / 2 cx, cy = search_crop.width / 2, search_crop.height / 2
best = min( best = min(
@@ -142,7 +146,7 @@ def process_face_mode(
_save_jpeg(aligned, os.path.join(output_dir, f"{count}.jpg")) _save_jpeg(aligned, os.path.join(output_dir, f"{count}.jpg"))
return aligned.size return aligned.size
except Exception as e: except Exception as e:
logger.debug(f"InsightFace re-detection failed for {asset.get('id')}: {e}") logger.debug("InsightFace re-detection failed for %s: %s", asset.get("id"), e)
# Landmark alignment from Immich metadata (Immich does not currently # Landmark alignment from Immich metadata (Immich does not currently
# expose landmarks, so this path is a future-proofing fallback) # expose landmarks, so this path is a future-proofing fallback)
@@ -170,49 +174,4 @@ def process_face_mode(
return face_crop.size return face_crop.size
def process_object_mode(
img: Image.Image,
config: dict,
output_dir: str,
count: int,
) -> bool:
"""Detect and crop objects using YOLO."""
try:
model = get_yolo_model()
target_class = config.get("object_class", "dog")
import torch
if os.getenv("FORCE_CPU", "").lower() in ("true", "1", "yes"):
device = "cpu"
elif hasattr(torch, "xpu") and torch.xpu.is_available():
device = "xpu"
else:
device = None # YOLO auto-selects (CUDA/ROCm/CPU)
results = model(img, verbose=False, device=device)
found = False
class_idx = 0 # Sequential counter per target class (Issue #10)
for box in (box for r in results for box in r.boxes):
cls_id = int(box.cls[0])
conf = float(box.conf[0])
if 0 <= cls_id < len(model.names) and model.names[cls_id] == target_class and conf > 0.5:
x1, y1, x2, y2 = box.xyxy[0].tolist()
_save_jpeg(
img.crop((x1, y1, x2, y2)),
os.path.join(output_dir, f"{count}_{class_idx}.jpg"),
)
class_idx += 1
found = True
return found
except Exception as e:
logger.error(f"YOLO processing failed: {e}")
return False
def process_full_mode(img: Image.Image, output_dir: str, count: int) -> bool:
"""Save full image."""
_save_jpeg(img, os.path.join(output_dir, f"{count}.jpg"))
return True
+122 -46
View File
@@ -5,7 +5,6 @@ from dataclasses import dataclass
from datetime import datetime, timedelta, timezone from datetime import datetime, timedelta, timezone
from io import BytesIO from io import BytesIO
import numpy as np
import requests import requests
from PIL import Image, ImageOps from PIL import Image, ImageOps
@@ -14,19 +13,42 @@ from .config import Config, get_headers
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
MAX_PAGES = 1000 # Safety limit for pagination MAX_PAGES = 1000 # Safety limit for pagination
_MAX_ASSETS_PER_PERSON = 5000 # Stop fetching after this many — diversity pool is capped at 3000 anyway
@dataclass @dataclass
class FaceData: class FaceData:
"""Pre-computed face data from Immich.""" """Pre-computed face data from Immich."""
embedding: np.ndarray | None
bbox: tuple[float, float, float, float] # (x1, y1, x2, y2) bbox: tuple[float, float, float, float] # (x1, y1, x2, y2)
confidence: float | None confidence: float | None
image_width: int image_width: int
image_height: int image_height: int
def get_immich_version() -> tuple[int, int, int] | None:
"""Fetch Immich server version from GET /api/server/version.
Returns (major, minor, patch) or None if unreachable or unparseable.
"""
try:
resp = requests.get(
f"{Config.IMMICH_URL}/api/server/version",
headers=get_headers(),
timeout=5,
)
if resp.ok:
data = resp.json()
major, minor, patch = data.get("major"), data.get("minor"), data.get("patch")
if major is None or minor is None or patch is None:
logger.debug("Unexpected Immich version schema: %s", data)
return None
return (int(major), int(minor), int(patch))
return None
except Exception:
return None
def get_people() -> list[dict]: def get_people() -> list[dict]:
"""Fetch all people from Immich.""" """Fetch all people from Immich."""
try: try:
@@ -39,9 +61,13 @@ def get_people() -> list[dict]:
logger.error("Immich API key is invalid or expired (401 Unauthorized). Update API_KEY.") logger.error("Immich API key is invalid or expired (401 Unauthorized). Update API_KEY.")
return [] return []
resp.raise_for_status() resp.raise_for_status()
return resp.json().get("people", []) data = resp.json()
except (requests.RequestException, ValueError) as e: if not isinstance(data, dict):
logger.error(f"Failed to fetch people from Immich: {e}") logger.error("Unexpected response shape from Immich /people: %r", type(data))
return []
return data.get("people") or []
except (requests.RequestException, ValueError, AttributeError) as e:
logger.error("Failed to fetch people from Immich: %s", e)
return [] return []
@@ -61,20 +87,32 @@ def merge_people(survivor_id: str, merge_ids: list[str]) -> bool:
resp.raise_for_status() resp.raise_for_status()
return True return True
except requests.RequestException as e: except requests.RequestException as e:
logger.error(f"Failed to merge people into {survivor_id}: {e}") logger.error("Failed to merge people into %s: %s", survivor_id, e)
return False return False
def fetch_all_assets(person: dict) -> list[dict]: def fetch_all_assets(person: dict) -> tuple[list[dict], int]:
"""Fetch all assets for a person with pagination.""" """Fetch all assets for a person with pagination.
Returns (assets, total_raw) where assets is the list of valid dict items
and total_raw is the raw item count across pages that had at least one valid
dict. All-garbage pages (every item non-dict) stop pagination and are not
counted. total_raw is a lower bound in two cases: a network error interrupts
pagination (a warning is logged), or an all-garbage page terminates it early
(a warning is logged and later pages are not fetched).
"""
name = person.get("name", "Unknown") name = person.get("name", "Unknown")
person_id = person["id"] person_id = person.get("id")
if not person_id:
logger.error("Person dict missing 'id' field for %s — skipping asset fetch", name)
return [], 0
url = f"{Config.IMMICH_URL}/api/search/metadata" url = f"{Config.IMMICH_URL}/api/search/metadata"
page_size = 1000 page_size = 1000
logger.debug(f"Fetching assets for {name}...") logger.debug("Fetching assets for %s...", name)
assets = [] assets: list[dict] = []
total_raw = 0 # raw item count across pages that yielded at least one valid dict
for page in range(1, MAX_PAGES + 1): for page in range(1, MAX_PAGES + 1):
try: try:
resp = requests.post( resp = requests.post(
@@ -85,31 +123,64 @@ def fetch_all_assets(person: dict) -> list[dict]:
) )
if not resp.ok: if not resp.ok:
logger.error(f"Error fetching assets for {name} (page {page}): {resp.status_code}") logger.error("Error fetching assets for %s (page %s): %s", name, page, resp.status_code)
break break
page_assets = resp.json().get("assets", []) body = resp.json()
if not isinstance(body, dict):
logger.error("Unexpected response shape fetching assets for %s (page %s): %r", name, page, type(body))
break
page_assets = body.get("assets", [])
# Immich ≥2.x returns {"assets": {"items": [...]}};
# earlier versions returned {"assets": [...]} directly.
if isinstance(page_assets, dict): if isinstance(page_assets, dict):
page_assets = page_assets.get("items", []) page_assets = page_assets.get("items") or []
if not page_assets: page_count = len(page_assets) # raw count for termination check before filtering
# Single pass: partition valid assets from unexpected non-dict items
valid_assets, skipped_count = [], 0
for item in page_assets:
if isinstance(item, dict):
valid_assets.append(item)
else:
skipped_count += 1
if skipped_count:
logger.warning("%s: skipping %s non-dict item(s) in page %s", name, skipped_count, page)
if not valid_assets:
if page_count > 0:
logger.warning(
"%s: page %s returned %s item(s) but none were valid dicts — stopping pagination",
name, page, page_count,
)
break break
assets.extend(page_assets) # Count page_count (not just valid items) so that non-dict items from a
logger.debug(f"Fetched page {page}, total: {len(assets)}") # transient schema issue on a mixed page don't cause MIN_FACE_COUNT to
# skip a real person. Pages where every item is a non-dict are excluded —
# they indicate a structural problem and break above without contributing.
total_raw += page_count
assets.extend(valid_assets)
logger.debug("Fetched page %s, total: %s", page, len(assets))
if len(page_assets) < page_size: if page_count < page_size or len(assets) >= _MAX_ASSETS_PER_PERSON:
break break
except (requests.RequestException, ValueError) as e: except (requests.RequestException, ValueError) as e:
logger.error(f"Exception fetching assets for {name}: {e}") logger.error("Exception fetching assets for %s (page %s): %s", name, page, e)
if page > 1:
logger.warning(
"%s: pagination interrupted at page %s — total_raw=%s may undercount actual assets",
name, page, total_raw,
)
break break
return assets return assets, total_raw
def fetch_face_data(asset_id: str, person_id: str | None = None) -> FaceData | None: def fetch_face_data(asset_id: str, person_id: str | None = None) -> FaceData | None:
"""Fetch pre-computed face data (embedding, bbox, confidence) from Immich. """Fetch pre-computed face data (bbox, confidence) from Immich.
Queries GET /api/faces?id={asset_id} to retrieve face detection results Queries GET /api/faces?id={asset_id} to retrieve face detection results
that Immich already computed using InsightFace Buffalo_L. that Immich already computed using InsightFace Buffalo_L.
@@ -119,7 +190,7 @@ def fetch_face_data(asset_id: str, person_id: str | None = None) -> FaceData | N
person_id: Optional person ID to match the specific face person_id: Optional person ID to match the specific face
Returns: Returns:
FaceData with embedding, bbox, and confidence, or None if unavailable FaceData with bbox and confidence, or None if unavailable
""" """
try: try:
resp = requests.get( resp = requests.get(
@@ -130,29 +201,25 @@ def fetch_face_data(asset_id: str, person_id: str | None = None) -> FaceData | N
) )
if not resp.ok: if not resp.ok:
logger.debug(f"Face data endpoint returned {resp.status_code} for {asset_id}") logger.debug("Face data endpoint returned %s for %s", resp.status_code, asset_id)
return None return None
faces = resp.json() faces = resp.json()
if not faces: if not isinstance(faces, list) or not faces:
return None return None
# Match the target person if specified # Match the target person if specified; never fall back to a different person's face.
face = None face = None
if person_id: if person_id:
face = next( face = next(
(f for f in faces if (f.get("person") or {}).get("id") == person_id), (f for f in faces if isinstance(f, dict) and (f.get("person") or {}).get("id") == person_id),
None, None,
) )
else:
face = faces[0] if isinstance(faces[0], dict) else None
if face is None: if face is None:
face = faces[0] # Fall back to first/largest face return None
# Extract embedding if available
embedding = None
if "embedding" in face:
embedding = np.array(face["embedding"], dtype=np.float32)
# Extract bounding box
bbox = ( bbox = (
face.get("boundingBoxX1", 0), face.get("boundingBoxX1", 0),
face.get("boundingBoxY1", 0), face.get("boundingBoxY1", 0),
@@ -162,7 +229,6 @@ def fetch_face_data(asset_id: str, person_id: str | None = None) -> FaceData | N
score = face.get("score") score = face.get("score")
return FaceData( return FaceData(
embedding=embedding,
bbox=bbox, bbox=bbox,
confidence=score if score is not None else face.get("confidence"), confidence=score if score is not None else face.get("confidence"),
image_width=face.get("imageWidth", 0), image_width=face.get("imageWidth", 0),
@@ -170,10 +236,10 @@ def fetch_face_data(asset_id: str, person_id: str | None = None) -> FaceData | N
) )
except requests.RequestException as e: except requests.RequestException as e:
logger.debug(f"Failed to fetch face data for {asset_id}: {e}") logger.debug("Failed to fetch face data for %s: %s", asset_id, e)
return None return None
except (AttributeError, KeyError, TypeError, ValueError) as e: except (AttributeError, KeyError, TypeError, ValueError) as e:
logger.debug(f"Failed to parse face data for {asset_id}: {e}") logger.debug("Failed to parse face data for %s: %s", asset_id, e)
return None return None
@@ -194,9 +260,9 @@ def fetch_full_image(asset_id: str, timeout: int = 60) -> Image.Image | None:
try: try:
return ImageOps.exif_transpose(Image.open(BytesIO(resp.content))) return ImageOps.exif_transpose(Image.open(BytesIO(resp.content)))
except Exception: except Exception:
logger.debug(f"PIL can't open original for {asset_id}, falling back to preview") logger.debug("PIL can't open original for %s, falling back to preview", asset_id)
except requests.RequestException: except requests.RequestException:
logger.debug(f"Original request failed for {asset_id}, falling back to preview") logger.debug("Original request failed for %s, falling back to preview", asset_id)
# Fall back to preview thumbnail (always JPEG) # Fall back to preview thumbnail (always JPEG)
try: try:
@@ -208,22 +274,26 @@ def fetch_full_image(asset_id: str, timeout: int = 60) -> Image.Image | None:
if resp.ok: if resp.ok:
return ImageOps.exif_transpose(Image.open(BytesIO(resp.content))) return ImageOps.exif_transpose(Image.open(BytesIO(resp.content)))
except Exception as e: except Exception as e:
logger.error(f"Failed to fetch image {asset_id}: {e}") logger.error("Failed to fetch image %s: %s", asset_id, e)
return None return None
def filter_recent_assets(assets: list[dict], years: int | None = None) -> list[dict]: def filter_recent_assets(assets: list[dict], years: int | None = None) -> list[dict]:
"""Filter assets to keep only those from the last N years.""" """Filter assets to keep only those from the last N years. Pass years=0 to include all."""
years = years or Config.YEARS_FILTER if years is None:
years = Config.YEARS_FILTER
if not years:
return list(assets)
cutoff = datetime.now(timezone.utc) - timedelta(days=365 * years) cutoff = datetime.now(timezone.utc) - timedelta(days=365 * years)
logger.debug(f"Filtering assets older than {years} years ({cutoff})") logger.debug("Filtering assets older than %s years (%s)", years, cutoff)
recent, skipped = [], 0 recent, skipped, bad_timestamp = [], 0, 0
for asset in assets: for asset in assets:
created_at_str = asset.get("fileCreatedAt") created_at_str = asset.get("fileCreatedAt")
if not created_at_str: if not isinstance(created_at_str, str) or not created_at_str:
bad_timestamp += 1
continue continue
try: try:
@@ -233,9 +303,15 @@ def filter_recent_assets(assets: list[dict], years: int | None = None) -> list[d
recent.append(asset) recent.append(asset)
else: else:
skipped += 1 skipped += 1
except ValueError: except (ValueError, TypeError):
bad_timestamp += 1
continue continue
logger.debug(f"Retained {len(recent)} assets (filtered {skipped} old assets).") if bad_timestamp:
logger.warning(
"filter_recent_assets: %s asset(s) had missing or unparseable fileCreatedAt"
" and were excluded from the pool.", bad_timestamp
)
logger.debug("Retained %s assets (filtered %s old assets).", len(recent), skipped)
return recent return recent
+113 -104
View File
@@ -8,7 +8,7 @@ from rich.progress import BarColumn, Progress, SpinnerColumn, TaskProgressColumn
from rich.prompt import Confirm, IntPrompt, Prompt from rich.prompt import Confirm, IntPrompt, Prompt
from rich.table import Table from rich.table import Table
from .config import Config from .config import Config, _getenv_bool, _getenv_int, _getenv_optional_int
from .diversity import select_diverse_assets from .diversity import select_diverse_assets
from .embeddings import is_embedding_available, load_embedding_model from .embeddings import is_embedding_available, load_embedding_model
from .frigate_api import get_frigate_face_counts from .frigate_api import get_frigate_face_counts
@@ -20,18 +20,16 @@ logger = logging.getLogger(__name__)
# Strategy presets: (limit, mode_name) # Strategy presets: (limit, mode_name)
STRATEGY_PRESETS = { STRATEGY_PRESETS = {
"1": ("auto", "Auto Diversity"), "1": ("auto", "Adaptive Diversity"),
"2": (30, "Standard (30)"), "2": (30, "Standard (30)"),
"3": (100, "Broad (100)"), "3": (100, "Broad (100)"),
} }
def _get_strategy_choice(has_embedding: bool, entity_type: str) -> tuple[int | str, str]: def _get_strategy_choice(has_embedding: bool) -> tuple[int | str, str]:
"""Prompt user for training strategy and return (limit, selection_mode).""" """Prompt user for training strategy and return (limit, selection_mode)."""
model_name = "InsightFace" if entity_type == "face" else "SigLIP"
if has_embedding: if has_embedding:
rprint(" [bold]1.[/bold] Auto (Objective Diversity) [green][Recommended][/green]") rprint(" [bold]1.[/bold] Adaptive Diversity [green][Recommended][/green]")
rprint(" [dim]• Dynamically selects images until redundancy starts[/dim]") rprint(" [dim]• Dynamically selects images until redundancy starts[/dim]")
rprint(" [bold]2.[/bold] Standard (30 images)") rprint(" [bold]2.[/bold] Standard (30 images)")
rprint(" [bold]3.[/bold] Broad (100 images)") rprint(" [bold]3.[/bold] Broad (100 images)")
@@ -51,7 +49,7 @@ def _get_strategy_choice(has_embedding: bool, entity_type: str) -> tuple[int | s
return 30, "smart" return 30, "smart"
# Fallback when embedding model not available # Fallback when embedding model not available
rprint(f" [yellow]Note: {model_name} not available. Using Time Spread.[/yellow]") rprint(" [yellow]Note: InsightFace not available. Using Time Spread.[/yellow]")
rprint(" [bold]1.[/bold] Standard (30 images) [green][Recommended][/green]") rprint(" [bold]1.[/bold] Standard (30 images) [green][Recommended][/green]")
rprint(" [bold]2.[/bold] Broad (100 images)") rprint(" [bold]2.[/bold] Broad (100 images)")
rprint(" [bold]3.[/bold] Custom Count") rprint(" [bold]3.[/bold] Custom Count")
@@ -67,33 +65,42 @@ def _get_strategy_choice(has_embedding: bool, entity_type: str) -> tuple[int | s
def _resolve_strategy(strategy: str, has_embedding: bool) -> tuple[int | str, str]: def _resolve_strategy(strategy: str, has_embedding: bool) -> tuple[int | str, str]:
"""Resolve env var strategy to (limit, selection_mode) without prompts.""" """Resolve env var strategy to (limit, selection_mode) without prompts."""
custom_limit = os.environ.get("LIMIT", "").strip() if strategy == "skip":
return 0, "skip"
if not has_embedding: if not has_embedding:
limit = int(custom_limit) if custom_limit else 30 limit = _getenv_int("LIMIT", 30)
if limit <= 0:
logger.warning("LIMIT=%s is invalid — ignoring and using default 30", limit)
limit = 30
return limit, "time" return limit, "time"
if custom_limit: custom_limit = _getenv_optional_int("LIMIT")
return int(custom_limit), "smart" if custom_limit is not None:
if custom_limit > 0:
return custom_limit, "smart"
logger.warning("LIMIT=%s is invalid — ignoring and using auto strategy", custom_limit)
strategy_map = { strategy_map = {
"auto": ("auto", "smart"), "adaptive": ("auto", "smart"),
"auto": ("auto", "smart"), # legacy alias for adaptive
"standard": (30, "smart"), "standard": (30, "smart"),
"broad": (100, "smart"), "broad": (100, "smart"),
} }
return strategy_map.get(strategy, ("auto", "smart")) result = strategy_map.get(strategy)
if result is None:
logger.warning("Unrecognised STRATEGY=%r — falling back to auto", strategy)
return ("auto", "smart")
return result
def _perform_selection( def _perform_selection(
assets: list, limit: int | str, name: str, selection_mode: str, entity_type: str, person_id: str | None = None assets: list, limit: int | str, name: str, selection_mode: str, person_id: str | None = None
) -> list: ) -> list:
"""Run diversity selection with progress display.""" """Run diversity selection with progress display."""
if selection_mode == "smart": if selection_mode == "smart":
model_display = "InsightFace (face embeddings)" if entity_type == "face" else "SigLIP (visual embeddings)" rprint("\n[cyan]Using InsightFace (face embeddings) for diversity analysis...[/cyan]")
rprint(f"\n[cyan]Using {model_display} for diversity analysis...[/cyan]")
# Pre-load model explicitly (separate from availability check) load_embedding_model()
load_embedding_model(entity_type)
with Progress( with Progress(
SpinnerColumn(), SpinnerColumn(),
@@ -108,7 +115,6 @@ def _perform_selection(
limit, limit,
name, name,
selection_mode=selection_mode, selection_mode=selection_mode,
entity_type=entity_type,
person_id=person_id, person_id=person_id,
progress_callback=lambda c, t: progress.update(task, completed=c, total=t), progress_callback=lambda c, t: progress.update(task, completed=c, total=t),
) )
@@ -119,45 +125,51 @@ def _perform_selection(
rprint(f"\n[cyan]Using time-spread selection for {limit} images...[/cyan]") rprint(f"\n[cyan]Using time-spread selection for {limit} images...[/cyan]")
with console.status(f"[bold]Selecting {limit} images evenly distributed over time...[/bold]"): with console.status(f"[bold]Selecting {limit} images evenly distributed over time...[/bold]"):
selected = select_diverse_assets( selected = select_diverse_assets(assets, limit, name, selection_mode="time", person_id=person_id)
assets, limit, name, selection_mode="time", entity_type=entity_type, person_id=person_id
)
rprint(f" [green]Selected {len(selected)} images using time spread.[/green]") rprint(f" [green]Selected {len(selected)} images using time spread.[/green]")
return selected return selected
def _build_job(
person: dict,
assets: list,
limit: int | str,
selection_mode: str,
quality_replacement: bool = False,
) -> dict | None:
"""Select from pre-filtered assets and build a job dict. No terminal I/O."""
if not assets:
return None
name = person["name"]
selected = _perform_selection(assets, limit, name, selection_mode, person_id=person["id"])
if not selected:
return None
return {
"person": person,
"assets": selected,
"limit": len(selected),
"config": {"name": name, "quality_replacement": quality_replacement},
}
def _configure_person(person: dict, people: list[dict]) -> dict | None: def _configure_person(person: dict, people: list[dict]) -> dict | None:
"""Configure training for a single person. Returns job dict or None.""" """Configure training for a single person. Returns job dict or None."""
name = person["name"] name = person["name"]
console.print(f"\nSelected: [bold green]{name}[/bold green]") console.print(f"\nSelected: [bold green]{name}[/bold green]")
# Select training mode
rprint("\n[bold cyan]Training Mode:[/bold cyan]")
rprint(" [bold]1.[/bold] Face (Frigate Face Recognition)")
rprint(" [bold]2.[/bold] Object (Frigate Object Classification)")
mode_choice = Prompt.ask("Choice", choices=["1", "2"], default="1")
entity_type = "face" if mode_choice == "1" else "object"
config = {"name": name, "mode": entity_type, "quality_replacement": Config.QUALITY_REPLACEMENT}
if entity_type == "object":
config["object_class"] = Prompt.ask("Enter Object Class (e.g. dog, cat, car)", default="dog")
# Fetch and filter assets
years = IntPrompt.ask("Filter images older than (years)", default=Config.YEARS_FILTER) years = IntPrompt.ask("Filter images older than (years)", default=Config.YEARS_FILTER)
console.print(f"Scanning for {name} ({entity_type})...") console.print(f"Scanning for {name}...")
with console.status("[bold green]Fetching assets...[/bold green]"): with console.status("[bold green]Fetching assets...[/bold green]"):
all_assets = fetch_all_assets(person) all_assets, total_raw = fetch_all_assets(person)
recent_assets = filter_recent_assets(all_assets, years=years) recent_assets = filter_recent_assets(all_assets, years=years)
rprint(f" Found [bold]{len(all_assets)}[/bold] total, [bold]{len(recent_assets)}[/bold] in range ({years} years).") rprint(f" Found [bold]{total_raw}[/bold] total, [bold]{len(recent_assets)}[/bold] in range ({years} years).")
# Filter out assets already uploaded to Frigate. # Ask before strategy so the post-dedup count can inform the choice
# In interactive mode, ask — use the env var only as the default so it can retry_env = _getenv_bool("RETRY_REJECTED", False)
# still be pre-set (e.g. RETRY_REJECTED=true) without forcing the answer.
retry_env = os.environ.get("RETRY_REJECTED", "false").lower() in ("true", "1", "yes")
retry_rejected = Confirm.ask("Include previously rejected images?", default=retry_env) retry_rejected = Confirm.ask("Include previously rejected images?", default=retry_env)
before_dedup = len(recent_assets) before_dedup = len(recent_assets)
new_asset_ids = set(filter_already_uploaded([a["id"] for a in recent_assets], retry_rejected=retry_rejected)) new_asset_ids = set(filter_already_uploaded([a["id"] for a in recent_assets], retry_rejected=retry_rejected))
recent_assets = [a for a in recent_assets if a["id"] in new_asset_ids] recent_assets = [a for a in recent_assets if a["id"] in new_asset_ids]
@@ -169,21 +181,26 @@ def _configure_person(person: dict, people: list[dict]) -> dict | None:
rprint(" [dim]Skipping (0 new images after dedup).[/dim]") rprint(" [dim]Skipping (0 new images after dedup).[/dim]")
return None return None
# Strategy selection has_embedding = is_embedding_available()
has_embedding = is_embedding_available(entity_type)
rprint(f"\n[bold cyan]Select Training Strategy for {name}:[/bold cyan]") rprint(f"\n[bold cyan]Select Training Strategy for {name}:[/bold cyan]")
limit, selection_mode = _get_strategy_choice(has_embedding)
limit, selection_mode = _get_strategy_choice(has_embedding, entity_type)
if selection_mode == "skip": if selection_mode == "skip":
return None return None
# Perform selection job = _build_job(person, recent_assets, limit, selection_mode, quality_replacement=Config.QUALITY_REPLACEMENT)
selected_assets = _perform_selection( if job is None:
recent_assets, limit, name, selection_mode, entity_type, person_id=person["id"] rprint(" [dim]Skipping (0 images selected).[/dim]")
) return None
rprint(f" [green]Queued {len(selected_assets)} images for {name}.[/green]") rprint(f" [green]Queued {job['limit']} images for {name}.[/green]")
return {"person": person, "assets": selected_assets, "limit": len(selected_assets), "config": config} return job
def _valid_people(people: list[dict]) -> list[dict]:
return sorted(
[p for p in people if (p.get("name") or "").strip() and p.get("id")],
key=lambda x: x["name"],
)
def interactive_configure(people: list[dict]) -> list[dict]: def interactive_configure(people: list[dict]) -> list[dict]:
@@ -192,7 +209,7 @@ def interactive_configure(people: list[dict]) -> list[dict]:
Supports multi-person batch mode — after configuring one person, Supports multi-person batch mode — after configuring one person,
prompts to add another. prompts to add another.
""" """
valid_people = sorted([p for p in people if p.get("name")], key=lambda x: x["name"]) valid_people = _valid_people(people)
if not valid_people: if not valid_people:
rprint("[red]No people found with names in Immich.[/red]") rprint("[red]No people found with names in Immich.[/red]")
@@ -203,9 +220,9 @@ def interactive_configure(people: list[dict]) -> list[dict]:
while True: while True:
# Select person # Select person
console.print("\n[bold cyan]Select Person to Train:[/bold cyan]") console.print("\n[bold cyan]Select Person to Train:[/bold cyan]")
queued_ids = {j["person"]["id"] for j in jobs}
for idx, p in enumerate(valid_people, 1): for idx, p in enumerate(valid_people, 1):
# Mark already-queued people marker = " [dim](queued)[/dim]" if p.get("id") in queued_ids else ""
marker = " [dim](queued)[/dim]" if any(j["person"]["id"] == p["id"] for j in jobs) else ""
console.print(f" [bold]{idx}.[/bold] {p['name']}{marker}") console.print(f" [bold]{idx}.[/bold] {p['name']}{marker}")
p_choice = IntPrompt.ask("Enter Number", choices=[str(i) for i in range(1, len(valid_people) + 1)]) p_choice = IntPrompt.ask("Enter Number", choices=[str(i) for i in range(1, len(valid_people) + 1)])
@@ -224,31 +241,22 @@ def interactive_configure(people: list[dict]) -> list[dict]:
def auto_configure(people: list[dict]) -> list[dict]: def auto_configure(people: list[dict]) -> list[dict]:
"""Non-interactive: configure jobs for all named people automatically.""" """Non-interactive: configure jobs for all named people automatically."""
valid_people = sorted([p for p in people if p.get("name")], key=lambda x: x["name"]) valid_people = _valid_people(people)
if not valid_people: if not valid_people:
rprint("[red]No people found with names in Immich.[/red]") rprint("[red]No people found with names in Immich.[/red]")
return [] return []
mode = os.environ.get("TRAINING_MODE", "face")
strategy = os.environ.get("STRATEGY", "auto") strategy = os.environ.get("STRATEGY", "auto")
skip = os.environ.get("SKIP_PEOPLE", "").split(",") if os.environ.get("SKIP_PEOPLE") else [] skip = {s.strip().casefold() for s in os.environ.get("SKIP_PEOPLE", "").split(",") if s.strip()}
only = os.environ.get("ONLY_PEOPLE", "").split(",") if os.environ.get("ONLY_PEOPLE") else [] only = {s.strip().casefold() for s in os.environ.get("ONLY_PEOPLE", "").split(",") if s.strip()}
if only: if only:
valid_people = [p for p in valid_people if p["name"] in only] valid_people = [p for p in valid_people if p["name"].casefold() in only]
if skip: if skip:
valid_people = [p for p in valid_people if p["name"] not in skip] valid_people = [p for p in valid_people if p["name"].casefold() not in skip]
# Filter by minimum face count (Issue #6: previously unimplemented)
min_face_count = Config.MIN_FACE_COUNT min_face_count = Config.MIN_FACE_COUNT
if min_face_count > 0:
valid_people = [p for p in valid_people if p.get("assetCount", 0) >= min_face_count]
if valid_people:
rprint(
f" Filtered to {len(valid_people)} people with"
f" ≥{min_face_count} assets (MIN_FACE_COUNT={min_face_count})"
)
frigate_counts = get_frigate_face_counts() frigate_counts = get_frigate_face_counts()
# Persist each count to tracker so the last known value survives Frigate downtime # Persist each count to tracker so the last known value survives Frigate downtime
@@ -259,28 +267,19 @@ def auto_configure(people: list[dict]) -> list[dict]:
jobs = [] jobs = []
for person in valid_people: for person in valid_people:
name = person["name"] name = person["name"]
entity_type = mode
config = {"name": name, "mode": entity_type} all_assets, total_raw = fetch_all_assets(person)
if entity_type == "object":
config["object_class"] = os.environ.get("OBJECT_CLASS", "dog")
all_assets = fetch_all_assets(person)
recent_assets = filter_recent_assets(all_assets, years=Config.YEARS_FILTER) recent_assets = filter_recent_assets(all_assets, years=Config.YEARS_FILTER)
rprint(f" {name}: {len(all_assets)} total, {len(recent_assets)} recent") rprint(f" {name}: {total_raw} total, {len(recent_assets)} recent")
# Filter out assets already uploaded to Frigate # MIN_FACE_COUNT guard: skip people with too few Immich assets.
retry_rejected = os.environ.get("RETRY_REJECTED", "false").lower() in ("true", "1", "yes") # Uses total_raw so that non-dict items from a transient Immich schema
before_dedup = len(recent_assets) # issue on a mixed page don't shrink the count below the threshold.
new_asset_ids = set(filter_already_uploaded([a["id"] for a in recent_assets], retry_rejected=retry_rejected)) # Done here (after fetch) rather than upfront because Immich v2.7.5+
recent_assets = [a for a in recent_assets if a["id"] in new_asset_ids] # dropped assetCount from the /api/people response.
skipped = before_dedup - len(recent_assets) if min_face_count > 0 and total_raw < min_face_count:
if skipped: rprint(f" [dim]Skipping {name} ({total_raw} assets < MIN_FACE_COUNT={min_face_count}).[/dim]")
rprint(f" [dim]Skipped {skipped} assets already uploaded to Frigate.[/dim]")
if not recent_assets:
rprint(f" [dim]Skipping {name} (0 new images after dedup).[/dim]")
continue continue
# Enforce MAX_AUTO_IMAGES against the tracked file count only. # Enforce MAX_AUTO_IMAGES against the tracked file count only.
@@ -304,33 +303,46 @@ def auto_configure(people: list[dict]) -> list[dict]:
else: else:
quality_replacement_only = False quality_replacement_only = False
config["quality_replacement"] = quality_replacement_only or Config.QUALITY_REPLACEMENT quality_replacement = quality_replacement_only or Config.QUALITY_REPLACEMENT
has_embedding = is_embedding_available(entity_type) has_embedding = is_embedding_available()
limit, selection_mode = _resolve_strategy(strategy, has_embedding) limit, selection_mode = _resolve_strategy(strategy, has_embedding)
# Cap selection to remaining capacity (no cap when replacement-only — executor # Cap selection to remaining capacity (no cap when replacement-only — executor
# decides per-image whether to swap; any candidate could be an improvement). # decides per-image whether to swap; any candidate could be an improvement).
auto_cap = None
if not quality_replacement_only: if not quality_replacement_only:
if limit == "auto": if limit == "auto":
if already_uploaded > 0: if already_uploaded > 0:
auto_cap = capacity # Switch from open-ended auto to a fixed budget at remaining capacity
# so the diversity selector stops at the right count instead of
# selecting more than MAX_AUTO_IMAGES and overflowing the cap.
# First runs keep limit="auto" so FPS adaptive early-stop can fire.
limit = capacity
else: else:
limit = min(limit, capacity) limit = min(limit, capacity)
if selection_mode == "skip": if selection_mode == "skip":
continue continue
selected_assets = _perform_selection( retry_rejected = _getenv_bool("RETRY_REJECTED", False)
recent_assets, limit, name, selection_mode, entity_type, person_id=person["id"] before_dedup = len(recent_assets)
) new_asset_ids = set(filter_already_uploaded([a["id"] for a in recent_assets], retry_rejected=retry_rejected))
if auto_cap is not None: recent_assets = [a for a in recent_assets if a["id"] in new_asset_ids]
selected_assets = selected_assets[:auto_cap] skipped = before_dedup - len(recent_assets)
if skipped:
rprint(f" [dim]Skipped {skipped} assets already uploaded to Frigate.[/dim]")
if selected_assets: if not recent_assets:
rprint(f" [green]Queued {len(selected_assets)} images for {name}.[/green]") rprint(f" [dim]Skipping {name} (0 new images after dedup).[/dim]")
jobs.append({"person": person, "assets": selected_assets, "limit": len(selected_assets), "config": config}) continue
job = _build_job(person, recent_assets, limit, selection_mode, quality_replacement=quality_replacement)
if job is None:
rprint(f" [dim]Skipping {name} (0 images selected).[/dim]")
continue
rprint(f" [green]Queued {job['limit']} images for {name}.[/green]")
jobs.append(job)
return jobs return jobs
@@ -339,20 +351,17 @@ def _show_preview(jobs: list[dict]) -> None:
"""Show a summary table of all queued jobs before execution.""" """Show a summary table of all queued jobs before execution."""
table = Table(title="📋 Training Job Preview", show_header=True, header_style="bold cyan") table = Table(title="📋 Training Job Preview", show_header=True, header_style="bold cyan")
table.add_column("Person", style="bold") table.add_column("Person", style="bold")
table.add_column("Mode", style="dim")
table.add_column("Images", justify="right") table.add_column("Images", justify="right")
table.add_column("Date Range", style="dim") table.add_column("Date Range", style="dim")
for job in jobs: for job in jobs:
name = job["person"]["name"] name = job["person"]["name"]
mode = job["config"].get("mode", "face")
count = str(job["limit"]) count = str(job["limit"])
# Date range
dates = sorted(a.get("fileCreatedAt", "")[:10] for a in job["assets"] if a.get("fileCreatedAt")) dates = sorted(a.get("fileCreatedAt", "")[:10] for a in job["assets"] if a.get("fileCreatedAt"))
date_range = f"{dates[0]} → {dates[-1]}" if len(dates) >= 2 else (dates[0] if dates else "—") date_range = f"{dates[0]} → {dates[-1]}" if len(dates) >= 2 else (dates[0] if dates else "—")
table.add_row(name, mode, count, date_range) table.add_row(name, count, date_range)
console.print() console.print()
console.print(table) console.print(table)
-4
View File
@@ -13,12 +13,9 @@ console = Console()
NOISY_LOGGERS = ( NOISY_LOGGERS = (
"urllib3", "urllib3",
"PIL", "PIL",
"ultralytics",
"insightface", "insightface",
"onnxruntime", "onnxruntime",
"matplotlib", "matplotlib",
"transformers",
"torch",
) )
@@ -51,7 +48,6 @@ def setup_logging(verbose: bool = False) -> logging.Logger:
# Suppress Python warnings from ML libraries # Suppress Python warnings from ML libraries
warnings.filterwarnings("ignore", category=UserWarning, module="onnxruntime") warnings.filterwarnings("ignore", category=UserWarning, module="onnxruntime")
warnings.filterwarnings("ignore", category=FutureWarning, module="transformers")
return root return root
+29 -4
View File
@@ -14,6 +14,11 @@ from PIL import Image
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
def _laplacian_var(img_np: np.ndarray) -> float:
gray = cv2.cvtColor(img_np, cv2.COLOR_RGB2GRAY) if img_np.ndim == 3 else img_np
return float(cv2.Laplacian(gray, cv2.CV_64F).var())
@dataclass @dataclass
class QualityResult: class QualityResult:
"""Result of quality assessment on a face/image crop.""" """Result of quality assessment on a face/image crop."""
@@ -32,8 +37,7 @@ def check_blur(img_np: np.ndarray, threshold: float = 100.0) -> tuple[bool, str]
Lower variance = blurrier image. ArcFace needs clear facial features. Lower variance = blurrier image. ArcFace needs clear facial features.
""" """
gray = cv2.cvtColor(img_np, cv2.COLOR_RGB2GRAY) if img_np.ndim == 3 else img_np variance = _laplacian_var(img_np)
variance = cv2.Laplacian(gray, cv2.CV_64F).var()
if variance < threshold: if variance < threshold:
return False, f"Blurry (laplacian={variance:.1f}, threshold={threshold})" return False, f"Blurry (laplacian={variance:.1f}, threshold={threshold})"
return True, "" return True, ""
@@ -115,8 +119,7 @@ def assess_quality(
reasons = [] reasons = []
# Compute laplacian variance once (used by check_blur and stored as blur_score) # Compute laplacian variance once (used by check_blur and stored as blur_score)
gray = cv2.cvtColor(img_np, cv2.COLOR_RGB2GRAY) if img_np.ndim == 3 else img_np blur_score = _laplacian_var(img_np)
blur_score = float(cv2.Laplacian(gray, cv2.CV_64F).var())
checks = [ checks = [
( (
@@ -138,3 +141,25 @@ def assess_quality(
return QualityResult(passed=len(reasons) == 0, reasons=reasons, blur_score=blur_score) return QualityResult(passed=len(reasons) == 0, reasons=reasons, blur_score=blur_score)
def blur_score_from_image(img: Image.Image, max_dim: int = 1440) -> float | None:
"""Compute Laplacian-variance blur score, capped at max_dim px to normalise scale.
Caps resolution so full-res and thumbnail scores are comparable — Laplacian
variance grows with pixel count, making uncapped full-res scores much larger
than thumbnail scores for the same perceived sharpness.
Returns None on error so callers can distinguish a failed measurement from a
legitimately low (near-zero) score.
"""
try:
score_img = img.convert("RGB") if img.mode != "RGB" else img
if score_img.width > max_dim or score_img.height > max_dim:
if score_img is img:
score_img = score_img.copy()
score_img.thumbnail((max_dim, max_dim), Image.LANCZOS)
return _laplacian_var(np.array(score_img))
except Exception as exc:
logger.debug("blur_score_from_image failed: %s", exc)
return None
+137
View File
@@ -0,0 +1,137 @@
"""Frigate upload post-processing: reconciliation and asset enrichment."""
import logging
import time
from .frigate_api import get_frigate_person_files
from .immich_api import fetch_face_data
from .upload_tracker import record_frigate_files_batch
logger = logging.getLogger(__name__)
# Exponential back-off delays (seconds) when polling Frigate after uploads.
# Frigate processes the upload queue asynchronously, so files aren't
# immediately visible in GET /api/faces — we wait progressively longer
# rather than hammering the API.
_RECONCILE_POLL_DELAYS = (1, 2, 4, 8)
def reconcile_frigate_mappings(
person_name: str,
known_files_before: set[str],
uploaded: list[tuple[str, str | None]],
) -> None:
"""Map Frigate filenames to asset IDs after a batch of uploads.
Polls until all expected new files appear in the Frigate API, then maps
them to asset IDs by filename timestamp order (Frigate processes the
upload queue in FIFO order, so earlier uploads get earlier timestamps).
KNOWN LIMITATION — race condition with external uploads:
If another client uploads a face file for this person concurrently, the
count of new files will exceed `len(uploaded)` and we bail out entirely
(the "> target" branch). That's safe — we never record a wrong mapping —
but those uploads become permanently unmapped (they won't be eligible for
quality replacement). The right fix is a Frigate API that returns the
filename in the upload response, removing the need for any post-upload
diffing. Until then, the external-upload guard keeps mappings correct at
the cost of occasionally missing them when another client is active.
"""
target = len(uploaded)
new_files: set[str] = set()
# Check before the first sleep so a fast Frigate response returns immediately.
for delay in (None, *_RECONCILE_POLL_DELAYS):
if delay is not None:
time.sleep(delay)
fresh = get_frigate_person_files(person_name)
if fresh is None:
logger.warning(
"%s: Frigate API unreachable during mapping reconciliation"
" — quality replacement won't target these files",
person_name,
)
return
new_files = set(fresh) - known_files_before
if len(new_files) >= target:
break
if len(new_files) == target:
def _ts(fname: str) -> float:
try:
return float(fname.rsplit("_", 1)[-1].rsplit(".", 1)[0])
except (ValueError, IndexError):
return float("inf")
logger.debug(
"%s: mapping %s file(s) by filename timestamp — assumes Frigate processes"
" uploads in FIFO order; mapping may be wrong if that ever changes",
person_name,
target,
)
mappings = {
frigate_file: asset_id
for (_, asset_id), frigate_file in zip(uploaded, sorted(new_files, key=lambda f: (_ts(f), f)))
if asset_id
}
record_frigate_files_batch(person_name, mappings)
elif len(new_files) > target:
logger.warning(
"%s: %s new Frigate files for %s uploads"
" (external upload detected) — skipping file mapping;"
" these files are permanently unmapped",
person_name,
len(new_files),
target,
)
else:
logger.warning(
"%s: only %s of %s expected Frigate files"
" appeared after reconciliation — mapping skipped;"
" these files are permanently unmapped",
person_name,
len(new_files),
target,
)
def enrich_asset_with_face_data(asset: dict, person: dict) -> dict:
"""Enrich an asset dict with face bounding box data from the Immich faces API.
The search/metadata endpoint does not include face bounding box data,
so we fetch it from GET /api/faces?id={asset_id} and inject it into
the asset's "people" field so process_face_mode can find it.
Returns the enriched asset dict (modifies in place and returns it).
"""
person_id = person["id"]
face_data = fetch_face_data(asset["id"], person_id=person_id)
if face_data is None:
logger.debug("No face data returned for %s in asset %s", person.get("name"), asset.get("id"))
# Clean any None entries from the people list (can come from Immich API)
if "people" in asset:
asset["people"] = [p for p in asset["people"] if p is not None]
return asset
# Skip zero-area bounding boxes (face detection failed or no face found)
if face_data.bbox == (0, 0, 0, 0):
logger.debug("Zero-area bounding box for %s in asset %s", person.get("name"), asset.get("id"))
# Clean any None entries from the people list (can come from Immich API)
if "people" in asset:
asset["people"] = [p for p in asset["people"] if p is not None]
return asset
face_info = {
"boundingBoxX1": face_data.bbox[0],
"boundingBoxY1": face_data.bbox[1],
"boundingBoxX2": face_data.bbox[2],
"boundingBoxY2": face_data.bbox[3],
"imageWidth": face_data.image_width,
"imageHeight": face_data.image_height,
}
# Inject into asset so process_face_mode can find it via asset["people"]
asset["people"] = [{"id": person_id, "faces": [face_info]}]
asset["face_confidence"] = face_data.confidence
return asset
+215 -94
View File
@@ -1,6 +1,6 @@
"""Persistent tracker for Immich asset IDs already uploaded/rejected by Frigate. """Persistent tracker for Immich asset IDs already uploaded/rejected by Frigate.
Two separate JSON files in CACHE_DIR: Two separate JSON files in DATA_DIR:
frigate_uploaded_ids.json — successfully uploaded assets frigate_uploaded_ids.json — successfully uploaded assets
frigate_rejected_ids.json — assets Frigate rejected (e.g. no face detected) frigate_rejected_ids.json — assets Frigate rejected (e.g. no face detected)
@@ -32,47 +32,105 @@ import logging
import os import os
from pathlib import Path from pathlib import Path
from .frigate_api import delete_frigate_person_files from .frigate_api import _get_frigate_url, delete_frigate_person_files
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
UPLOAD_TRACKER_FILE = "frigate_uploaded_ids.json" UPLOAD_TRACKER_FILE = "frigate_uploaded_ids.json"
REJECT_TRACKER_FILE = "frigate_rejected_ids.json" REJECT_TRACKER_FILE = "frigate_rejected_ids.json"
# Write-through in-memory cache keyed by the resolved file path.
# Reduces per-call JSON reads from O(calls) to O(1) after the first load.
# Keyed by full path so tests with isolated tmp dirs never share entries.
_cache: dict[str, dict] = {}
_deferred: set[str] = set() # paths whose disk writes are batched until flush_batch()
_dirty: set[str] = set() # deferred paths that received at least one _save during the batch
def _tracker_path(filename: str) -> Path: def _tracker_path(filename: str) -> Path:
try: try:
from .config import Config from .config import Config
return Path(Config.CACHE_DIR) / filename return Path(Config.DATA_DIR) / filename
except (ImportError, AttributeError): except (ImportError, AttributeError):
return Path(filename) return Path(filename)
def _load(filename: str) -> dict: def _load(filename: str) -> dict:
path = _tracker_path(filename) path = _tracker_path(filename)
if not path.exists(): key = str(path)
return {} if key in _cache:
return _cache[key]
data: dict = {}
if path.exists():
try:
with open(path) as f:
data = json.load(f)
except (json.JSONDecodeError, OSError) as e:
logger.warning(f"Could not load tracker {filename}: {e}")
_cache[key] = data
return data
def _write_to_disk(path: Path, data: dict) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
tmp = path.with_suffix(".tmp")
try: try:
with open(path) as f: with open(tmp, "w") as f:
return json.load(f) json.dump(data, f, indent=2)
except (json.JSONDecodeError, OSError) as e: os.replace(tmp, path)
logger.warning(f"Could not load tracker {filename}: {e}") except Exception:
return {} tmp.unlink(missing_ok=True)
raise
def _save(filename: str, data: dict) -> None: def _save(filename: str, data: dict) -> None:
path = _tracker_path(filename) path = _tracker_path(filename)
path.parent.mkdir(parents=True, exist_ok=True) key = str(path)
with open(path, "w") as f: if key in _deferred:
json.dump(data, f, indent=2) _cache[key] = data # accumulate in cache; disk write deferred until flush_batch()
_dirty.add(key)
return
_write_to_disk(path, data)
_cache[key] = data # update cache only after successful write
def begin_batch(filename: str) -> None:
"""Defer tracker disk writes for filename. All _save calls accumulate in the
in-memory cache until flush_batch() is called. Use around per-person upload loops
to reduce N writes to 1.
If a previous batch for this file was interrupted before flush_batch() was called
(e.g. an exception escaped the upload loop), the leftover cache state is flushed
to disk here before starting fresh so that partial progress is not silently lost.
"""
path = _tracker_path(filename)
key = str(path)
if key in _deferred and key in _dirty:
try:
_write_to_disk(path, _cache[key])
except Exception:
logger.warning(
"begin_batch: could not flush leftover deferred state for %s"
" — partial progress may be lost",
path,
)
_deferred.discard(key)
_dirty.discard(key)
_deferred.add(key)
def flush_batch(filename: str) -> None:
"""Write the accumulated cache state for filename to disk."""
path = _tracker_path(filename)
key = str(path)
if key in _dirty and key in _cache:
_write_to_disk(path, _cache[key])
_deferred.discard(key)
_dirty.discard(key)
def _flat_key(filename: str) -> str: def _flat_key(filename: str) -> str:
return "uploaded_asset_ids" if "uploaded" in filename else "rejected_asset_ids" return "uploaded_asset_ids" if filename == UPLOAD_TRACKER_FILE else "rejected_asset_ids"
def _load_flat(filename: str) -> set[str]:
return set(_load(filename).get(_flat_key(filename), []))
def _get_ids(entry: list | dict) -> list[str]: def _get_ids(entry: list | dict) -> list[str]:
@@ -86,12 +144,14 @@ def _migrate_entry(entry: list | dict) -> dict:
"""Ensure by_person entry is in the current dict format.""" """Ensure by_person entry is in the current dict format."""
if isinstance(entry, list): if isinstance(entry, list):
return {"asset_ids": sorted(entry), "scores": {}, "frigate_scores": {}, "frigate_files": {}, "crop_dims": {}} return {"asset_ids": sorted(entry), "scores": {}, "frigate_scores": {}, "frigate_files": {}, "crop_dims": {}}
entry.setdefault("asset_ids", []) # Copy top-level and all nested dicts so callers' mutations never reach the cache.
entry.setdefault("scores", {}) result = dict(entry)
entry.setdefault("frigate_scores", {}) result["asset_ids"] = list(result.get("asset_ids", []))
entry.setdefault("frigate_files", {}) result["scores"] = dict(result.get("scores", {}))
entry.setdefault("crop_dims", {}) result["frigate_scores"] = dict(result.get("frigate_scores", {}))
return entry result["frigate_files"] = dict(result.get("frigate_files", {}))
result["crop_dims"] = dict(result.get("crop_dims", {}))
return result
def _mark( def _mark(
@@ -102,35 +162,46 @@ def _mark(
crop_dims: tuple[int, int] | None = None, crop_dims: tuple[int, int] | None = None,
frigate_score: float | None = None, frigate_score: float | None = None,
) -> None: ) -> None:
if not person_name:
logger.warning("_mark called with empty person_name for asset %s — asset not recorded", asset_id)
return
data = _load(filename) data = _load(filename)
flat_key = _flat_key(filename) by_person = dict(data.get("by_person", {}))
flat = set(data.get(flat_key, [])) entry = _migrate_entry(by_person.get(person_name, {}))
flat.add(asset_id) ids = set(entry["asset_ids"])
data[flat_key] = sorted(flat) ids.add(asset_id)
if person_name: entry["asset_ids"] = sorted(ids)
by_person = data.setdefault("by_person", {}) if score is not None:
entry = _migrate_entry(by_person.get(person_name, {})) entry["scores"][asset_id] = round(score, 4)
ids = set(entry["asset_ids"]) if crop_dims is not None:
ids.add(asset_id) entry["crop_dims"][asset_id] = [crop_dims[0], crop_dims[1]]
entry["asset_ids"] = sorted(ids) if frigate_score is not None:
if score is not None: entry["frigate_scores"][asset_id] = round(frigate_score, 4)
entry["scores"][asset_id] = round(score, 4) by_person[person_name] = entry
if crop_dims is not None: new_data = dict(data)
entry["crop_dims"][asset_id] = [crop_dims[0], crop_dims[1]] new_data["by_person"] = by_person
if frigate_score is not None: _save(filename, new_data)
entry["frigate_scores"][asset_id] = round(frigate_score, 4) logger.debug("Marked %s in %s (%s)", asset_id, filename, person_name)
by_person[person_name] = entry
_save(filename, data)
# ── Public API ──────────────────────────────────────────────────────────────── # ── Public API ────────────────────────────────────────────────────────────────
def load_uploaded_ids() -> set[str]: def load_uploaded_ids() -> set[str]:
return _load_flat(UPLOAD_TRACKER_FILE) """Return all asset IDs recorded as uploaded. Derives from by_person (primary)
plus any legacy flat list still present in old tracker files."""
data = _load(UPLOAD_TRACKER_FILE)
ids = {aid for e in data.get("by_person", {}).values() for aid in _get_ids(e)}
ids.update(data.get("uploaded_asset_ids", [])) # backward compat with pre-0.6.1 files
return ids
def load_rejected_ids() -> set[str]: def load_rejected_ids() -> set[str]:
return _load_flat(REJECT_TRACKER_FILE) """Return all asset IDs recorded as rejected. Derives from by_person (primary)
plus any legacy flat list still present in old tracker files."""
data = _load(REJECT_TRACKER_FILE)
ids = {aid for e in data.get("by_person", {}).values() for aid in _get_ids(e)}
ids.update(data.get("rejected_asset_ids", [])) # backward compat with pre-0.6.1 files
return ids
def mark_uploaded( def mark_uploaded(
@@ -141,24 +212,32 @@ def mark_uploaded(
frigate_score: float | None = None, frigate_score: float | None = None,
) -> None: ) -> None:
_mark(UPLOAD_TRACKER_FILE, asset_id, person_name, score=score, crop_dims=crop_dims, frigate_score=frigate_score) _mark(UPLOAD_TRACKER_FILE, asset_id, person_name, score=score, crop_dims=crop_dims, frigate_score=frigate_score)
logger.debug(f"Marked {asset_id} as uploaded ({person_name})")
def mark_rejected(asset_id: str, person_name: str | None = None) -> None: def mark_rejected(asset_id: str, person_name: str | None = None) -> None:
_mark(REJECT_TRACKER_FILE, asset_id, person_name) _mark(REJECT_TRACKER_FILE, asset_id, person_name)
logger.debug(f"Marked {asset_id} as rejected ({person_name})")
def record_frigate_file(person_name: str, frigate_filename: str, asset_id: str) -> None: def record_frigate_file(person_name: str, frigate_filename: str, asset_id: str) -> None:
"""Record the mapping from a Frigate training filename to an Immich asset ID.""" """Record a single Frigate filename → asset_id mapping."""
data = _load(UPLOAD_TRACKER_FILE) record_frigate_files_batch(person_name, {frigate_filename: asset_id})
by_person = data.setdefault("by_person", {})
def record_frigate_files_batch(person_name: str, mappings: dict[str, str]) -> None:
"""Record multiple Frigate filename → asset_id mappings in a single load/save."""
if not mappings:
return
src = _load(UPLOAD_TRACKER_FILE)
by_person = dict(src.get("by_person", {}))
entry = _migrate_entry(by_person.get(person_name, {})) entry = _migrate_entry(by_person.get(person_name, {}))
entry["frigate_files"][frigate_filename] = asset_id entry["frigate_files"].update(mappings)
by_person[person_name] = entry by_person[person_name] = entry
data = dict(src)
data["by_person"] = by_person
_save(UPLOAD_TRACKER_FILE, data) _save(UPLOAD_TRACKER_FILE, data)
logger.debug(f"Mapped Frigate file {frigate_filename} → {asset_id} ({person_name})") logger.debug(f"Batch-mapped {len(mappings)} Frigate file(s) for {person_name}")
def remove_frigate_file(person_name: str, frigate_filename: str) -> None: def remove_frigate_file(person_name: str, frigate_filename: str) -> None:
@@ -167,15 +246,26 @@ def remove_frigate_file(person_name: str, frigate_filename: str) -> None:
Does NOT unmark the source asset_id — the deletion was deliberate and Does NOT unmark the source asset_id — the deletion was deliberate and
we don't want to re-upload the inferior image on the next run. we don't want to re-upload the inferior image on the next run.
""" """
data = _load(UPLOAD_TRACKER_FILE) remove_frigate_files_batch(person_name, [frigate_filename])
by_person = data.get("by_person", {})
entry = _migrate_entry(by_person.get(person_name, {}))
asset_id = entry["frigate_files"].pop(frigate_filename, None) def remove_frigate_files_batch(person_name: str, frigate_filenames: list[str]) -> None:
if asset_id: """Remove multiple Frigate filenames in a single load/save."""
entry["frigate_scores"].pop(asset_id, None) src = _load(UPLOAD_TRACKER_FILE)
raw = src.get("by_person", {}).get(person_name)
if raw is None:
return
entry = _migrate_entry(raw)
for fn in frigate_filenames:
asset_id = entry["frigate_files"].pop(fn, None)
if asset_id is not None and asset_id not in entry["frigate_files"].values():
entry["frigate_scores"].pop(asset_id, None)
by_person = dict(src.get("by_person", {})) # copy so assignment does not mutate the cache
by_person[person_name] = entry by_person[person_name] = entry
data = dict(src)
data["by_person"] = by_person
_save(UPLOAD_TRACKER_FILE, data) _save(UPLOAD_TRACKER_FILE, data)
logger.debug(f"Removed Frigate file mapping {frigate_filename} ({person_name})") logger.debug(f"Removed {len(frigate_filenames)} Frigate file mapping(s) for {person_name}")
def get_tracked_frigate_file_count(person_name: str) -> int: def get_tracked_frigate_file_count(person_name: str) -> int:
@@ -203,9 +293,11 @@ def get_tracked_frigate_filenames(person_name: str) -> set[str]:
def has_frigate_scores(person_name: str) -> bool: def has_frigate_scores(person_name: str) -> bool:
"""Return True if any mapped file for this person has a stored Frigate recognition score.""" """Return True if any mapped file for this person has a stored Frigate recognition score."""
data = _load(UPLOAD_TRACKER_FILE) data = _load(UPLOAD_TRACKER_FILE)
entry = _migrate_entry(data.get("by_person", {}).get(person_name, {})) raw = data.get("by_person", {}).get(person_name)
frigate_files = entry.get("frigate_files", {}) if not raw or isinstance(raw, list):
frigate_scores = entry.get("frigate_scores", {}) return False
frigate_files = raw.get("frigate_files", {})
frigate_scores = raw.get("frigate_scores", {})
return any(asset_id in frigate_scores for asset_id in frigate_files.values()) return any(asset_id in frigate_scores for asset_id in frigate_files.values())
@@ -215,11 +307,12 @@ def _pick_mapped_file(
data = _load(UPLOAD_TRACKER_FILE) data = _load(UPLOAD_TRACKER_FILE)
entry = _migrate_entry(data.get("by_person", {}).get(person_name, {})) entry = _migrate_entry(data.get("by_person", {}).get(person_name, {}))
scores = entry.get(score_key, {}) scores = entry.get(score_key, {})
candidates = [ seen_assets: set[str] = set()
(ff, asset_id, scores[asset_id]) candidates = []
for ff, asset_id in entry.get("frigate_files", {}).items() for ff, asset_id in entry.get("frigate_files", {}).items():
if (exclude is None or ff not in exclude) and asset_id in scores if (exclude is None or ff not in exclude) and asset_id in scores and asset_id not in seen_assets:
] seen_assets.add(asset_id)
candidates.append((ff, asset_id, scores[asset_id]))
if not candidates: if not candidates:
return None return None
return max(candidates, key=lambda x: x[2]) if highest else min(candidates, key=lambda x: x[2]) return max(candidates, key=lambda x: x[2]) if highest else min(candidates, key=lambda x: x[2])
@@ -250,16 +343,6 @@ def get_most_redundant_mapped_file(
return _pick_mapped_file(person_name, "frigate_scores", highest=True, exclude=exclude) return _pick_mapped_file(person_name, "frigate_scores", highest=True, exclude=exclude)
def get_frigate_filename_for_asset(person_name: str, asset_id: str) -> str | None:
"""Return the Frigate training filename mapped to this asset ID, or None."""
data = _load(UPLOAD_TRACKER_FILE)
entry = _migrate_entry(data.get("by_person", {}).get(person_name, {}))
for frigate_filename, aid in entry["frigate_files"].items():
if aid == asset_id:
return frigate_filename
return None
def find_by_crop_dimension(size: int) -> list[dict]: def find_by_crop_dimension(size: int) -> list[dict]:
"""Return all tracked crops whose width or height matches `size` pixels. """Return all tracked crops whose width or height matches `size` pixels.
@@ -272,9 +355,13 @@ def find_by_crop_dimension(size: int) -> list[dict]:
entry = _migrate_entry(raw_entry) entry = _migrate_entry(raw_entry)
scores = entry.get("scores", {}) scores = entry.get("scores", {})
frigate_files = entry.get("frigate_files", {}) frigate_files = entry.get("frigate_files", {})
asset_to_frigate = {v: k for k, v in frigate_files.items()} asset_to_frigate: dict[str, str] = {}
for fn, aid in frigate_files.items():
asset_to_frigate.setdefault(aid, fn) # first-seen wins; plain inversion silently drops duplicates
frigate_scores = entry.get("frigate_scores", {}) frigate_scores = entry.get("frigate_scores", {})
for asset_id, dims in entry.get("crop_dims", {}).items(): for asset_id, dims in entry.get("crop_dims", {}).items():
if not isinstance(dims, (list, tuple)) or len(dims) < 2:
continue
w, h = dims[0], dims[1] w, h = dims[0], dims[1]
if w == size or h == size: if w == size or h == size:
results.append({ results.append({
@@ -292,11 +379,38 @@ def find_by_crop_dimension(size: int) -> list[dict]:
def update_frigate_count(person_name: str, count: int) -> None: def update_frigate_count(person_name: str, count: int) -> None:
"""Record Frigate's authoritative training image count for a person.""" """Record Frigate's authoritative training image count for a person."""
data = _load(UPLOAD_TRACKER_FILE) data = _load(UPLOAD_TRACKER_FILE)
by_person = data.setdefault("by_person", {}) by_person = dict(data.get("by_person", {}))
entry = _migrate_entry(by_person.get(person_name, {})) entry = _migrate_entry(by_person.get(person_name, {}))
entry["frigate_count"] = count entry["frigate_count"] = count
by_person[person_name] = entry by_person[person_name] = entry
_save(UPLOAD_TRACKER_FILE, data) new_data = dict(data)
new_data["by_person"] = by_person
_save(UPLOAD_TRACKER_FILE, new_data)
def reset_all_people() -> None:
"""Reset all tracking data in two writes (O(P) Frigate API calls, O(1) disk writes).
Preferred over calling reset_person() in a loop when RESET_PERSON=* — that
approach is O(P²) because each call rebuilds the flat list from all remaining entries.
"""
upload_data = _load(UPLOAD_TRACKER_FILE)
frigate_url = _get_frigate_url()
if not frigate_url:
logger.info("FRIGATE_URL not set — skipping Frigate file deletion")
for person_name, raw_entry in upload_data.get("by_person", {}).items():
entry = _migrate_entry(raw_entry)
frigate_filenames = list(entry.get("frigate_files", {}).keys())
if not frigate_filenames:
continue
if frigate_url:
if delete_frigate_person_files(person_name, frigate_filenames):
logger.info(f"Deleted {len(frigate_filenames)} Frigate file(s) for {person_name}")
else:
logger.warning(f"Could not delete Frigate files for {person_name} — tracker reset proceeding anyway")
_save(UPLOAD_TRACKER_FILE, {})
_save(REJECT_TRACKER_FILE, {})
logger.info("Reset all tracking data")
def reset_person(person_name: str) -> None: def reset_person(person_name: str) -> None:
@@ -311,7 +425,7 @@ def reset_person(person_name: str) -> None:
entry = _migrate_entry(upload_data.get("by_person", {}).get(person_name, {})) entry = _migrate_entry(upload_data.get("by_person", {}).get(person_name, {}))
frigate_filenames = list(entry.get("frigate_files", {}).keys()) frigate_filenames = list(entry.get("frigate_files", {}).keys())
if frigate_filenames: if frigate_filenames:
if not os.environ.get("FRIGATE_URL", "").strip(): if not _get_frigate_url():
logger.info(f"FRIGATE_URL not set — skipping Frigate file deletion for {person_name}") logger.info(f"FRIGATE_URL not set — skipping Frigate file deletion for {person_name}")
elif delete_frigate_person_files(person_name, frigate_filenames): elif delete_frigate_person_files(person_name, frigate_filenames):
logger.info(f"Deleted {len(frigate_filenames)} Frigate file(s) for {person_name}") logger.info(f"Deleted {len(frigate_filenames)} Frigate file(s) for {person_name}")
@@ -319,16 +433,23 @@ def reset_person(person_name: str) -> None:
logger.warning(f"Could not delete Frigate files for {person_name} — tracker reset proceeding anyway") logger.warning(f"Could not delete Frigate files for {person_name} — tracker reset proceeding anyway")
changed = False changed = False
tracker_files = ((UPLOAD_TRACKER_FILE, upload_data), (REJECT_TRACKER_FILE, _load(REJECT_TRACKER_FILE))) for filename in (UPLOAD_TRACKER_FILE, REJECT_TRACKER_FILE):
for filename, data in tracker_files: src = upload_data if filename == UPLOAD_TRACKER_FILE else _load(REJECT_TRACKER_FILE)
flat_key = _flat_key(filename) by_person = dict(src.get("by_person", {})) # copy so pop() does not mutate the cache
by_person = data.get("by_person", {})
tracker_entry = by_person.pop(person_name, None) tracker_entry = by_person.pop(person_name, None)
if tracker_entry is not None: if tracker_entry is not None:
person_ids = set(_get_ids(tracker_entry)) data = dict(src)
flat = set(data.get(flat_key, [])) - person_ids
data[flat_key] = sorted(flat)
data["by_person"] = by_person data["by_person"] = by_person
flat_key = _flat_key(filename)
person_ids = set(_get_ids(tracker_entry))
if person_ids and flat_key in data and not isinstance(data[flat_key], list):
logger.warning(
"reset_person: %s has unexpected type for %s (%s) — skipping flat-list cleanup;"
" all persons' legacy IDs in this field are unaffected but unreadable",
filename, flat_key, type(data[flat_key]).__name__,
)
elif person_ids and flat_key in data:
data[flat_key] = sorted(set(data[flat_key]) - person_ids)
_save(filename, data) _save(filename, data)
changed = True changed = True
if changed: if changed:
@@ -344,14 +465,14 @@ def get_person_summary() -> dict[str, dict]:
names = set(uploaded_data) | set(rejected_data) names = set(uploaded_data) | set(rejected_data)
result = {} result = {}
for name in sorted(names): for name in sorted(names):
u_entry = uploaded_data.get(name, {}) u_entry = _migrate_entry(uploaded_data.get(name, {}))
r_entry = rejected_data.get(name, {}) r_entry = _migrate_entry(rejected_data.get(name, {}))
result[name] = { result[name] = {
"uploaded": len(_get_ids(u_entry)), "uploaded": len(u_entry["asset_ids"]),
"rejected": len(_get_ids(r_entry)), "rejected": len(r_entry["asset_ids"]),
"frigate_count": u_entry.get("frigate_count") if isinstance(u_entry, dict) else None, "frigate_count": u_entry.get("frigate_count"),
"scores": u_entry.get("scores", {}) if isinstance(u_entry, dict) else {}, "scores": u_entry["scores"],
"frigate_files": u_entry.get("frigate_files", {}) if isinstance(u_entry, dict) else {}, "frigate_files": u_entry["frigate_files"],
} }
return result return result