Compare commits

..
6 Commits
Author SHA1 Message Date
flanandClaude Sonnet 4.6 326fdbdf38 release: merge dev → main for 0.4.1
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 19:58:19 +00:00
flanandClaude Sonnet 4.6 405413490b fix: RESET_PERSON now deletes managed Frigate files before clearing tracker
Previously reset_person wiped the local tracker but left existing Frigate
training files as orphans, causing the next run to upload a full new batch
on top of them. Now deletes all winnow-managed files from Frigate first so
the next run starts truly clean. Manually-added Frigate files are never
touched.

Also fixes a spurious warning when FRIGATE_URL is unset: the deletion step
is now skipped at info level rather than logging a misleading error. Moves
the deferred import to top-level and eliminates a double disk read.

Bumps to 0.4.1. Also fixes ruff lint violations in executor.py (import
sort, line length) and promotes the "winnow only touches files it uploaded"
callout to the README intro.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 19:58:08 +00:00
flanandClaude Sonnet 4.6 91e0858aa6 refactor: merge dev → main — post-0.4.0 cleanup
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 18:55:42 +00:00
flanandClaude Sonnet 4.6 e67f2d9638 refactor: cleanup audit findings — dedup helpers, prune orphan scores, cache has_frigate_scores
- upload_tracker: extract _pick_mapped_file() private helper; get_lowest_quality_mapped_file
  and get_most_redundant_mapped_file are now one-liners over the same body
- upload_tracker: remove_frigate_file now also prunes the corresponding frigate_scores entry,
  preventing unbounded accumulation of orphaned score entries across replacement cycles
- frigate_api: get_frigate_face_counts delegates to get_all_frigate_person_files, eliminating
  the duplicated "name != 'train' and isinstance(files, list)" filter body
- executor: cache has_frigate_scores(name) as person_has_fscores before the per-file loop;
  refresh it after each remove_frigate_file call and after each scored upload, eliminating
  two redundant disk reads per at-cap file iteration
- executor: casefold() both sides of the recognize_face person-name comparison so a Frigate
  casing normalization or manual-registration casing mismatch does not silently suppress scoring

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-13 18:55:18 +00:00
github-actions[bot] 71df0e81de chore: update lockfiles 2026-06-13 18:36:36 +00:00
github-actions[bot] bbbac18207 chore: update lockfiles 2026-06-13 18:35:54 +00:00
7 changed files with 95 additions and 64 deletions
+7
View File
@@ -7,6 +7,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased] ## [Unreleased]
## [0.4.1] - 2026-06-13
### Fixed
- **`RESET_PERSON` no longer creates duplicate Frigate files**: previously, resetting a person only wiped the local tracker — existing Frigate training files were left as unmanaged orphans, causing the next run to upload a full new batch on top of them. `reset_person` now deletes all winnow-managed files for that person from Frigate before clearing the tracker. Manually-added Frigate files are unaffected.
- **No spurious warning when `FRIGATE_URL` is unset and `RESET_PERSON` is used**: the deletion step is now skipped silently at info level rather than logging a misleading "could not delete" warning.
## [0.4.0] - 2026-06-13 ## [0.4.0] - 2026-06-13
### Added ### Added
+3 -1
View File
@@ -11,6 +11,8 @@
Frigate's face recognition is only as good as its training data — and the key quality metric is **diversity**, not volume. A hundred photos from the same week teach the model one lighting condition. What you need is a spread: different years, different angles, different lighting, different contexts. Your photo library already has that data. winnow finds and delivers the right subset automatically. Frigate's face recognition is only as good as its training data — and the key quality metric is **diversity**, not volume. A hundred photos from the same week teach the model one lighting condition. What you need is a spread: different years, different angles, different lighting, different contexts. Your photo library already has that data. winnow finds and delivers the right subset automatically.
> **winnow only touches files it uploaded.** Faces added to Frigate manually through its UI are never deleted, replaced, or modified — not by quality replacement, not by `RESET_PERSON`, not by stale cleanup. If you have a curated training set you want to keep, it is safe.
--- ---
## How It Works ## How It Works
@@ -222,7 +224,7 @@ In scheduled mode the process (and loaded models) stays resident between runs. T
| :--- | :--- | :--- | | :--- | :--- | :--- |
| `DRY_RUN` | `false` | Preview selection without downloading or uploading | | `DRY_RUN` | `false` | Preview selection without downloading or uploading |
| `RETRY_REJECTED` | `false` | Re-attempt assets previously rejected by Frigate | | `RETRY_REJECTED` | `false` | Re-attempt assets previously rejected by Frigate |
| `RESET_PERSON` | *(unset)* | Clear upload and rejection history for one person by name | | `RESET_PERSON` | *(unset)* | Clear upload history for one person and delete their winnow-managed Frigate training files so the next run starts fresh. Manually added Frigate files are never touched |
### Scheduling ### Scheduling
+1 -1
View File
@@ -1,6 +1,6 @@
[project] [project]
name = "winnow" name = "winnow"
version = "0.4.0" version = "0.4.1"
description = "Selects diverse, high-quality photos from Immich as training data for Frigate face recognition and object classification." description = "Selects diverse, high-quality photos from Immich as training data for Frigate face recognition and object classification."
license = "AGPL-3.0-or-later" license = "AGPL-3.0-or-later"
requires-python = ">=3.13" requires-python = ">=3.13"
Generated
+1 -1
View File
@@ -2348,7 +2348,7 @@ wheels = [
[[package]] [[package]]
name = "winnow" name = "winnow"
version = "0.3.3" version = "0.4.0"
source = { editable = "." } source = { editable = "." }
dependencies = [ dependencies = [
{ name = "croniter" }, { name = "croniter" },
+16 -5
View File
@@ -13,7 +13,12 @@ from rich import print as rprint
from rich.progress import BarColumn, Progress, SpinnerColumn, TaskProgressColumn, TextColumn from rich.progress import BarColumn, Progress, SpinnerColumn, TaskProgressColumn, TextColumn
from .config import Config, get_headers from .config import Config, get_headers
from .frigate_api import delete_frigate_person_files, get_all_frigate_person_files, get_frigate_person_files, recognize_face from .frigate_api import (
delete_frigate_person_files,
get_all_frigate_person_files,
get_frigate_person_files,
recognize_face,
)
from .image_processing import process_face_mode, process_full_mode, process_object_mode from .image_processing import process_face_mode, process_full_mode, process_object_mode
from .immich_api import fetch_face_data, fetch_full_image from .immich_api import fetch_face_data, fetch_full_image
from .log_config import console from .log_config import console
@@ -395,6 +400,7 @@ def upload_to_frigate(jobs: list[dict]) -> None:
actually_uploaded: list[tuple[str, str | None]] = [] actually_uploaded: list[tuple[str, str | None]] = []
failed_deletes: set[str] = set() failed_deletes: set[str] = set()
min_quality_score_for_slot: float | None = None min_quality_score_for_slot: float | None = None
person_has_fscores: bool = has_frigate_scores(name)
for fname in person_files: for fname in person_files:
fpath = os.path.join(person_dir, fname) fpath = os.path.join(person_dir, fname)
@@ -427,9 +433,9 @@ def upload_to_frigate(jobs: list[dict]) -> None:
# handles this conservatively by skipping that candidate until the next run. # handles this conservatively by skipping that candidate until the next run.
pre_fscore: float | None = None pre_fscore: float | None = None
if Config.ENABLE_FRIGATE_SCORES and pre_run_count > 0: if Config.ENABLE_FRIGATE_SCORES and pre_run_count > 0:
if not at_cap or has_frigate_scores(name): if not at_cap or person_has_fscores:
_result = recognize_face(fpath) _result = recognize_face(fpath)
if _result is not None and _result[0] == name: if _result is not None and (_result[0] or "").casefold() == name.casefold():
pre_fscore = _result[1] pre_fscore = _result[1]
# Ceiling check: skip if the existing training set already covers this # Ceiling check: skip if the existing training set already covers this
@@ -450,7 +456,7 @@ def upload_to_frigate(jobs: list[dict]) -> None:
progress.advance(upload_task) progress.advance(upload_task)
continue continue
using_fscore = has_frigate_scores(name) and Config.ENABLE_FRIGATE_SCORES using_fscore = person_has_fscores and Config.ENABLE_FRIGATE_SCORES
if using_fscore: if using_fscore:
candidate_score = pre_fscore candidate_score = pre_fscore
if candidate_score is None: if candidate_score is None:
@@ -476,8 +482,10 @@ def upload_to_frigate(jobs: list[dict]) -> None:
) )
if delete_frigate_person_files(name, [target_frigate_file]): if delete_frigate_person_files(name, [target_frigate_file]):
remove_frigate_file(name, target_frigate_file) remove_frigate_file(name, target_frigate_file)
person_has_fscores = has_frigate_scores(name)
effective_count -= 1 effective_count -= 1
min_quality_score_for_slot = None # clear any blur-mode slot floor — Frigate uses a different score metric # clear any blur-mode slot floor — Frigate uses a different score metric
min_quality_score_for_slot = None
else: else:
logger.warning(f"Failed to delete {target_frigate_file} for {name}, skipping replacement") logger.warning(f"Failed to delete {target_frigate_file} for {name}, skipping replacement")
failed_deletes.add(target_frigate_file) failed_deletes.add(target_frigate_file)
@@ -507,6 +515,7 @@ def upload_to_frigate(jobs: list[dict]) -> None:
) )
if delete_frigate_person_files(name, [target_frigate_file]): if delete_frigate_person_files(name, [target_frigate_file]):
remove_frigate_file(name, target_frigate_file) remove_frigate_file(name, target_frigate_file)
person_has_fscores = has_frigate_scores(name)
effective_count -= 1 effective_count -= 1
min_quality_score_for_slot = score_map.get(fname) min_quality_score_for_slot = score_map.get(fname)
else: else:
@@ -538,6 +547,8 @@ def upload_to_frigate(jobs: list[dict]) -> None:
crop_dims=dims_map.get(fname), crop_dims=dims_map.get(fname),
frigate_score=pre_fscore, frigate_score=pre_fscore,
) )
if pre_fscore is not None:
person_has_fscores = True
actually_uploaded.append((fname, asset_id)) actually_uploaded.append((fname, asset_id))
break break
+14 -18
View File
@@ -22,24 +22,6 @@ def _get_faces_data() -> dict | None:
return None return None
def get_frigate_face_counts() -> dict[str, int] | None:
"""Return {person_name: training_image_count} from Frigate's train directory.
Returns None if FRIGATE_URL is not set or the API is unreachable, so callers
can distinguish "API unavailable" from "person has 0 images."
"""
data = _get_faces_data()
if data is None:
return None
# Response: {person_name: [file, ...], "train": [...], ...}
# "train" is a flat pending list, not a person — skip it.
return {
name: len(files)
for name, files in data.items()
if name != "train" and isinstance(files, list)
}
def get_all_frigate_person_files() -> dict[str, list[str]] | None: def get_all_frigate_person_files() -> dict[str, list[str]] | None:
"""Return {person_name: [filename, ...]} for every person in Frigate. """Return {person_name: [filename, ...]} for every person in Frigate.
@@ -49,6 +31,8 @@ def get_all_frigate_person_files() -> dict[str, list[str]] | None:
data = _get_faces_data() data = _get_faces_data()
if data is None: if data is None:
return None return None
# Response: {person_name: [file, ...], "train": [...], ...}
# "train" is a flat pending list, not a person — skip it.
return { return {
name: files name: files
for name, files in data.items() for name, files in data.items()
@@ -56,6 +40,18 @@ def get_all_frigate_person_files() -> dict[str, list[str]] | None:
} }
def get_frigate_face_counts() -> dict[str, int] | None:
"""Return {person_name: training_image_count} from Frigate's train directory.
Returns None if FRIGATE_URL is not set or the API is unreachable, so callers
can distinguish "API unavailable" from "person has 0 images."
"""
all_files = get_all_frigate_person_files()
if all_files is None:
return None
return {name: len(files) for name, files in all_files.items()}
def get_frigate_person_files(person_name: str) -> list[str] | None: def get_frigate_person_files(person_name: str) -> list[str] | None:
"""Return the list of training filenames for a person in Frigate. """Return the list of training filenames for a person in Frigate.
+53 -38
View File
@@ -29,8 +29,11 @@ Frigate's UI are never mapped here and are never touched by quality replacement.
import json import json
import logging import logging
import os
from pathlib import Path from pathlib import Path
from .frigate_api import delete_frigate_person_files
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
UPLOAD_TRACKER_FILE = "frigate_uploaded_ids.json" UPLOAD_TRACKER_FILE = "frigate_uploaded_ids.json"
@@ -167,7 +170,9 @@ def remove_frigate_file(person_name: str, frigate_filename: str) -> None:
data = _load(UPLOAD_TRACKER_FILE) data = _load(UPLOAD_TRACKER_FILE)
by_person = data.get("by_person", {}) by_person = data.get("by_person", {})
entry = _migrate_entry(by_person.get(person_name, {})) entry = _migrate_entry(by_person.get(person_name, {}))
entry["frigate_files"].pop(frigate_filename, None) asset_id = entry["frigate_files"].pop(frigate_filename, None)
if asset_id:
entry["frigate_scores"].pop(asset_id, None)
by_person[person_name] = entry by_person[person_name] = entry
_save(UPLOAD_TRACKER_FILE, data) _save(UPLOAD_TRACKER_FILE, data)
logger.debug(f"Removed Frigate file mapping {frigate_filename} ({person_name})") logger.debug(f"Removed Frigate file mapping {frigate_filename} ({person_name})")
@@ -204,6 +209,22 @@ def has_frigate_scores(person_name: str) -> bool:
return any(asset_id in frigate_scores for asset_id in frigate_files.values()) return any(asset_id in frigate_scores for asset_id in frigate_files.values())
def _pick_mapped_file(
person_name: str, score_key: str, *, highest: bool, exclude: set[str] | None = None
) -> tuple[str, str, float] | None:
data = _load(UPLOAD_TRACKER_FILE)
entry = _migrate_entry(data.get("by_person", {}).get(person_name, {}))
scores = entry.get(score_key, {})
candidates = [
(ff, asset_id, scores[asset_id])
for ff, asset_id in entry.get("frigate_files", {}).items()
if (exclude is None or ff not in exclude) and asset_id in scores
]
if not candidates:
return None
return max(candidates, key=lambda x: x[2]) if highest else min(candidates, key=lambda x: x[2])
def get_lowest_quality_mapped_file( def get_lowest_quality_mapped_file(
person_name: str, exclude: set[str] | None = None person_name: str, exclude: set[str] | None = None
) -> tuple[str, str, float] | None: ) -> tuple[str, str, float] | None:
@@ -213,21 +234,7 @@ def get_lowest_quality_mapped_file(
Used for quality replacement when no Frigate scores are available. Used for quality replacement when no Frigate scores are available.
Pass `exclude` to skip files that failed to delete this run. Pass `exclude` to skip files that failed to delete this run.
""" """
data = _load(UPLOAD_TRACKER_FILE) return _pick_mapped_file(person_name, "scores", highest=False, exclude=exclude)
entry = _migrate_entry(data.get("by_person", {}).get(person_name, {}))
frigate_files = entry.get("frigate_files", {})
blur_scores = entry.get("scores", {})
candidates = [
(ff, asset_id, blur_scores[asset_id])
for ff, asset_id in frigate_files.items()
if (exclude is None or ff not in exclude)
and asset_id in blur_scores
]
if not candidates:
return None
return min(candidates, key=lambda x: x[2])
def get_most_redundant_mapped_file( def get_most_redundant_mapped_file(
@@ -240,21 +247,7 @@ def get_most_redundant_mapped_file(
= the most redundant file and therefore the best replacement target. = the most redundant file and therefore the best replacement target.
Pass `exclude` to skip files that failed to delete this run. Pass `exclude` to skip files that failed to delete this run.
""" """
data = _load(UPLOAD_TRACKER_FILE) return _pick_mapped_file(person_name, "frigate_scores", highest=True, exclude=exclude)
entry = _migrate_entry(data.get("by_person", {}).get(person_name, {}))
frigate_files = entry.get("frigate_files", {})
frigate_scores = entry.get("frigate_scores", {})
candidates = [
(ff, asset_id, frigate_scores[asset_id])
for ff, asset_id in frigate_files.items()
if (exclude is None or ff not in exclude)
and asset_id in frigate_scores
]
if not candidates:
return None
return max(candidates, key=lambda x: x[2])
def get_frigate_filename_for_asset(person_name: str, asset_id: str) -> str | None: def get_frigate_filename_for_asset(person_name: str, asset_id: str) -> str | None:
@@ -307,19 +300,41 @@ def update_frigate_count(person_name: str, count: int) -> None:
def reset_person(person_name: str) -> None: def reset_person(person_name: str) -> None:
"""Remove all uploaded and rejected records for a given person.""" """Remove all uploaded and rejected records for a given person.
for filename in (UPLOAD_TRACKER_FILE, REJECT_TRACKER_FILE):
data = _load(filename) Also deletes winnow-managed Frigate training files so the next run starts
clean rather than uploading on top of orphaned files. Manually-added Frigate
files (not in frigate_files) are never touched. Proceeds with tracker reset
even if Frigate is unreachable.
"""
upload_data = _load(UPLOAD_TRACKER_FILE)
entry = _migrate_entry(upload_data.get("by_person", {}).get(person_name, {}))
frigate_filenames = list(entry.get("frigate_files", {}).keys())
if frigate_filenames:
if not os.environ.get("FRIGATE_URL", "").strip():
logger.info(f"FRIGATE_URL not set — skipping Frigate file deletion for {person_name}")
elif delete_frigate_person_files(person_name, frigate_filenames):
logger.info(f"Deleted {len(frigate_filenames)} Frigate file(s) for {person_name}")
else:
logger.warning(f"Could not delete Frigate files for {person_name} — tracker reset proceeding anyway")
changed = False
tracker_files = ((UPLOAD_TRACKER_FILE, upload_data), (REJECT_TRACKER_FILE, _load(REJECT_TRACKER_FILE)))
for filename, data in tracker_files:
flat_key = _flat_key(filename) flat_key = _flat_key(filename)
by_person = data.get("by_person", {}) by_person = data.get("by_person", {})
entry = by_person.pop(person_name, None) tracker_entry = by_person.pop(person_name, None)
if entry is not None: if tracker_entry is not None:
person_ids = set(_get_ids(entry)) person_ids = set(_get_ids(tracker_entry))
flat = set(data.get(flat_key, [])) - person_ids flat = set(data.get(flat_key, [])) - person_ids
data[flat_key] = sorted(flat) data[flat_key] = sorted(flat)
data["by_person"] = by_person data["by_person"] = by_person
_save(filename, data) _save(filename, data)
logger.info(f"Reset tracking data for {person_name}") changed = True
if changed:
logger.info(f"Reset tracking data for {person_name}")
else:
logger.debug(f"reset_person: no tracking data found for {person_name}")
def get_person_summary() -> dict[str, dict]: def get_person_summary() -> dict[str, dict]: