Tests cover the core ML pipeline algorithms in diversity.py — previously
untested. No network or model dependencies; all pure-function or
numpy-only paths:
- Face bbox and confidence extraction from Immich metadata, including
person_id filtering and missing-data edge cases
- Face crop scaling: verifies bbox coordinates are correctly scaled when
the thumbnail dimensions differ from the metadata image dimensions
- Near-duplicate dedup: removal below cosine threshold, quality-score
preference between duplicates, zero-quality-score treated as zero not
missing (falsy bug guard)
- K-Medoids: correct medoid count, distinctness, valid index range, and
full-N edge case
- Adaptive threshold: positive output, floor at 0.05 for identical
embeddings, single-point, scales with embedding spread
- Time-spread fallback: exact count, all-under-limit passthrough,
auto→30 default, first/last inclusion
- Cluster-aware selection: exact limit, subset invariant, auto-stop on
tight cluster, hard-example confidence weighting accepted