test_cut_pending_isolates_per_episode_failures matched a bare "2" against
the full audio file path to target episode 2's failure — but pytest's
auto-numbered tmp_path can itself contain that digit, so the test failed
depending on run order. Now matches the deterministic per-episode filename.
README/CLAUDE.md/PLAN.md still described the hark relationship as an open
question; M5 was actually decided (library dependency, not a merge) — bring
them in line with what hark's own docs already say.