Project brief, plan, and seed show list

This commit is contained in:
flan
2026-07-10 14:48:05 +00:00
commit 0925c6e37b
3 changed files with 108 additions and 0 deletions
+49
View File
@@ -0,0 +1,49 @@
# hark
Cross-podcast topic index and discovery service. Working title "hark" — renaming is cheap,
don't get attached.
## What this is
A homelab web service (NOT a mobile app, NOT an AntennaPod fork) that:
1. **Topic index (first milestone):** resolves podcast episodes in subject-per-episode genres
(true crime, history, disasters, scams/fraud, biographies, espionage, cults, mysteries) to
the real-world case/event/person they cover, so you can ask "who covered the Dyatlov Pass
incident?" and compare treatments across shows.
2. **Discovery:** related-show and notable-episode recommendations via topic/embedding
similarity.
3. **Episode scoring (later):** metric-based interestingness ratings, tiltmeter-style
(auditable, defined metrics, calibrated against the owner's actual listening).
Origin: ideas #2 and #3 in a private ideas repo (git.onetick.ninja) — read
`~/project-ideas/README.md` for the full assessments and reasoning.
## Architecture decisions (already made — don't relitigate)
- Standalone service on the homelab, shaped like tiltmeter: scheduled ingest → pipeline →
SQLite → API/UI. The owner's player stays AntennaPod.
- **Input integration:** AntennaPod syncs subscriptions + play history to Nextcloud (gpodder
sync app) on truenas; hark reads from that API. OPML import as fallback. (Not wired in M0.)
- **Output integration:** hark generates custom RSS feeds (e.g. "top episodes about topics
you like", "best of candidate shows") that get subscribed to in AntennaPod like any podcast.
No app modification anywhere.
- Feed URLs resolve via the keyless iTunes Search API; Podcast Index API can be added later
(needs a registered key). Episode metadata comes from plain RSS.
- Topic extraction: LLM extraction from episode title/description, canonicalized against
Wikidata. Transcripts/Whisper are explicitly OUT of scope until much later — these genres
name their subject in the metadata.
- Topics can belong to multiple genres (Titanic = history + disaster); never force one bucket.
## Conventions
- Python 3.12+, `uv` + `pyproject.toml`, src layout. SQLite for storage. Keep dependencies
minimal (feedparser/httpx-level, no frameworks until the API milestone).
- CHANGELOG.md in Keep a Changelog format; SemVer.
- **No AI/Claude attribution in commit messages** (no Co-Authored-By). Disclose AI use in the
README instead. Commit messages describe actual changes, concise; never reference prompts
or instructions.
- Significant multi-commit features go on a feature branch; small increments can go on main
while the project is pre-0.1.
- No remote configured yet — do NOT create forge repos or add remotes; the owner will decide
hosting.
+50
View File
@@ -0,0 +1,50 @@
# hark — plan
Milestones. Each one ships something usable and gets a CHANGELOG version.
## M0 — scaffold + ingest (current)
- Project scaffold: uv/pyproject, src layout, pytest.
- SQLite schema: shows, episodes, topics, episode_topics (extraction fields nullable —
populated in M1).
- Feed resolution: show names in `feeds.txt` → feed URLs via iTunes Search API (keyless).
- RSS ingest: fetch + parse feeds, upsert shows/episodes (id, title, description, pubdate,
duration, audio URL). Idempotent re-runs.
- CLI: `hark resolve`, `hark ingest`, `hark stats`.
- Unit tests with feed fixtures (no network in tests).
## M1 — topic extraction + index
- LLM extraction of subject entities from title/description (stub interface in M0; model
wiring decided when we get here).
- Canonicalization against Wikidata (aliases: "BTK" = "Dennis Rader"); multi-part/serial
episode handling; multi-genre topics.
- Topic pages: "who covered X" — the core query.
## M2 — discovery
- Embedding similarity over episode topics → related shows, notable back-catalog episodes.
- Candidate-show pipeline: cheap signals first, deeper analysis only for shows that pass.
## M3 — AntennaPod loop
- Read subscriptions/history from Nextcloud gpodder sync (truenas).
- Generate custom RSS feeds as the recommendation delivery channel.
## M4 — episode scoring (tiltmeter-style)
- Defined interestingness metrics, calibration loop against owner ratings.
- Per-topic treatment comparison (depth, sensationalism) — needs transcripts for fidelity;
revisit Whisper here, not before.
## Seed shows (feeds.txt)
Start with well-known subject-per-episode shows across two genres, e.g.: Casefile,
Casual Criminalist, Swindled, The Rest Is History, Short History Of, Cautionary Tales.
Resolve their real feed URLs via iTunes Search API at runtime — do not hand-copy URLs.
## Open questions (owner input needed, don't block on these)
- Hosting (private Gitea vs GitHub) — no remote until decided.
- Which LLM/provider for extraction (M1 decision).
- GPU/Whisper feasibility on this LXC (M4 decision; CUDA device nodes may not be exposed).
+9
View File
@@ -0,0 +1,9 @@
# Seed shows — one per line. Resolved to feed URLs via iTunes Search API (`hark resolve`).
# true crime
Casefile True Crime
The Casual Criminalist
Swindled
# history
The Rest Is History
Short History Of
Cautionary Tales