diff --git a/Benchmarks.md b/Benchmarks.md new file mode 100644 index 0000000..7ee9812 --- /dev/null +++ b/Benchmarks.md @@ -0,0 +1,90 @@ +# Benchmarks + +Inference latency and throughput for winnow's two embedding models, measured on real hardware. Run with the `scripts/benchmark.py` script (included in the repo). + +--- + +## Hardware + +| | | +| :-- | :-- | +| **CPU** | AMD / Intel (see per-run notes) | +| **GPU** | NVIDIA GeForce RTX 2070 SUPER | +| **Driver** | 570.172.08 | +| **Container CUDA** | 12.8.1 | +| **Host OS** | Debian 13 (TrueNAS LXC) | + +--- + +## InsightFace Buffalo_L — face mode + +Detection (RetinaFace det_10g) + ArcFace embedding (w600k_r50) on a 640×640 image. This is the pipeline winnow runs for every face crop it evaluates. + +### RTX 2070 SUPER + +| Mode | Model load | Median latency | Throughput | +| :-- | --: | --: | --: | +| GPU (`FORCE_CPU=false`) | — | — | — | +| CPU (`FORCE_CPU=true`) | 4.3 s | 102 ms | 9.8 img/s | + +*GPU results pending CUDA 12.8.1 image build.* + +### Notes + +- Latency measured over 30 runs after 5 warmup iterations. +- Input: 640×640 synthetic image. The detection network processes the full input regardless of whether faces are found; timing is representative of real-world single-image throughput. +- 320×320 input: CPU median 101 ms — detection runtime is dominated by the fixed model overhead, not image size at these resolutions. + +--- + +## SigLIP ViT-B/16 — object mode + +`google/siglip-base-patch16-224` — 224×224 Vision Transformer used for object-mode diversity selection. Supports batched inference; GPU benefit scales with batch size. + +### RTX 2070 SUPER + +#### GPU (`FORCE_CPU=false`) + +| Batch | ms/batch | ms/img | img/s | p95/img | +| --: | --: | --: | --: | --: | +| 1 | — | — | — | — | +| 4 | — | — | — | — | +| 8 | — | — | — | — | +| 16 | — | — | — | — | +| 32 | — | — | — | — | + +*GPU results pending.* + +#### CPU (`FORCE_CPU=true`) + +| Batch | ms/batch | ms/img | img/s | p95/img | +| --: | --: | --: | --: | --: | +| 1 | 253 | 253 | 4.0 | 315 | +| 4 | 821 | 205 | 4.9 | 214 | +| 8 | 1566 | 196 | 5.1 | 214 | +| 16 | 3181 | 199 | 5.0 | 205 | +| 32 | 6175 | 193 | 5.2 | 199 | + +Model load: **18.8 s** (CPU; first load, no cache) + +CPU batching saturates quickly — throughput barely improves past batch 4 (~5 img/s ceiling). On GPU, large batches are expected to see significant throughput gains. + +--- + +## Running the benchmark + +```bash +# Inside the container — GPU mode: +docker exec winnow python /app/scripts/benchmark.py + +# CPU-only mode: +docker exec -e FORCE_CPU=true winnow python /app/scripts/benchmark.py + +# Or directly with docker run: +docker run --rm --gpus all \ + --entrypoint /app/.venv/bin/python \ + -v /your/models:/insightface \ + -e INSIGHTFACE_HOME=/insightface \ + ghcr.io/sudolulo/winnow:latest \ + /app/scripts/benchmark.py +``` diff --git a/Home.md b/Home.md index 2e2380e..4652f93 100644 --- a/Home.md +++ b/Home.md @@ -9,6 +9,7 @@ It runs fully headless in Docker, is configured entirely through environment var ## Pages - [[Setup]] — installation, Docker Compose configuration, GPU passthrough, environment variables +- [[Benchmarks]] — GPU vs CPU inference latency and throughput (InsightFace, SigLIP) - [[Troubleshooting]] — common failures and fixes - [[FAQ]] — frequently asked questions diff --git a/Setup.md b/Setup.md index b7197b5..83ab2c8 100644 --- a/Setup.md +++ b/Setup.md @@ -25,7 +25,7 @@ This is the base URL of your Frigate instance, e.g. `http://192.168.1.10:5000`. | Tag | Arch | Acceleration | | :-- | :-- | :-- | -| `:latest` | amd64 + arm64 | NVIDIA CUDA 13.3 (amd64) · requires [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html) | +| `:latest` | amd64 + arm64 | NVIDIA CUDA 12.8 (amd64) · requires [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html) · minimum driver 570 | | `:rocm` | amd64 | AMD ROCm · pass `/dev/kfd` + `/dev/dri` | | `:intel` | amd64 | Intel Arc / iGPU via OpenVINO · pass `/dev/dri`, set `OPENVINO_DEVICE=GPU` | | `:cpu` | amd64 + arm64 | CPU only · ~2 GB smaller · no GPU required | @@ -151,7 +151,11 @@ environment: - OPENVINO_DEVICE=GPU # omit to run OpenVINO inference on CPU (default) ``` -Set `OPENVINO_DEVICE=GPU` to target the Intel GPU. Omitting it (or setting `CPU`) runs OpenVINO on CPU — useful if you want Intel's optimised runtime without GPU passthrough. +`OPENVINO_DEVICE` selects the compute device within the OpenVINO runtime: +- `GPU` — Intel iGPU or Arc GPU (requires `/dev/dri` passthrough) +- `CPU` *(default)* — CPU cores, via OpenVINO's CPU plugin which uses AVX512 and multi-threading. This is *not* the iGPU — it runs on the CPU itself, but OpenVINO's optimised kernels are meaningfully faster than standard CPUExecutionProvider. + +Use `CPU` if you want OpenVINO's faster CPU path without passing through the GPU device. --- @@ -216,7 +220,7 @@ docker exec winnow cat /app/frigate_train/winnow.log | Variable | Default | Description | | :--- | :--- | :--- | | `FORCE_CPU` | `false` | Disable GPU — fall back to CPU for all inference | -| `OPENVINO_DEVICE` | `CPU` | Intel variant only: set `GPU` to use Arc or iGPU | +| `OPENVINO_DEVICE` | `CPU` | Intel variant only: `GPU` = iGPU/Arc (requires `/dev/dri`); `CPU` = OpenVINO CPU plugin (optimised CPU kernels, not the iGPU) | | `ENABLE_CACHE` | `true` | Cache computed embeddings to disk (speeds up re-runs on the same library) | | `CACHE_DIR` | `.if_cache` | Path for embedding cache and upload tracker files | | `HF_HOME` | *(system)* | HuggingFace model cache path (SigLIP) | diff --git a/Troubleshooting.md b/Troubleshooting.md index 74b83ec..ef4d723 100644 --- a/Troubleshooting.md +++ b/Troubleshooting.md @@ -118,7 +118,7 @@ Everything works on CPU but embedding computation is slower — typically tens o - render ``` -2. **Ensure `OPENVINO_DEVICE=GPU` is set** — without it, OpenVINO defaults to CPU even with device passthrough. +2. **Ensure `OPENVINO_DEVICE=GPU` is set** — without it, OpenVINO runs on the CPU (via its optimised CPU plugin), not the iGPU, even with `/dev/dri` passed through. 3. **Verify OpenVINO sees the device**: ```bash