fix: second audit — the delete check did nothing on a real box, and five guards were untestable
The most important finding is that the FIRST audit's fix was wrong. _can_delete() asked `callable(getattr(service, "delete"))`. But CRUDService defines `delete` on the BASE class and dispatches to self.do_delete at call time, so a bound `delete` exists on every CRUDService subclass whether or not it still implements one. The check was therefore answering "is this a CRUDService?" — precisely the weaker "is the namespace registered?" question its own docstring said must never be asked. It would still have picked a gutted pool.snapshot and failed every delete. It now walks the MRO and ignores middlewared.service.* plumbing, so only a PLUGIN class defining delete/do_delete counts. The test double was equally wrong: it modelled a gutted service as object(), a shape middlewared cannot produce, so the test passed against a fake it could never have caught in the field. It is now CRUDService-shaped. Also: - The recursive delete's fast path returned [] without confirming anything was destroyed. A delete that returns cleanly is not proof — iX has already gutted pool.snapshot.do_update on master into a no-op that returns None. cleanup_task read "no survivors" as a clean sweep, dropped the sidecar (the only record), and would have orphaned ~250 snapshots per run, silently. It confirms against ZFS now, and the by-name sweep trusts ZFS rather than the API's return value. - When ZFS cannot be read, the sweep no longer claims success. The two mistakes are not symmetric: a false survivor self-heals (sidecar kept, next run reclaims, record clears), a lost record does not. - _write_sidecar swallowed OSError. The sidecar is the only record the snapshots exist; failing to write it must never be invisible. - stage_nested now refuses UP FRONT when middleware has no usable snapshot delete, rather than discovering it after restic has already run. Tests. The autouse fixture added in the last commit did not work: `runner=_run`, `mounts_file="/proc/self/mounts"` and `sleep=time.sleep` are frozen into __defaults__ at def time, so monkeypatching the module attribute never reached them. 19 tests were still reading the real mount table — one matching name from running a real `umount` on the NAS — and the retry loop really slept. All three are late-bound now; the suite reads nothing outside tmp_path and runs in 1.1s. Every mutation the audit reported as SURVIVING now fails the suite: the naive delete check, the unconfirmed fast path, the malformed-row guard, a disconnected GC, eager service resolution, compat's method check, compat's unknown handling, a single-quoted filtered query in apply.sh, and the get-service assumption. Also: fingerprint() folded `unknown` problems into a broken module, so one transient 429 rewrote the bug report and the next clean run rewrote it back. Problems are state-tagged; only definite breakage is digested. Verified on TrueNAS 26.0.0-BETA.1: zvol-orphan case 0 orphans, 292-dataset backup 0 orphans / 0 leaked mounts / 0 stale sidecars, byte-identical restore of a 4-deep child dataset.
This commit is contained in:
@@ -54,6 +54,46 @@ worse than no alert, because one day it carries a security fix.
|
||||
`zfs.dataset.query`, which returns all 270 datasets. The bug existed only in the
|
||||
unreleased TrueNAS 26 port.
|
||||
|
||||
- **The patch now owns the snapshot sweep even when it does not stage anything.**
|
||||
Stock decides whether to take a *recursive* snapshot by its own rule, and on
|
||||
TrueNAS 26 that rule stopped being ours.
|
||||
|
||||
Up to 25.10, stock's `create_snapshot` called `get_dataset_recursive()` — the same
|
||||
function this module vendors — so "stock went recursive" and "we have something to
|
||||
stage" were the *same question*, and stock's non-recursive delete was correct for
|
||||
everything the patch declined to stage. TrueNAS 26 uses `filesystem.statfs`:
|
||||
`recursive = (path == the dataset's mountpoint)`. The two rules now disagree for a
|
||||
dataset whose only descendants are **ZVOLs** or **legacy/none-mountpoint** datasets
|
||||
— stock snapshots it recursively, while the patch sees nothing to stage.
|
||||
|
||||
The patch then handed the snapshot back to stock, which destroys the parent only.
|
||||
With no staging tree there was no sidecar, and the garbage collector only ever ran
|
||||
from the staging path — so nothing on the box would ever have found the children.
|
||||
Reproduced on the test VM: one orphaned snapshot per zvol, on every run, forever,
|
||||
with the backup reporting success. Ownership of the sweep is no longer conditional
|
||||
on staging.
|
||||
|
||||
- **The runtime resolved a *namespace*; the checker verified a *method*.** Those are
|
||||
different questions, and the gap is a false "ok". `get_service()` only proves a
|
||||
namespace is registered — it says nothing about whether `delete` still exists on it.
|
||||
So if iX guts the method while keeping the service (they have already done exactly
|
||||
that to `pool.snapshot.do_update` on master), `tools/compat.py` would fall through
|
||||
to `zfs.snapshot`, report the box healthy, and let the patch apply — while the
|
||||
runtime picked `pool.snapshot` and failed *every* delete, orphaning the whole tree.
|
||||
Both sides now ask the same question, and a test binds the two lists together.
|
||||
|
||||
- `query_filesystems()` **dropped malformed `zfs list` rows silently** — the last
|
||||
remaining silent-omission path, and a direct contradiction of this module's cardinal
|
||||
rule. It raises now. A missing `zfs` binary raised `FileNotFoundError` rather than
|
||||
`ZfsError`; also fixed.
|
||||
|
||||
- The snapshot retry loop **discarded the delete error** and reported every survivor
|
||||
as "(still busy?)" — naming the one cause that is benign and self-healing, and
|
||||
hiding the ones that are permanent. It keeps and reports the real error.
|
||||
|
||||
- The staging-failure handler could **lose the original exception** if its own cleanup
|
||||
sweep raised. An error handler must not be able to lose the error.
|
||||
|
||||
## v0.6.1 — 2026-07-13
|
||||
### Fixed
|
||||
|
||||
|
||||
Reference in New Issue
Block a user