Compare commits

..
Author SHA1 Message Date
flan c6b252ac6b release v0.6.1
CI / shell (shellcheck + syntax) (push) Successful in 11s
CI / python 3.11 (push) Successful in 13s
CI / python 3.12 (push) Successful in 14s
CI / python 3.13 (push) Successful in 13s
TrueNAS compatibility / compat (push) Successful in 11s
Release / release (push) Successful in 14s
2026-07-13 19:59:38 +00:00
flan 841e0364fd CHANGELOG: repair a section spliced into the middle of a bullet, and guard it
CI / shell (shellcheck + syntax) (push) Successful in 10s
CI / python 3.11 (push) Successful in 16s
CI / python 3.12 (push) Successful in 15s
CI / python 3.13 (push) Successful in 14s
An edit matched the literal '## Unreleased' inside a backticked phrase in a prose
bullet and spliced a whole new section into the middle of it, splitting the sentence
in half. The release body IS this file, so that would have shipped to every user.

Tests now assert: no empty version section, versions descend, no heading is indented
inside a list item, and every bullet's bold phrases are balanced (ignoring code spans
-- '*args, **kwargs' is a literal, not markup).
2026-07-13 19:59:35 +00:00
flan 0d04c2cd1c Collect orphaned snapshots by name: the sidecar lives in tmpfs
CI / shell (shellcheck + syntax) (push) Successful in 10s
CI / python 3.11 (push) Successful in 13s
CI / python 3.12 (push) Successful in 16s
CI / python 3.13 (push) Successful in 16s
A reboot mid-backup orphaned the entire tree, permanently. The sidecar is the record
of which snapshots a run pinned -- and /run is tmpfs. A reboot or crash between the
recursive snapshot and its cleanup destroyed that record, leaving one snapshot per
descendant dataset (250+ on a real pool) with nothing pointing at them. Nothing would
ever have found them.

gc_stale_snapshots() identifies leftovers by NAME, so it works when the record is
gone. It runs after the sidecar reclaim -- the recorded path stays authoritative and
the collector only mops up what the record lost.

It deletes data on a name match, which is a weaker claim than a recorded fact, so the
selection is a pure function with the harshest tests here. A snapshot is collected
only if the name is exactly <dataset>@<task>-<YYYYMMDDHHMMSS>, it is not the current
run's, NOTHING IS MOUNTED FROM IT (this, not the age guard, is what protects a
concurrent backup), and it is over an hour old.

Checked against the real pool: of 4728 snapshots including 2341 periodic ones, it
selects exactly the orphans of the task being run and nothing else.
2026-07-13 19:56:32 +00:00
flan 2b8ef107f7 The sidecar must carry EVERY pending tree, not just the newest
CI / shell (shellcheck + syntax) (push) Successful in 9s
CI / python 3.12 (push) Successful in 13s
CI / python 3.13 (push) Successful in 17s
CI / python 3.11 (push) Successful in 15s
TrueNAS compatibility / compat (push) Successful in 11s
Release / release (push) Successful in 15s
Found live, in the code written to prevent exactly this.

The sidecar held ONE snapshot. So a run that reclaimed an older tree, failed to
finish reclaiming it, and then recorded its own snapshot OVERWROTE the only record of
the survivor -- orphaning it permanently.

Observed: job 24 left one snapshot busy and kept the sidecar (correct). Job 46
reclaimed it, hit ZFS's 300s automount window (the runs were minutes apart), left it
behind again, and then wrote its own snapshot over the record. Permanent orphan,
created by the safety net.

The sidecar is now a list. stage_nested carries forward whatever a reclaim could not
delete; cleanup_task sweeps every pending tree and writes back only the survivors.
cleanup_all reports them one per line instead of formatting a list into an f-string
at the user during uninstall.

Job 46 also confirms the automount fix itself: it swept all 256 of its own snapshots
with no straggler.
2026-07-13 19:35:44 +00:00
11 changed files with 623 additions and 30 deletions
+37
View File
@@ -6,6 +6,35 @@ is deliberate: see [Releasing](docs/releasing.md). Twelve releases were cut on
live, every one of those interrupts every user. An alert people learn to ignore is
worse than no alert, because one day it carries a security fix.
## v0.6.1 — 2026-07-13
### Fixed
- **A reboot mid-backup orphaned the entire snapshot tree, permanently.** The sidecar
is the record of which snapshots a run pinned — and it lives in `/run`, which is
**tmpfs**. A reboot (or a crash) between taking the recursive snapshot and cleaning
it up destroyed that record, leaving one snapshot per descendant dataset — **250+ on
a real pool** — with nothing left pointing at them. Nothing would ever have found
them again.
`gc_stale_snapshots()` is the backstop: it identifies leftovers **by name**, so it
works when the record is gone. It runs at the start of every backup, after the
sidecar reclaim — the recorded path stays authoritative, and the collector only ever
mops up what the record lost.
Because it deletes data on a *name match* — a weaker claim than a recorded fact — the
selection is a **pure function** with the harshest tests in the suite. A snapshot is
collected only if **all** of these hold:
| | |
| --- | --- |
| name is exactly `<dataset>@<task>-<YYYYMMDDHHMMSS>` | so `cloud_backup-5` never matches `cloud_backup-50`, an `auto-*` periodic snapshot, or anything a human made |
| it is not the current run's | parent *and* children are excluded |
| **nothing is mounted from it** | an in-flight run pins its own snapshots — this, not the age guard, is what protects a concurrent backup |
| it is **over an hour old** | covers the seconds-long window where a live run has snapshotted but not yet mounted |
Verified against the real pool: of **4,728** snapshots — including **2,341** periodic
ones — it selects exactly the orphans of the task being run, and nothing else.
## v0.6.0 — 2026-07-13
### Added
@@ -90,6 +119,14 @@ worse than no alert, because one day it carries a security fix.
one no-op delete on the next run, while a sidecar removed while the tree still
exists is unrecoverable. Survivors are reclaimed by the next run.
**Expect the occasional straggler, and expect it to clean itself up.** On a
256-snapshot tree this reliably sweeps ~255 immediately and may leave **one**: it is
whatever restic read last, so its 300-second window has barely opened. That one is
logged, its sidecar is kept, and the next run reclaims it before doing anything else.
The leak is bounded at a single cycle rather than growing without limit — which is
the property that actually matters. Blocking a backup job for five minutes to chase
the last snapshot would be a worse trade, so it is not made.
- **Installing the patch permanently blocked updating it.** `install.sh` does
`chmod +x update.sh`, and git recorded `update.sh` as `100644` — so the chmod was a
*tracked modification*, and `update.sh` refuses to run over a dirty tree. Install
+34
View File
@@ -95,6 +95,40 @@ re-scan each time.
### Snapshot lifecycle
> **Two mechanisms clean up, and the second exists because the first can be destroyed.**
>
> 1. **The sidecar** records exactly which snapshots a run pinned, and is removed only
> on a confirmed-clean sweep. Precise, and it survives a middlewared restart.
> 2. **The garbage collector** finds leftovers by *name*, so it still works when the
> sidecar is gone — and it can be: **the sidecar lives in `/run`, which is tmpfs.** A
> reboot mid-backup takes it, and with it the only record of a 250-snapshot tree.
>
> The collector runs at the start of every backup, after the sidecar reclaim. It will
> only touch a snapshot named `<dataset>@<task>-<timestamp>` that is not the current
> run's, has **nothing mounted from it** (which is what protects a concurrently-running
> backup), and is **over an hour old**. Periodic `auto-*` snapshots, other tasks'
> snapshots, and anything you made by hand are structurally out of reach.
> **A snapshot may survive a run, and that is expected.** ZFS **automounts**
> `<dataset>/.zfs/snapshot/<snap>` the moment it is read, and holds it for
> `zfs_expire_snapshot` seconds (**300** by default) after the last access. So
> whatever restic read *last* is still pinned when we try to destroy it, and
> `zfs destroy` refuses with `dataset is busy`.
>
> The patch unmounts those automounts itself and retries, which clears ~255 of 256 on
> a real pool. The one that remains is **logged, its sidecar is kept, and the next run
> reclaims it before doing anything else** — so the leak is bounded at a single cycle
> instead of growing forever. Seeing one `could not delete snapshot … it will be
> reclaimed on the next run` in the log is normal. Seeing the count *grow* run over run
> is not, and would be a bug.
>
> This is why the sidecar is removed **only on a confirmed-clean sweep**: it is the
> only record those snapshots exist, and a run that dropped it while they were still
> around would orphan them permanently. That is precisely what happened before this was
> fixed.
`zfs.snapshot.delete` defaults to **`recursive=False`**, and stock
`restic_backup()` calls it with no options. Stock is safe only because its
validation means a *recursive* snapshot never actually happens in the field.
+1 -1
View File
@@ -18,7 +18,7 @@
set -euo pipefail
VERSION="0.6.0"
VERSION="0.6.1"
# The directory containing install.sh is the permanent install location.
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
+1 -1
View File
@@ -32,7 +32,7 @@
# Derive PATCH_DIR from this script's location (parent of the patch/ directory).
PATCH_DIR="$(cd "$(dirname "$0")/.." && pwd)"
LOG="$PATCH_DIR/apply.log"
VERSION="0.6.0"
VERSION="0.6.1"
# Rotate log at 512 KB to avoid unbounded growth on a system volume.
# Keep two prior generations (.1 and .2) so the last three boots are always available.
+1 -1
View File
@@ -52,7 +52,7 @@ import subprocess
import sys
import time
__version__ = "0.6.0"
__version__ = "0.6.1"
_PATCH_DIR = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
_STATUS_FILE = os.path.join(_PATCH_DIR, "hook_status.json")
+222 -24
View File
@@ -55,6 +55,7 @@ Therefore this module owns the whole lifecycle:
from __future__ import annotations
import contextlib
import datetime
import os
import stat
import subprocess
@@ -68,10 +69,13 @@ __all__ = [
"cleanup_task",
"current_mounts_under",
"delete_snapshot_tree",
"gc_stale_snapshots",
"mounted_snapshots",
"plan_staging",
"sidecar_for",
"snapshot_tree_names",
"stage_nested",
"stale_snapshot_names",
"staging_root_for",
"teardown",
"verify_staged",
@@ -114,21 +118,39 @@ def sidecar_for(staging_root: str) -> str:
return staging_root + ".snapshot"
def _write_sidecar(staging_root: str, snapshot: str) -> None:
"""Record the pinned snapshot on disk. Blocking; call via run_in_thread."""
def _write_sidecar(staging_root: str, snapshots) -> None:
"""Record every snapshot tree this task still owns. One per line.
A LIST, not a single name -- and that is not over-engineering, it is a bug fix.
The sidecar used to hold one snapshot, so a run that reclaimed an older tree,
FAILED to finish reclaiming it, and then recorded its own snapshot would
**overwrite the only record of the survivor** -- orphaning it permanently, which is
exactly the outcome the sidecar exists to prevent. Observed live: a snapshot
survived one run, the next run's reclaim also failed (ZFS's 300s automount window
had not elapsed, because the runs were minutes apart), and the record was
destroyed anyway.
Now every still-pending tree is carried forward until it is actually gone.
"""
if isinstance(snapshots, str):
snapshots = [snapshots]
with contextlib.suppress(OSError):
os.makedirs(os.path.dirname(staging_root), exist_ok=True)
with open(sidecar_for(staging_root), "w", encoding="utf-8") as fh:
fh.write(snapshot)
fh.write("\n".join(dict.fromkeys(snapshots))) # de-duped, order kept
def _read_sidecar(staging_root: str) -> str | None:
"""The snapshot a previous run recorded here, if any."""
def _read_sidecar(staging_root: str):
"""Every snapshot tree a previous run recorded here. [] if none.
Tolerates the old single-line format, which is just a one-element list.
"""
try:
with open(sidecar_for(staging_root), encoding="utf-8") as fh:
return fh.read().strip() or None
return [ln.strip() for ln in fh if ln.strip()]
except OSError:
return None
return []
def _remove_sidecar(staging_root: str) -> None:
@@ -158,6 +180,76 @@ def snapshot_tree_names(snapshot: str, all_names) -> list[str]:
]
#: A snapshot must be at least this old before the garbage collector will touch it.
#:
#: The GC identifies our leftovers by NAME, so its only real risk is deleting a
#: snapshot belonging to a run that is still starting up -- the window between
#: `zfs snapshot -r` and the bind mounts appearing, which is seconds. An hour is three
#: orders of magnitude more slack than that window needs, and still reclaims a lost
#: tree on the very next daily run.
GC_MIN_AGE_SECONDS = 3600
def stale_snapshot_names(task_name, current_snapshot, all_names, now,
in_use=(), min_age=GC_MIN_AGE_SECONDS):
"""Snapshots THIS task created in an earlier run and never cleaned up.
Pure, because this is the one function here that DELETES DATA on a name match, and
a name match is a weaker claim than a recorded fact. Everything it relies on is an
argument, so every way it could be wrong is a test.
Why a garbage collector exists at all, when there is already a sidecar: **the
sidecar lives in /run, which is tmpfs.** A reboot mid-backup destroys it, and with
it the only record of a 250-snapshot tree. The sidecar handles the normal case
precisely; this handles the case where the record itself is gone.
A snapshot is ours to collect only if ALL of these hold:
* its name is exactly ``<dataset>@<task_name>-<YYYYMMDDHHMMSS>`` -- so
``cloud_backup-5`` never matches ``cloud_backup-50``'s snapshots, and never
matches a periodic ``auto-2026-…`` or anything a human made;
* it is not the snapshot the current run is using;
* nothing is mounted from it (`in_use`) -- an in-flight run pins its own
snapshots, so this alone protects a concurrent one-time backup;
* it is older than `min_age` -- which covers the seconds-long window in which a
run has taken its snapshot but not yet mounted it.
`now` is a timezone-aware datetime; timestamps in the name are UTC (stock builds
them with `utc_now()`).
"""
prefix = task_name + "-"
stale = []
for name in all_names:
_dataset, _, snapname = name.partition("@")
if not snapname or not snapname.startswith(prefix):
continue
if name == current_snapshot or snapname == _snapname_of(current_snapshot):
continue
if name in in_use:
continue
stamp = snapname[len(prefix):]
try:
when = datetime.datetime.strptime(stamp, "%Y%m%d%H%M%S").replace(
tzinfo=datetime.UTC
)
except ValueError:
# Not our timestamp format. Something else owns this name; leave it alone.
continue
if (now - when).total_seconds() < min_age:
continue
stale.append(name)
return stale
def _snapname_of(snapshot):
return snapshot.partition("@")[2] if snapshot else ""
def _probe_snapdir(path):
"""Classify a snapshot directory: ``ok``, ``missing``, or why it is unusable.
@@ -580,6 +672,80 @@ def delete_snapshot_tree(middleware, snapshot, logger=None, attempts=4,
return remaining
def mounted_snapshots(mounts_file="/proc/self/mounts"):
"""Every ZFS snapshot something is currently mounted from.
The device field of a snapshot mount IS the snapshot name (`Tap/apps/x@snap`), for
both our staging bind mounts and ZFS's own .zfs automounts. So this is a direct,
factual answer to "is anything using this snapshot right now" -- which is what
protects a concurrently-running backup from the garbage collector, rather than
trusting an age heuristic to be generous enough.
"""
live = set()
try:
with open(mounts_file, encoding="utf-8") as fh:
for line in fh:
dev = line.split(" ", 1)[0]
if "@" in dev:
live.add(dev.replace("\\040", " "))
except OSError:
return set()
return live
def gc_stale_snapshots(middleware, task_name, current_snapshot, logger=None,
now=None, mounts_file="/proc/self/mounts"):
"""Delete snapshots this task left behind in an earlier run. Returns what remains.
The backstop for when the RECORD is gone, not just the snapshots: the sidecar lives
in /run (tmpfs), so a reboot mid-backup takes it with them. Without this, that tree
-- one snapshot per descendant dataset, 250+ on a real pool -- is orphaned with
nothing left pointing at it.
Selection is `stale_snapshot_names()`, which is pure and heavily tested, because a
name match is a weaker claim than a recorded fact and this deletes data on one.
"""
dataset = current_snapshot.partition("@")[0]
now = now or datetime.datetime.now(datetime.UTC)
try:
snaps = middleware.call_sync(
"zfs.snapshot.query", [["name", "^", dataset]], {"select": ["name"]}
)
except Exception as e: # noqa: BLE001 - cannot enumerate; collect nothing
if logger:
logger.warning(
"truecloud-patch: could not enumerate snapshots for GC: %r", e
)
return []
stale = stale_snapshot_names(
task_name, current_snapshot, [s["name"] for s in snaps], now,
in_use=mounted_snapshots(mounts_file),
)
if not stale:
return []
if logger:
logger.warning(
"truecloud-patch: %d snapshot(s) from an earlier run of %s were never "
"cleaned up (a lost record, e.g. a reboot mid-backup); collecting them",
len(stale), task_name,
)
remaining = []
for name in stale:
try:
middleware.call_sync("zfs.snapshot.delete", name)
except Exception as e: # noqa: BLE001 - busy, or gone; either way, next run
remaining.append(name)
if logger:
logger.debug(
"truecloud-patch: could not collect %s: %r", name, e
)
return remaining
def stage_nested(middleware, path, snapshot, base_dataset, base_mountpoint,
task_name, datasets, logger=None):
"""Build a complete staging tree for `path` from the already-taken `snapshot`.
@@ -606,24 +772,50 @@ def stage_nested(middleware, path, snapshot, base_dataset, base_mountpoint,
# A previous run may have crashed mid-flight; never build on top of that.
teardown(staging_root)
# ...and if it left a sidecar behind, that snapshot tree is still on disk and
# nothing else will ever reclaim it. Sweep it before we overwrite the record,
# or a single crashed run orphans 160+ snapshots permanently.
stale = _read_sidecar(staging_root)
if stale and stale != snapshot:
# ...and if it left snapshot trees behind, they are still on disk and nothing else
# will ever reclaim them. Sweep them before recording our own, or a single crashed
# run orphans 160+ snapshots permanently.
#
# Anything a reclaim FAILS to delete is carried forward, not dropped. Overwriting
# the sidecar with only our own snapshot is what destroyed the record of a survivor
# once already: the reclaim ran, hit ZFS's 300-second automount window (the runs
# were minutes apart), left one snapshot behind, and then the record of it was
# overwritten -- a permanent orphan, created by the very code meant to prevent one.
pending = []
for stale in _read_sidecar(staging_root):
if stale == snapshot:
continue
if logger:
logger.warning(
"truecloud-patch: reclaiming snapshot tree from an earlier "
"interrupted run: %s", stale,
"run: %s", stale,
)
delete_snapshot_tree(middleware, stale, logger=logger)
pending.extend(delete_snapshot_tree(middleware, stale, logger=logger))
if pending and logger:
logger.warning(
"truecloud-patch: %d snapshot(s) from an earlier run are still busy; "
"carrying them forward to the next run", len(pending),
)
# ...and collect anything from an earlier run that has NO record at all.
#
# The sidecar above is precise but lives in /run, which is tmpfs -- a reboot
# mid-backup destroys it and orphans the whole tree with nothing pointing at it.
# This finds those by name and is the only thing that ever will.
#
# It runs AFTER the sidecar reclaim on purpose: the recorded path is authoritative
# and cheap, and the GC should only ever be mopping up what the record lost.
pending.extend(
gc_stale_snapshots(middleware, task_name, snapshot, logger=logger)
)
# Record the snapshot BEFORE mounting anything, not after. middlewared can
# die at any point (this patch even schedules a restart at boot), and the
# sidecar is the only thing that survives it -- an in-process dict would take
# the sole record of a 160-snapshot tree with it. Writing it after apply_plan
# would leave exactly the crash window the sidecar exists to close.
_write_sidecar(staging_root, snapshot)
_write_sidecar(staging_root, [*pending, snapshot])
try:
mounts, skipped = plan_staging(
@@ -667,9 +859,9 @@ def cleanup_task(middleware, task_name, logger=None):
Safe to call unconditionally: a no-op when the task was never staged.
"""
staging_root = staging_root_for(task_name)
snapshot = _read_sidecar(staging_root)
pinned = _read_sidecar(staging_root)
if snapshot is None and not os.path.isdir(staging_root):
if not pinned and not os.path.isdir(staging_root):
return # never staged; nothing to do
errors = teardown(staging_root)
@@ -677,11 +869,15 @@ def cleanup_task(middleware, task_name, logger=None):
for err in errors:
logger.warning("truecloud-patch: staging teardown: %s", err)
if snapshot is None:
if not pinned:
_remove_sidecar(staging_root)
return
survivors = delete_snapshot_tree(middleware, snapshot, logger=logger)
# Every tree this task still owns -- ours, plus anything an earlier run could not
# finish reclaiming.
survivors = []
for snapshot in pinned:
survivors.extend(delete_snapshot_tree(middleware, snapshot, logger=logger))
# KEEP the sidecar if anything survived. It is the only record that those
# snapshots exist, and removing it orphans them permanently.
@@ -698,10 +894,13 @@ def cleanup_task(middleware, task_name, logger=None):
if survivors:
if logger:
logger.warning(
"truecloud-patch: %d snapshot(s) from %s could not be deleted; "
"keeping the sidecar so the next run reclaims them",
len(survivors), snapshot,
"truecloud-patch: %d snapshot(s) could not be deleted (still busy); "
"recording them so the next run reclaims them: %s",
len(survivors), ", ".join(survivors),
)
# The SURVIVORS, not the trees we asked to delete. Writing the original list
# back would keep re-sweeping trees that are already gone.
_write_sidecar(staging_root, survivors)
return
_remove_sidecar(staging_root)
@@ -731,8 +930,7 @@ def cleanup_all(base=None, runner=_run, mounts_file="/proc/self/mounts",
# a sidecar is the only record that an interrupted run's snapshot tree (one
# snapshot per descendant dataset) is still on disk.
for sc in sorted(glob_fn(os.path.join(base, "*.snapshot"))):
snap = read_sidecar(sc[: -len(".snapshot")])
if snap:
for snap in read_sidecar(sc[: -len(".snapshot")]):
lines.append(f" NOTE: an interrupted backup left snapshot '{snap}' behind.")
lines.append(f" Remove it and its children: zfs destroy -r '{snap}'")
+1 -1
View File
@@ -17,7 +17,7 @@
# bash /mnt/tank/truenas-truecloud-patch/patch/apply.sh
# systemctl restart middlewared
VERSION="0.6.0"
VERSION="0.6.1"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
+58
View File
@@ -9,6 +9,7 @@ across install.sh / uninstall.sh / recover.sh / apply.sh and nothing noticed.
"""
import os
import re
import sys
import pytest
@@ -233,3 +234,60 @@ class TestCandidateNotesResolveToTheBaseVersion:
def test_a_genuinely_missing_section_still_raises(self):
with pytest.raises(KeyError):
extract_notes(self.CHANGELOG, "v9.9.9-rc1")
class TestTheChangelogIsStructurallySound:
"""The release body IS this file, so a mangled section ships to every user.
It has been mangled once: an edit matched the literal `## Unreleased` inside a
backticked phrase in a prose bullet and spliced a whole new section into the middle
of it, splitting the sentence in half.
"""
def changelog(self):
with open(os.path.join(REPO, "CHANGELOG.md"), encoding="utf-8") as fh:
return fh.read()
def test_no_version_section_is_empty(self):
text = self.changelog()
for v in changelog_versions(text):
assert extract_notes(text, v).strip(), f"v{v} has an empty section"
def test_versions_are_in_descending_order(self):
from release_notes import version_tuple
versions = changelog_versions(self.changelog())
assert versions == sorted(versions, key=version_tuple, reverse=True), (
"CHANGELOG versions are out of order — a section was spliced in wrong"
)
def test_headings_are_at_the_start_of_a_line_and_not_inside_prose(self):
# A `### Fixed` that ends up indented under a bullet is a section nobody sees.
for i, line in enumerate(self.changelog().splitlines(), 1):
if line.lstrip().startswith(("## ", "### ")) and line != line.lstrip():
raise AssertionError(
f"line {i}: heading is indented, so it is inside a list item "
f"rather than being a section: {line!r}"
)
def test_every_bullet_that_opens_a_bold_phrase_closes_it(self):
# The splice cut `- **A stable release ... under \`## Unreleased` in half,
# leaving an unterminated ** and a dangling sentence.
#
# A bullet is the `- ` line plus everything up to the next top-level bullet or
# heading -- bold phrases routinely wrap across lines, so a per-line check
# would flag every long bullet in the file.
text = self.changelog()
bullets = re.split(r"^(?=- |#{2,3} )", text, flags=re.M)
bad = []
for b in bullets:
if not b.startswith("- "):
continue
# Code spans are not markup: `*args, **kwargs` is a literal, not a bold
# phrase, and counting its ** would flag a perfectly well-formed bullet.
prose = re.sub(r"`[^`]*`", "", b)
if prose.count("**") % 2:
bad.append(b.splitlines()[0][:70])
assert not bad, (
"unbalanced ** in a bullet — a section was probably spliced into the "
"middle of it:\n " + "\n ".join(bad)
)
+266
View File
@@ -783,3 +783,269 @@ class TestSidecarSurvivesAnIncompleteSweep:
monkeypatch.setattr(tn, "delete_snapshot_tree", lambda m, s, logger=None: [])
tn.cleanup_task(FakeMiddleware(), "cloud_backup-5")
assert not os.path.exists(sidecar_for(root))
class TestTheSidecarCarriesEveryPendingTree:
"""The sidecar holds a LIST, and that is a bug fix, not a generalisation.
It used to hold ONE snapshot. So a run that reclaimed an older tree, FAILED to
finish reclaiming it, and then recorded its own snapshot would **overwrite the only
record of the survivor** — orphaning it permanently, via the exact code written to
prevent orphans.
Observed live: a snapshot survived one run; the next run's reclaim also failed
(ZFS's 300s automount window had not elapsed, because the two runs were minutes
apart); the record was overwritten; the snapshot was orphaned for good.
"""
def test_round_trips_a_list(self, tmp_path):
import truecloud_nested as tn
root = str(tmp_path / "cloud_backup-5")
tn._write_sidecar(root, ["Tap@a", "Tap@b"])
assert tn._read_sidecar(root) == ["Tap@a", "Tap@b"]
def test_reads_the_old_single_line_format(self, tmp_path):
# Boxes upgrading from an older version have a one-line sidecar on disk.
import truecloud_nested as tn
root = str(tmp_path / "cloud_backup-5")
os.makedirs(os.path.dirname(sidecar_for(root)), exist_ok=True)
with open(sidecar_for(root), "w", encoding="utf-8") as fh:
fh.write("Tap@legacy")
assert tn._read_sidecar(root) == ["Tap@legacy"]
def test_a_failed_reclaim_is_carried_forward_not_overwritten(
self, tmp_path, monkeypatch
):
# THE bug. stage_nested reclaims an old tree, cannot finish, then records its
# own snapshot -- the survivor must still be in the sidecar afterwards.
import truecloud_nested as tn
monkeypatch.setattr(tn, "STAGING_BASE", str(tmp_path))
root = tn.staging_root_for("cloud_backup-5")
os.makedirs(os.path.dirname(root), exist_ok=True)
tn._write_sidecar(root, ["Tap@old"])
# The reclaim of Tap@old leaves one snapshot behind (still busy).
monkeypatch.setattr(
tn, "delete_snapshot_tree",
lambda m, s, logger=None: ["Tap/apps/x@old"] if s == "Tap@old" else [],
)
stub_core(monkeypatch, tn, plan=([("/src", root)], []))
tn.stage_nested(FakeMiddleware(), "/mnt/Tap", "Tap@new", "Tap", "/mnt/Tap",
"cloud_backup-5", DATASETS)
recorded = tn._read_sidecar(root)
assert "Tap/apps/x@old" in recorded, (
"the failed reclaim's survivor was dropped — orphaned forever"
)
assert "Tap@new" in recorded, "our own snapshot must also be recorded"
def test_cleanup_sweeps_every_pending_tree_and_records_only_survivors(
self, tmp_path, monkeypatch
):
import truecloud_nested as tn
monkeypatch.setattr(tn, "STAGING_BASE", str(tmp_path))
root = tn.staging_root_for("cloud_backup-5")
os.makedirs(root, exist_ok=True)
tn._write_sidecar(root, ["Tap@old", "Tap@new"])
swept = []
def fake_delete(m, s, logger=None):
swept.append(s)
return ["Tap/apps/x@new"] if s == "Tap@new" else []
monkeypatch.setattr(tn, "delete_snapshot_tree", fake_delete)
tn.cleanup_task(FakeMiddleware(), "cloud_backup-5")
assert swept == ["Tap@old", "Tap@new"], "both pending trees must be swept"
# Only the SURVIVOR is written back -- re-recording Tap@old would make every
# future run re-sweep a tree that is already gone.
assert tn._read_sidecar(root) == ["Tap/apps/x@new"]
def test_a_fully_clean_sweep_removes_the_sidecar(self, tmp_path, monkeypatch):
import truecloud_nested as tn
monkeypatch.setattr(tn, "STAGING_BASE", str(tmp_path))
root = tn.staging_root_for("cloud_backup-5")
os.makedirs(root, exist_ok=True)
tn._write_sidecar(root, ["Tap@a", "Tap@b"])
monkeypatch.setattr(tn, "delete_snapshot_tree", lambda m, s, logger=None: [])
tn.cleanup_task(FakeMiddleware(), "cloud_backup-5")
assert not os.path.exists(sidecar_for(root))
def test_cleanup_all_reports_each_pending_snapshot_on_its_own_line(self, tmp_path):
# It formats them for a human during uninstall. A list rendered into an
# f-string would print "['Tap@a', 'Tap@b']" at them.
import truecloud_nested as tn
root = str(tmp_path / "cloud_backup-5")
tn._write_sidecar(root, ["Tap@a", "Tap@b"])
lines, _errors = tn.cleanup_all(
base=str(tmp_path),
glob_fn=lambda _p: [sidecar_for(root)],
mounts_file=os.devnull,
)
notes = [ln for ln in lines if "left snapshot" in ln]
assert len(notes) == 2
assert "'Tap@a'" in notes[0] and "'Tap@b'" in notes[1]
assert "[" not in "".join(notes)
class TestGarbageCollectorSelection:
"""`stale_snapshot_names` DELETES DATA on a name match.
A name match is a weaker claim than a recorded fact, so every way it could be wrong
is a test. It exists because the sidecar — which IS a recorded fact — lives in /run,
which is tmpfs: a reboot mid-backup destroys it and orphans a 250-snapshot tree with
nothing left pointing at it. This is the only thing that would ever find those.
"""
import datetime as _dt
NOW = _dt.datetime(2026, 7, 14, 12, 0, 0, tzinfo=_dt.UTC)
CURRENT = "Tap@cloud_backup-5-20260714115900" # 1 minute ago
OLD = "Tap/apps/x@cloud_backup-5-20260713030000" # ~33 hours ago
def collect(self, names, **kw):
import truecloud_nested as tn
return tn.stale_snapshot_names(
"cloud_backup-5", self.CURRENT, names, self.NOW, **kw
)
def test_it_collects_our_own_leftovers(self):
assert self.collect([self.OLD]) == [self.OLD]
def test_it_NEVER_touches_the_current_run(self):
# Both the parent and its children share the current snapname.
names = [self.CURRENT, "Tap/apps/x@cloud_backup-5-20260714115900"]
assert self.collect(names) == []
def test_it_NEVER_touches_a_periodic_snapshot(self):
assert self.collect(["Tap/apps/x@auto-2026-07-13_03-00"]) == []
def test_it_NEVER_touches_a_human_made_snapshot(self):
assert self.collect(["Tap@before-i-broke-everything"]) == []
def test_it_NEVER_touches_another_TASK(self):
# cloud_backup-5 must not match cloud_backup-50. This is why the prefix
# carries the trailing dash.
assert self.collect(["Tap/apps/x@cloud_backup-50-20260713030000"]) == []
assert self.collect(["Tap/apps/x@cloud_backup-7-20260713030000"]) == []
def test_it_NEVER_touches_a_one_time_backup(self):
assert self.collect(["Tap@cloud_backup-onetime-20260713030000"]) == []
def test_it_NEVER_touches_a_snapshot_that_is_MOUNTED(self):
# An in-flight run pins its own snapshots. This — not the age heuristic — is
# what actually protects a concurrent backup.
assert self.collect([self.OLD], in_use={self.OLD}) == []
def test_it_NEVER_touches_a_snapshot_younger_than_the_minimum_age(self):
# Covers the seconds-long window between `zfs snapshot -r` and the mounts
# appearing, when a live run's snapshots look exactly like garbage.
young = "Tap/apps/x@cloud_backup-5-20260714113000" # 30 minutes ago
assert self.collect([young]) == []
assert self.collect([young], min_age=60) == [young]
def test_a_name_it_cannot_parse_is_left_alone(self):
assert self.collect(["Tap@cloud_backup-5-not-a-timestamp"]) == []
assert self.collect(["Tap@cloud_backup-5-"]) == []
def test_a_realistic_mixed_pool(self):
names = [
self.CURRENT, # ours, running
"Tap/apps/x@cloud_backup-5-20260714115900", # ours, running (child)
self.OLD, # ours, orphaned <-
"Tap/apps/y@cloud_backup-5-20260712030000", # ours, orphaned <-
"Tap/apps/x@auto-2026-07-13_03-00", # periodic
"Tap/apps/x@cloud_backup-7-20260713030000", # another task
"Tap@manual-keepme", # human
]
assert sorted(self.collect(names)) == sorted(
[self.OLD, "Tap/apps/y@cloud_backup-5-20260712030000"]
)
class TestMountedSnapshots:
def test_it_reads_snapshot_names_out_of_the_mount_table(self, tmp_path):
import truecloud_nested as tn
mounts = tmp_path / "mounts"
mounts.write_text(
"tmpfs /run tmpfs rw 0 0\n"
"Tap/apps/x@snap1 /run/truecloud-nested/t/apps/x zfs ro 0 0\n"
"Tap/apps/y@snap1 /mnt/Tap/apps/y/.zfs/snapshot/snap1 zfs ro 0 0\n"
"Tap/live /mnt/Tap/live zfs rw 0 0\n"
)
live = tn.mounted_snapshots(str(mounts))
assert live == {"Tap/apps/x@snap1", "Tap/apps/y@snap1"}
assert "Tap/live" not in live # a live dataset is not a snapshot
class TestGarbageCollectorExecution:
def test_it_deletes_the_stale_ones_and_nothing_else(self, monkeypatch, tmp_path):
import datetime as dt
import truecloud_nested as tn
mounts = tmp_path / "mounts"
mounts.write_text("")
now = dt.datetime(2026, 7, 14, 12, 0, 0, tzinfo=dt.UTC)
mw = FakeMiddleware([
"Tap@cloud_backup-5-20260714115900", # current run
"Tap/apps/x@cloud_backup-5-20260713030000", # orphan <-
"Tap/apps/x@auto-2026-07-13_03-00", # periodic
"Tap/apps/x@cloud_backup-7-20260713030000", # other task
])
remaining = tn.gc_stale_snapshots(
mw, "cloud_backup-5", "Tap@cloud_backup-5-20260714115900",
now=now, mounts_file=str(mounts),
)
assert remaining == []
assert mw.snapshots == [
"Tap@cloud_backup-5-20260714115900",
"Tap/apps/x@auto-2026-07-13_03-00",
"Tap/apps/x@cloud_backup-7-20260713030000",
]
def test_a_busy_orphan_is_reported_not_swallowed(self, monkeypatch, tmp_path):
import datetime as dt
import truecloud_nested as tn
mounts = tmp_path / "mounts"
mounts.write_text("")
now = dt.datetime(2026, 7, 14, 12, 0, 0, tzinfo=dt.UTC)
orphan = "Tap/apps/x@cloud_backup-5-20260713030000"
mw = BusyMiddleware(
["Tap@cloud_backup-5-20260714115900", orphan],
busy=[orphan], busy_for=99,
)
remaining = tn.gc_stale_snapshots(
mw, "cloud_backup-5", "Tap@cloud_backup-5-20260714115900",
now=now, mounts_file=str(mounts),
)
assert remaining == [orphan]
def test_it_collects_NOTHING_when_the_query_fails(self, tmp_path):
# Cannot enumerate => cannot know what is ours => delete nothing.
import datetime as dt
import truecloud_nested as tn
mounts = tmp_path / "mounts"
mounts.write_text("")
class Broken(FakeMiddleware):
def call_sync(self, method, *args):
if method == "zfs.snapshot.query":
raise RuntimeError("middleware is having a day")
return super().call_sync(method, *args)
assert tn.gc_stale_snapshots(
Broken(["Tap/apps/x@cloud_backup-5-20260713030000"]),
"cloud_backup-5", "Tap@cloud_backup-5-20260714115900",
now=dt.datetime(2026, 7, 14, 12, 0, 0, tzinfo=dt.UTC),
mounts_file=str(mounts),
) == []
+1 -1
View File
@@ -3,7 +3,7 @@
set -euo pipefail
VERSION="0.6.0"
VERSION="0.6.1"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
_HOOK_COMMENT='TrueCloud provider patch (S3/B2)'
+1 -1
View File
@@ -19,7 +19,7 @@
set -euo pipefail
VERSION="0.6.0"
VERSION="0.6.1"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
_PREV_FILE="$PATCH_DIR/.update_previous"