Compare commits

...
13 Commits
Author SHA1 Message Date
flan 092bdeae29 v0.4.1: fix three real bugs in update.sh found by auditing it
Release-candidate tags would have been installed as stable
-----------------------------------------------------------
git's version sort ranks v0.5.0-rc1 ABOVE v0.5.0 (verified empirically), and the
release workflow deliberately supports rc/beta/alpha tags. update.sh would have
offered an RC as "the newest release". Tag selection is now filtered to plain
vX.Y.Z.

update.sh would have died mid-update on an untracked file
----------------------------------------------------------
The dirty-tree guard uses --untracked-files=no, so an untracked file that the
TARGET tracks slips past it -- and `git checkout` then aborts. Under set -e the
script died with a raw git error, after already recording the rollback point.

Not hypothetical: a hand-copied patch/wait_restart.sh blocked a pull on a real box
in exactly this way. It is now detected up front, by name. Gitignored files are
correctly not treated as blockers, since git overwrites those silently.

Special case: if update.sh ITSELF is the blocker, it was hand-copied in to
bootstrap -- and "delete update.sh, then re-run update.sh" is impossible. It now
says so and prints the git commands that bootstrap it properly.

--rollback skipped that check entirely and would have hit the identical failure.
The check is now a shared function used by both paths, and rollback also validates
that the recorded revision still exists.

Also: install.sh's chmod aborted under set -e if a listed file was missing (the
file set changes between versions, so --rollback must not be killed by a name this
version happens to know about), and --to with no value was silently ignored.

Verified end to end in a throwaway clone: forward v0.4.1 -> v0.4.2 and rollback
back, with files appearing and disappearing correctly; both guards fire.

132 tests, ruff and shellcheck -S style clean.
2026-07-13 16:30:50 +00:00
flan 347c415aa7 v0.4.0: add update.sh
Fetch a newer release and apply it, preserving the nested-snapshot opt-in setting.

  bash update.sh              # to the newest release, with a confirmation
  bash update.sh --check      # show what would happen; change nothing
  bash update.sh --rollback   # undo the last update

Deliberately NOT automated. This patch injects Python into middlewared and
re-applies itself at every boot, so an unattended pull would let any bad upstream
commit reach a box with no human in the loop and take effect on the next reboot.
v0.0.4 shipped exactly such a bug and took every app on the box down. The manual
step is the safety gate.

Design decisions worth keeping:

- Defaults to the newest RELEASE TAG, not main. main can be mid-refactor; a tag is
  the tested artifact. --main exists but says so loudly.
- Tags ordered by version, not date. Date order silently downgrades the box the
  first time a hotfix is tagged out of band: a v0.3.6 cut after v0.4.0 would sort
  as "newest".
- Refuses to run over a dirty working tree rather than merging across hand-edited
  or scp'd files. (Verified: the guard fires.)
- Shows the commits and release notes you do not have, read from the TARGET's
  CHANGELOG via tools/release_notes.py -- not a second copy of the extractor.
- Records the previous revision BEFORE moving, so --rollback works even if
  install.sh dies halfway.
- Repairs .git ownership, which past `sudo git pull`s leave root-owned and which
  then breaks every later non-root git command.

update.sh is covered by the version-drift check, so it cannot go stale the way
create_task.py's __version__ did.

Tested end to end in a throwaway clone: detects v0.3.2 -> v0.3.5, lists missing
commits, handles already-up-to-date, and the dirty-tree guard fires.

132 tests, ruff and shellcheck clean.
2026-07-13 16:20:37 +00:00
flan 45f957af23 v0.3.5: log the recursive-delete failure instead of swallowing it
delete_snapshot_tree tries one recursive delete first, then falls back to sweeping
the tree by name. The exception from the fast path was discarded.

That failure is usually benign -- stock's finally already removed the parent once
our mounts were released, which is exactly what the sweep exists to handle. But if
the cause were anything else, this was the only place it was ever visible, and it
went straight to /dev/null. The sweep would then report some different, downstream
symptom. It is now logged before falling through.

Also annotated the two remaining static-analysis findings as considered rather than
leaving them to be re-litigated every audit: subprocess is always invoked in list
form (no shell, so ZFS dataset names cannot inject), and a partial `systemctl` path
is moot in a script that only ever runs as root.

Extended ruleset (E,F,W,B,S,SIM,UP,C4,RET,ARG,A,ISC) and shellcheck -S style both
report zero. 132 tests.
2026-07-13 16:06:12 +00:00
flan 126756498c v0.3.4: one implementation of apply/revert (patch/mw_patch.py)
The "strip the TRUECLOUD_PATCH block" logic existed twice -- in apply.sh's heredoc
and in an inline heredoc in uninstall.sh -- and the uninstall copy was the untested
one. That is precisely how the two could have drifted apart, with apply.sh
reverting one set of files and uninstall.sh another.

Both now call patch/mw_patch.py. 17 new tests cover it, including that
revert_nested never touches restic.py: that file carries a TRUECLOUD_PATCH block
too, but it belongs to the providers module, and removing it would silently break
B2 backups.

apply.sh imports it fail-safe -- on ImportError the backend patch is skipped and
middlewared starts stock, which is this script's whole design principle. The
import uses sys.path.append, never insert(0): prepending would give patch/
precedence over the stdlib for that interpreter, so a future patch/json.py would
shadow the real json module and break the boot.

Also: the README's create_task.py example still taught `--password <secret>`, which
is how a security fix quietly fails to land. It now shows --password-stdin.

132 tests, ruff and shellcheck clean.
2026-07-13 16:02:53 +00:00
flan 60b3ac4557 v0.3.3: keep the restic repo password out of argv and shell history
Security
--------
create_task.py shelled out to `midclt call cloud_backup.create '<json>'`, and that
JSON carries the restic repository password -- so it sat in the subprocess's argv,
which is world-readable via ps, for the duration of the call. That password is the
encryption key for the entire cloud backup repository.

It now talks to the middleware through truenas_api_client, the library that backs
midclt itself, so the password never leaves this process's memory. Verified on a
live box: list-tasks and list-credentials work through the new transport.

--password is also no longer required, because passing a secret as a CLI argument
writes it to shell history permanently. --password-stdin reads it from stdin, and
with neither flag the tool prompts via getpass. --password still works but warns.

Fixed
-----
uninstall.sh could leave every patch installed. It reverted by unmounting the
overlay -- but apply.sh only mounts one when the target directory is read-only. On
a writable /usr it patches the real files in place, and uninstall would remove the
boot hook, report success, and leave the patch applied. It now strips the appended
blocks from the middleware files explicitly. This also covers the case where the
overlay unmount fails.

create_task.py's __version__ had been stuck at 0.2.0 through three releases. The
version-drift check added in v0.3.1 only looked at VERSION= in shell scripts, so it
missed the one file that actually shows a version to users (--version). The check
now covers __version__ too -- and caught this immediately.

118 tests, ruff and shellcheck clean.
2026-07-13 15:43:24 +00:00
flan 8aa9038226 Fix: --disable-nested-snapshots did not disable anything until reboot
apply.sh only ever ADDED patches; there was no revert path anywhere. Disabling
removed the opt-in marker and then merely skipped re-applying -- but the overlay
persists for the whole boot, so the previously patched cloud/{snapshot,crud}.py,
cloud_backup/sync.py and _truecloud_nested.py were all still on disk, and
middlewared re-imported them on the restart install.sh performs.

It printed "DISABLED (stock guard restored)" while the feature kept running until
the next reboot. Someone disabling it because they were worried about it would
have believed it was off.

apply.sh now reverts on every not-needed path (opt-out, or superseded by native
support): remove the module FIRST -- every injected block is guarded by
`if _tc_nested is not None`, so the stock guard comes back even if a later step
fails -- then strip the appended blocks from the three patched files.

restic.py also carries a TRUECLOUD_PATCH block but belongs to the providers
module; reverting it would silently break B2 backups, so it is explicitly
excluded. Verified: the three nested files restore byte-for-byte to stock, the
module is removed, and restic.py's block survives.

install.sh --disable also tears the staging tree down first, since those bind
mounts pin ZFS snapshots that could otherwise never be destroyed.

Updating WITHOUT the flag was always correct and is unchanged: the nested module
is never installed into middleware unless explicitly enabled.

109 tests, ruff and shellcheck clean.
2026-07-13 15:23:27 +00:00
flan ba533dc8ae v0.3.1: automated releases + version-drift check
The "Automated releases" entry was written under v0.3.0, but that commit landed
after the v0.3.0 tag -- the CHANGELOG was claiming the release contained something
it did not. Moved to its own version rather than left as a quiet inaccuracy.

Bumps VERSION to 0.3.1 across all four scripts, which the new consistency check
now enforces.
2026-07-13 15:07:29 +00:00
flan eb91a337cd Automate releases from tags
Releases were manual and had drifted: v0.2.0 and v0.2.1 were tagged but never
released, so the releases page jumped v0.1.0 -> v0.3.0 and hid the fix for the
incident that took every app down.

Pushing a v* tag now runs the full suite and cuts a GitHub release whose body is
the matching CHANGELOG.md section -- one source of truth for release notes, so
there is no second place for them to be wrong.

The workflow refuses to publish when:
  - the tests, ruff, or bash -n fail (a tagged commit is what people install; it
    must be at least as good as main)
  - the tag does not match the VERSION= declared by every script
  - CHANGELOG.md has no section for the tag, or the section is empty

That version check is not theoretical: VERSION= had drifted to three different
values across install.sh / uninstall.sh / recover.sh / apply.sh and nothing
noticed until this release. tests/test_release_notes.py now asserts the scripts
agree with each other and with the newest CHANGELOG entry, so the drift cannot
come back.

workflow_dispatch takes an existing tag, so releases can be backfilled for tags
that were pushed before this existed.

106 tests, ruff and shellcheck clean.
2026-07-13 15:05:46 +00:00
flan f3ea6b301c CHANGELOG: set v0.3.0 release date 2026-07-13 15:02:26 +00:00
flan 51bf5326d9 Refuse to write a bundle whose parens we unbalanced
Commit 47cdf72 shipped a pattern that matched one closing paren and emitted one,
netting an extra `)` in the Angular bundle:

    c(2,"filterByProviders",["STORJ_IX","S3","B2"]))("required",!0)
                                                  ^^ syntax error

The TrueNAS web UI went blank. Worse, MARKER was now present in the file, so
every later run reported "already patched" and skipped -- the patch could not
heal itself, and the bundle had to be hand-restored from the .pre-truecloud-patch
backup.

patch_ui.py now compares the parenthesis balance before and after substitution and
refuses to write if it changed. A bundle we cannot patch correctly is left exactly
as it was: an unpatched UI is a missing dropdown entry, a corrupted one is a dead
web UI.

Tests cover the real regression (verbatim 47cdf72 pattern) end to end: it still
matches, the balance still shifts, main() refuses, and the file on disk is
byte-for-byte unchanged. README documents the manual recovery for anyone who
already hit it.

93 tests, ruff and shellcheck clean.
2026-07-13 14:57:16 +00:00
flan 8aae261018 Nested snapshots validated in production; drop the untested caveat
An unattended scheduled backup of a live 252-dataset pool ran through the staging
tree end to end:

  task 5  /mnt/Tap  SUCCESS  18m14s

- 252 datasets recursively snapshotted; 173 bind mounts built and verified
- zero orphaned ZFS snapshots and zero stale mounts afterwards -- the failure
  that would otherwise have accumulated 251 snapshots on every single run
- the same task previously stalled at 74% for over 12 hours reading live files

The README said the mount --bind staging step had not been exercised by a live
backup run. That is no longer true, so it is removed rather than left to
understate the state of the code.

The advice to verify your own first backup actually contains child-dataset data
stays -- that one is not boilerplate.
2026-07-13 14:52:41 +00:00
flan 8a2028bfa7 Add test coverage for the Angular bundle patch
patch_ui.py rewrites minified third-party JavaScript by regex and had no tests.
It is the easiest place in this project to do real damage: a pattern that matches
nothing silently leaves the dropdown Storj-only, and one that consumes a paren
too many is a syntax error in the bundle that blanks the entire TrueNAS web UI.

Tests run the real patterns against verbatim snippets from a TrueNAS 25.x
chunk-*.js (the chained property(...)(...) form) and a 24.x literal array, and
assert: exactly one match, all three providers present, the paren balance is
UNCHANGED, surrounding code untouched, re-patching is a no-op, unrelated JS is
never matched, and every pattern stays anchored to filterByProviders.

The paren-balance assertion is the load-bearing one -- a plausible-but-wrong
pattern that eats both parens and re-emits none shifts the delta from 1 to 2 and
is caught.

Also restore the comment explaining why the 25.x pattern is shaped the way it is,
and correct the module docstring, which showed the binding with a single closing
paren; the real bundle wraps it in a chained property call.

90 tests, ruff and shellcheck clean.
2026-07-13 14:37:43 +00:00
flan 8421a34d8d Delete the snapshot tree atomically instead of 252 calls
delete_snapshot_tree removed the parent and every child snapshot individually.
On a real pool `zfs snapshot -r` creates one snapshot per descendant dataset --
252 on Tap -- so cleanup was 252 sequential middleware calls.

Slow, but the real problem is that it is not atomic: a job killed part-way
through the sweep leaves behind exactly the orphaned snapshots this function
exists to prevent.

zfs.snapshot.delete accepts {"recursive": True}, which destroys the parent and
all children in one call. Use that as the fast path and keep the name-by-name
sweep as the fallback -- it is still needed when the parent is already gone
(stock's finally can win the race once our mounts are released), which makes a
recursive delete fail while the children survive.

The test fake now emulates real `zfs destroy -r` semantics, so a test cannot pass
while the shipped code deletes only the parent.

76 tests, ruff and shellcheck clean.
2026-07-13 14:35:04 +00:00
20 changed files with 1775 additions and 90 deletions
+1 -1
View File
@@ -51,7 +51,7 @@ jobs:
run: python -m pip install --upgrade pip pytest ruff
- name: ruff
run: ruff check patch tests
run: ruff check patch tests tools
- name: pytest
run: pytest tests -v
+102
View File
@@ -0,0 +1,102 @@
name: Release
# Push a tag, get a release. The body always comes from CHANGELOG.md, so there is
# no second place to write release notes and therefore no second place for them to
# go stale.
#
# git tag -a v0.4.0 -m "v0.4.0" && git push origin v0.4.0
#
# workflow_dispatch re-cuts (or updates) the release for a tag that already
# exists, since re-pushing an existing tag triggers nothing.
#
# It checks out the TAG, because the tagged code is what people install and it has
# to pass its own tests. That means it only works for tags that actually contain
# this tooling (>= v0.3.0). Tags older than that were backfilled by hand.
on:
push:
tags: ["v*"]
workflow_dispatch:
inputs:
tag:
description: "Existing tag to create a release for (e.g. v0.2.1)"
required: true
type: string
permissions:
contents: write
jobs:
release:
runs-on: ubuntu-latest
steps:
- name: Resolve tag
id: tag
run: |
if [ "${{ github.event_name }}" = "workflow_dispatch" ]; then
echo "tag=${{ inputs.tag }}" >> "$GITHUB_OUTPUT"
else
echo "tag=${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
fi
- uses: actions/checkout@v4
with:
ref: ${{ steps.tag.outputs.tag }}
fetch-depth: 0
- uses: actions/setup-python@v5
with:
python-version: "3.13"
# Never publish a release for code that does not pass its own tests. A
# tagged commit is what people install; it has to be at least as good as
# main.
- name: install dev deps
run: python -m pip install --upgrade pip pytest ruff
- name: ruff
run: ruff check patch tests tools
- name: pytest
run: pytest tests -q
- name: shell syntax
run: |
fail=0
while IFS= read -r f; do
bash -n "$f" || { echo "::error file=$f::bash syntax error"; fail=1; }
done < <(find . -name '*.sh' -not -path './.git/*')
exit $fail
# Catches the failure mode this repo actually had: VERSION= drifted to
# three different values across the scripts, and nothing noticed.
- name: version matches tag and CHANGELOG has a section
run: python3 tools/release_notes.py check "${{ steps.tag.outputs.tag }}"
- name: extract release notes from CHANGELOG
run: |
python3 tools/release_notes.py notes "${{ steps.tag.outputs.tag }}" > /tmp/notes.md
echo "--- release body ---"
cat /tmp/notes.md
- name: create or update the release
env:
GH_TOKEN: ${{ github.token }}
TAG: ${{ steps.tag.outputs.tag }}
run: |
# Pre-1.0 and any -rc/-beta suffix ship as prereleases, not "Latest".
prerelease=""
case "$TAG" in
*-rc*|*-beta*|*-alpha*) prerelease="--prerelease" ;;
esac
if gh release view "$TAG" >/dev/null 2>&1; then
echo "Release $TAG exists — updating notes."
gh release edit "$TAG" --notes-file /tmp/notes.md
else
# shellcheck disable=SC2086
gh release create "$TAG" \
--title "$TAG" \
--notes-file /tmp/notes.md \
$prerelease
fi
+213 -1
View File
@@ -1,6 +1,196 @@
# Changelog
## v0.3.0 — 2026-07-12
## v0.4.1 — 2026-07-13
### Fixed
- **`update.sh` would have picked a release candidate as "the newest release".**
Git's version sort ranks `v0.5.0-rc1` *above* `v0.5.0` (verified), and the
release workflow deliberately supports rc/beta tags — so an RC would have been
installed as though it were the latest stable. Tag selection is now filtered to
plain `vX.Y.Z`.
- **`update.sh` would have died mid-update on an untracked file.** The dirty-tree
guard uses `--untracked-files=no`, so an untracked file that the *target* tracks
slipped past it — and `git checkout` then aborts. Under `set -e` the script died
with a raw git error, *after* recording the rollback point. This is exactly what
blocked a pull on a real box (a hand-copied `patch/wait_restart.sh`). It now
detects the collision up front and names the files. Gitignored files are
correctly *not* treated as blockers — git overwrites those silently.
Special case: if `update.sh` *itself* is the blocker, you hand-copied it in to
bootstrap — and "delete `update.sh`, then re-run `update.sh`" is impossible. It
now says so and prints the git commands that bootstrap it properly.
- **`--rollback` skipped that check entirely**, so it would have hit the identical
failure. The check is now a shared function used by both paths, and rollback also
validates that the recorded revision still exists (history can be rewritten).
- `install.sh`'s `chmod` aborted under `set -e` if any listed file was missing. The
file set changes between versions, so `update.sh --rollback` to an older revision
must not be killed by a filename this version happens to know about.
- `--to` with no value was silently ignored and fell back to the default target.
## v0.4.0 — 2026-07-13
### Added
- **`update.sh`** — fetch a newer release and apply it, preserving your
nested-snapshot opt-in setting.
```bash
bash update.sh # to the newest release, with a confirmation
bash update.sh --check # show what would happen; change nothing
bash update.sh --rollback # undo the last update
```
**Run it by hand. Never from cron or a systemd timer.** This patch injects
Python into middlewared and re-applies itself at every boot, so an unattended
pull would let any bad upstream commit reach your box with no human in the loop
and take effect on the next reboot. v0.0.4 shipped exactly such a bug and took
every app on the box down. The manual step *is* the safety gate.
Design:
- **Defaults to the newest release tag, not `main`.** `main` can be mid-refactor;
a tag is the tested artifact. `--main` exists but says so loudly.
- Tags are ordered by **version**, not by date — date order silently downgrades
the box the first time a hotfix is tagged out of band (a v0.3.6 released after
v0.4.0 would sort as "newest").
- **Refuses to run over a dirty working tree** rather than merging across
hand-edited or scp'd files.
- Shows the commits you don't have and the target's release notes (read from the
*target's* CHANGELOG, via `tools/release_notes.py` — not a second copy of the
extractor), then asks before doing anything.
- **Records the previous revision before moving**, so `--rollback` works even if
`install.sh` dies halfway.
- Repairs `.git` ownership, which past `sudo git pull`s leave root-owned and
which then breaks every later non-root git command.
- `update.sh` is covered by the version-drift check, so it cannot quietly go stale
the way `create_task.py.__version__` did.
## v0.3.5 — 2026-07-13
### Changed
- `delete_snapshot_tree` swallowed the error from its recursive-delete fast path.
That failure is *usually* just "parent already gone" — stock's `finally` winning
the race once our mounts are released, which the by-name sweep then handles. But
if the cause were anything else, this was the only place it was visible, and it
went straight to `/dev/null`. It is now logged before falling through.
- Annotated the two remaining static-analysis findings as considered-and-accepted
rather than leaving them to be re-litigated: `subprocess` is always called in
list form (no shell, so ZFS dataset names cannot inject), and the partial
`systemctl` path is moot in a script that only runs as root.
## v0.3.4 — 2026-07-13
### Changed
- **One implementation of apply/revert (`patch/mw_patch.py`).** The "strip the
`TRUECLOUD_PATCH` block" logic existed twice — in `apply.sh`'s heredoc and in an
inline heredoc in `uninstall.sh` — and the uninstall copy was the untested one.
That is exactly how the two could have drifted apart, with `apply.sh` reverting
one set of files and `uninstall.sh` another. Both now call the same tested
module (17 new tests, including that `revert_nested` never touches `restic.py`,
which belongs to the providers module and whose removal would silently break B2
backups).
`apply.sh` imports it fail-safe: if it cannot, the backend patch is skipped and
middlewared starts stock, which is this script's whole design principle. The
import uses `sys.path.append`, never `insert(0)` — prepending would give
`patch/` precedence over the stdlib for that interpreter, so a future
`patch/json.py` would shadow the real `json` and break the boot.
### Docs
- The README's `create_task.py` example still taught `--password <secret>`, which
is how a security fix quietly fails to land. It now shows `--password-stdin`.
## v0.3.3 — 2026-07-13
### Security
- **The restic repository password no longer passes through a process's argv.**
`create_task.py` shelled out to `midclt call cloud_backup.create '<json>'`, and
that JSON contains the repo password — so it appeared in the process's argv,
which is world-readable via `ps`, for the duration of the call. That password is
the encryption key for the entire cloud backup repository.
It now talks to the middleware through `truenas_api_client` (the library that
backs `midclt` itself), so the password never leaves the process's memory.
- **`--password` no longer required.** Passing a secret as a CLI argument writes it
to shell history permanently. `--password-stdin` reads it from stdin, and with
neither flag the tool prompts via `getpass`. `--password` still works but now
warns.
### Fixed
- **`uninstall.sh` could leave every patch installed.** It reverted by unmounting
the overlay — but `apply.sh` only mounts one when the target directory is
read-only. On a writable `/usr` it patches the real files in place, and uninstall
would remove the boot hook, report success, and leave the patch applied. It now
strips the appended blocks from the middleware files explicitly.
- **`create_task.py.__version__` had been stuck at `0.2.0`** for three releases.
The version-drift check added in v0.3.1 only looked at `VERSION=` in shell
scripts, so it missed the one file that actually shows a version to users
(`--version`). The check now covers `__version__` too — and caught this
immediately.
## v0.3.2 — 2026-07-13
### Fixed
- **`install.sh --disable-nested-snapshots` did not actually disable anything
until the next reboot.** `apply.sh` only ever *added* patches — there was no
revert path. Disabling removed the opt-in marker and then merely *skipped*
re-applying, but the overlay persists for the whole boot, so the previously
patched `plugins/cloud/{snapshot,crud}.py`, `plugins/cloud_backup/sync.py` and
`_truecloud_nested.py` were all still sitting there — and middlewared
re-imported them on the restart `install.sh` performs.
It printed *"DISABLED (stock guard restored)"* while the feature kept running.
Someone turning it off *because they were worried about it* would have believed
it was off.
`apply.sh` now actively reverts: it removes the module first (every injected
block is guarded by `if _tc_nested is not None`, so the stock guard is restored
even if a later step fails), then strips its appended blocks from the three
patched files. `restic.py` also carries a `TRUECLOUD_PATCH` block but belongs to
the *providers* module and is deliberately left alone — reverting it would break
B2 backups. `install.sh --disable` also tears down any staging tree first, since
those bind mounts pin ZFS snapshots that could otherwise never be destroyed.
Updating **without** the flag was always correct and is unchanged: the nested
module is never installed into middleware unless it is explicitly enabled.
## v0.3.1 — 2026-07-13
### Added
- **Automated releases.** Pushing a `v*` tag runs the full test suite and then
cuts a GitHub release whose body is the matching `CHANGELOG.md` section — so
release notes have exactly one source of truth, and no second place to go stale.
The workflow refuses to publish if the tests fail, if the tag does not match the
`VERSION=` declared by every script, or if the CHANGELOG has no section for it.
- **Version-drift check.** `VERSION=` had silently diverged to three different
values across `install.sh`, `uninstall.sh`, `recover.sh`, and `patch/apply.sh`,
and nothing noticed. CI now asserts every script agrees with the others and with
the newest CHANGELOG entry.
### Note
- Releases for `v0.2.0` and `v0.2.1` were backfilled — they had been tagged but
never released, so the releases page jumped v0.1.0 → v0.3.0 and hid the fix for
the boot race that took every app down.
## v0.3.0 — 2026-07-13
### Added
@@ -165,6 +355,17 @@
middlewared; there is now a regression test that executes apply.sh's own probe
code against the real wrapped source.
### Changed (production audit)
- **`delete_snapshot_tree` now uses a single recursive delete.** It previously
removed the parent and each child snapshot one at a time — 252 sequential
middleware calls on a real pool. That is slow, but the real problem is that it
is **not atomic**: a run killed part-way through the sweep leaves exactly the
orphaned snapshots the function exists to prevent. It now issues one
`zfs.snapshot.delete(..., {"recursive": True})` and falls back to the
name-by-name sweep only when that fails (e.g. stock's `finally` already removed
the parent, which leaves the children behind).
### Refactored
- Staging teardown had been copy-pasted into `uninstall.sh` and `recover.sh` —
@@ -176,6 +377,17 @@
middlewared-restart case (which empties it) that must not orphan a snapshot
tree. One record, on disk, or none.
### Validated in production
An unattended scheduled backup of a live 252-dataset pool (`/mnt/Tap`, TrueNAS
25.10) ran through the staging tree end to end:
- 252 datasets recursively snapshotted, 173 bind mounts built and verified
- completed in **18m14s**, `SUCCESS` — the same task previously stalled at 74%
for over 12 hours reading live files
- **zero** orphaned ZFS snapshots and **zero** stale mounts afterwards, which is
the failure mode that would otherwise have accumulated 251 snapshots per run
### Known issues
- Stock `restic_backup()` deletes the ZFS snapshot in its own `finally`, which
+37 -24
View File
@@ -38,8 +38,8 @@ knowing:
log after an update.
- If you file a TrueNAS bug report, **remove the patch first** and reproduce on a
stock system.
- **Test your restores.** That is true of any backup, but it matters more here —
see [Verifying it works](#verifying-it-works).
- **Test your restores.** True of any backup, but it matters more here — see
[Verifying it works](#verifying-it-works).
- Provided as-is, no warranty. See LICENSE.
The patch is two independent modules — **providers** (B2/S3) and **nested**
@@ -112,10 +112,14 @@ bash install.sh --disable-nested-snapshots
With neither flag `install.sh` leaves the setting alone, so `git pull && bash
install.sh` won't flip it. The providers module is unaffected either way.
The planner and the snapshot lifecycle have been validated against a real
250-dataset pool. The `mount --bind` staging step has not yet been exercised by a
live backup run, so confirm your first backup actually contains child-dataset
data before relying on it — see [Verifying it works](#verifying-it-works).
Validated end to end on a live 252-dataset pool: an unattended scheduled backup
of `/mnt/Tap` built a 173-mount staging tree, completed in **18m14s**, and left
**zero** orphaned snapshots and **zero** stale mounts behind. The same backup
previously stalled at 74% for over 12 hours reading live files.
Still: verify your own first run actually contains child-dataset data before you
rely on it — see [Verifying it works](#verifying-it-works). That advice is not
boilerplate; it is the specific thing this feature exists to make true.
TrueCloud Backup's **Take Snapshot** option makes restic read from a frozen ZFS
snapshot instead of live files. Without it the backup reads data *while apps are
@@ -248,6 +252,7 @@ mount | grep truecloud-nested # expect no output
| Backup fails: `dataset '…' has no snapshot '…'; refusing to back up an incomplete tree` | Working as designed — a descendant dataset was not covered by the snapshot. The backup is refused rather than silently omitting that data. |
| Backup fails: `snapshot '…' cannot be read (Permission denied)` | The snapshot exists but is unreadable. Middleware runs as root, so this indicates a real permissions problem, not a missing snapshot. |
| `cloud_backup-*` snapshots accumulating | The sweep is not running. Check `apply.log` for the nested patch applying, and confirm `sync.py` carries the `TRUECLOUD_PATCH` block. |
| Web UI blank after a patch | A bad pattern unbalanced the bundle. `apply.sh` now refuses to write in that case, but if you hit it on an older version: restore `chunk-*.js.pre-truecloud-patch` over the live chunk, then re-run `install.sh`. (`MARKER` makes an already-patched file skip, so the patch cannot heal a corrupted bundle by itself.) |
| Stale mounts under `/run/truecloud-nested` | A crashed run. The next backup tears them down. To clear them now: `python3 patch/truecloud_nested.py cleanup` (also run by `uninstall.sh` and `recover.sh`). It names any ZFS snapshot an interrupted run left pinned. |
## Supported providers after patching
@@ -336,27 +341,28 @@ Refresh your browser. S3 and B2 credentials now appear in the
## Updating
To update to a new version of the patch:
```bash
cd /mnt/tank/truenas-truecloud-patch
# If install.sh was previously run as root, the .git directory may be owned
# by root. Fix it first, or just pull as root:
sudo git pull # easiest option
# — or —
sudo chown -R $(whoami) .git && git pull
bash install.sh
bash update.sh # to the newest release, with a confirmation
bash update.sh --check # show what would happen; change nothing
bash update.sh --rollback # undo the last update
```
`install.sh` clears any stale kill switch, re-applies the updated patches,
and restarts middlewared. Run `python3 patch/create_task.py verify` afterwards
to confirm the patches loaded successfully.
It preserves your nested-snapshot opt-in setting, shows you the commits and
release notes you don't have yet, and asks before changing anything. It records
the previous revision *before* moving, so `--rollback` works even if `install.sh`
dies halfway.
Check [CHANGELOG.md](CHANGELOG.md) to see what changed between versions.
**Run it by hand. Never from cron or a systemd timer.** This patch injects Python
into `middlewared` and re-applies itself at every boot, so an unattended pull would
let any bad upstream commit reach your box with no human in the loop and take
effect on the next reboot. v0.0.4 shipped exactly such a bug and took every app on
the box down. The manual step *is* the safety gate — if you want convenience, watch
the [releases](https://github.com/sudolulo/truenas-truecloud-patch/releases) feed,
don't automate the pull.
---
It updates to the newest **release tag**, not `main` — `main` can be mid-refactor,
and a tag is the tested artifact. `--main` exists if you want unreleased code, and
says so loudly.
## Creating a task via CLI
@@ -371,18 +377,25 @@ host address or API key:
# List your cloud credentials to find the right ID
python3 /mnt/tank/truenas-truecloud-patch/patch/create_task.py list-credentials
# Create a task with a B2 credential (id=3)
# Create a task with a B2 credential (id=3).
# The restic repo password is read from stdin, so it never lands in your shell
# history — nor in any process's argv, where `ps` would expose it.
printf '%s' 'restic-repo-password' | \
python3 /mnt/tank/truenas-truecloud-patch/patch/create_task.py create \
--name "tank-to-b2" \
--path /mnt/tank/data \
--credential 3 \
--bucket my-bucket \
--folder backups/tank \
--password "restic-repo-password" \
--password-stdin \
--cache-path /mnt/tank/.restic-cache \
--keep-last 14
```
Omit `--password-stdin` and you'll be prompted for the password instead. `--password
<secret>` still works but warns: that password is the encryption key for the whole
repository, and a CLI argument persists in your shell history forever.
> **Always pass `--cache-path`.** Without it TrueNAS runs restic with `--no-cache`,
> which re-fetches all repo metadata from the provider every run — glacially slow
> on large repos. Point it at a writable dir on a pool with free space.
+18 -5
View File
@@ -18,7 +18,7 @@
set -euo pipefail
VERSION="0.3.0"
VERSION="0.4.1"
# The directory containing install.sh is the permanent install location.
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
@@ -95,8 +95,14 @@ fi
# ── Set permissions ───────────────────────────────────────────────────────────
echo "Setting permissions ..."
chmod +x "$PATCH_DIR/patch/apply.sh" "$PATCH_DIR/patch/create_task.py" \
"$PATCH_DIR/recover.sh" "$PATCH_DIR/uninstall.sh"
# Guard each path: under `set -e` a chmod on a missing file aborts the install.
# The file set changes between versions, so `update.sh --rollback` to an older
# revision must not be killed by a name this version happens to know about.
for _exe in patch/apply.sh patch/create_task.py recover.sh uninstall.sh update.sh; do
if [ -f "$PATCH_DIR/$_exe" ]; then
chmod +x "$PATCH_DIR/$_exe"
fi
done
echo "Done."
echo ""
@@ -160,10 +166,17 @@ case "$_nested_choice" in
off)
if [ -f "$_NESTED_MARKER" ]; then
rm -f "$_NESTED_MARKER"
echo "Nested-dataset snapshots: DISABLED (stock guard restored)."
# Tear down any staging tree first: those bind mounts PIN their ZFS
# snapshots, so leaving them would block those snapshots from ever
# being destroyed. apply.sh (below) then reverts the patched files.
python3 "$PATCH_DIR/patch/truecloud_nested.py" cleanup || \
echo " WARNING: staging mounts remain; unmount them manually."
echo "Nested-dataset snapshots: DISABLED."
echo " apply.sh will revert the patched middleware files and the stock"
echo " guard is restored when middlewared restarts (this script does that)."
echo " Any task that already has snapshot=true on a nested dataset will"
echo " fail validation on its next edit. Turn the option off on those"
echo " tasks, or re-run with --enable-nested-snapshots."
echo " tasks first, or re-run with --enable-nested-snapshots."
else
echo "Nested-dataset snapshots: already disabled."
fi
+41 -19
View File
@@ -32,7 +32,7 @@
# Derive PATCH_DIR from this script's location (parent of the patch/ directory).
PATCH_DIR="$(cd "$(dirname "$0")/.." && pwd)"
LOG="$PATCH_DIR/apply.log"
VERSION="0.3.0"
VERSION="0.4.1"
# Rotate log at 512 KB to avoid unbounded growth on a system volume.
# Keep two prior generations (.1 and .2) so the last three boots are always available.
@@ -468,14 +468,24 @@ if _tc_nested is not None:
"""
def patch_file(path, block):
with open(path, encoding="utf-8") as fh:
content = fh.read()
marker = "\n# TRUECLOUD_PATCH"
idx = content.find(marker)
base = content[:idx] if idx != -1 else content
with open(path, "w", encoding="utf-8") as fh:
fh.write(base.rstrip("\n") + "\n" + block)
# Single implementation of the block apply/revert logic (patch/mw_patch.py), so
# uninstall.sh and apply.sh cannot drift apart. Fail-safe: if it cannot be
# imported, skip the backend patch entirely -- middlewared then starts stock,
# which is the whole design principle of this script.
# APPEND, never insert(0): this dir would otherwise take precedence over the
# stdlib for this interpreter, so a future patch/json.py (say) would shadow the
# real json module and break the boot. Appending fails safe -- worst case our
# import misses and the backend patch is skipped.
sys.path.append(os.path.dirname(nested_src))
try:
from mw_patch import patch_file, revert_nested
except ImportError as _e:
print(f'WARNING: cannot import patch/mw_patch.py ({_e}) — skipping backend patch.')
print('WARNING: middlewared will start with stock (unpatched) modules.')
sys.exit(1)
# .../middlewared/plugins/cloud -> .../middlewared
mw_dir = os.path.dirname(os.path.dirname(cloud_dir))
b2_ok = restic_ok = False
nested_ok = False
@@ -517,16 +527,28 @@ else:
# guard LAST. If anything fails partway, the guard is still in place and the
# option stays unavailable -- we never expose "guard removed, traversal missing".
nested_detail = ''
if not nested_enabled:
nested_detail = 'disabled (opt-in; enable with: install.sh --enable-nested-snapshots)'
print('INFO: Nested-dataset snapshot support is disabled (opt-in feature).')
print('INFO: Enable with: bash install.sh --enable-nested-snapshots')
elif nested_native:
nested_detail = 'superseded: TrueNAS handles nested-dataset snapshots natively'
print('INFO: Nested module skipped — TrueNAS now handles nesting natively.')
elif not nested_needed:
nested_detail = 'not needed'
print('INFO: Nested module skipped.')
if not nested_needed:
# Not just "skip": actively revert. The overlay lives for the whole boot, so a
# previously-applied patch is still sitting there and middlewared would
# re-import it on restart. See revert_nested().
if not nested_enabled:
nested_detail = 'disabled (opt-in; enable with: install.sh --enable-nested-snapshots)'
print('INFO: Nested-dataset snapshot support is disabled (opt-in feature).')
elif nested_native:
nested_detail = 'superseded: TrueNAS handles nested-dataset snapshots natively'
print('INFO: Nested module skipped — TrueNAS now handles nesting natively.')
else:
nested_detail = 'not needed'
print('INFO: Nested module skipped.')
reverted = revert_nested(mw_dir)
if reverted:
print('OK: Reverted a previously-applied nested patch (' + ', '.join(reverted) + ').')
print(' The stock nesting guard is restored once middlewared restarts.')
nested_detail += ' — previous patch reverted'
if not nested_enabled:
print('INFO: Enable with: bash install.sh --enable-nested-snapshots')
else:
try:
snapshot_py = os.path.join(cloud_dir, 'snapshot.py')
+74 -24
View File
@@ -26,8 +26,9 @@ Create a task backed by a B2 credential (id=3):
--credential 3 \\
--bucket my-bucket \\
--folder backups/tank \\
--password "restic-repo-password" \\
--password-stdin \\
--keep-last 14
(pipe the password in: echo -n "s3cret" | python3 create_task.py create ... )
Create a task using an S3-compatible credential (Wasabi, R2, etc.):
python3 create_task.py create \\
@@ -36,7 +37,7 @@ Create a task using an S3-compatible credential (Wasabi, R2, etc.):
--credential 5 \\
--bucket my-bucket \\
--folder backups \\
--password "restic-repo-password"
--password-stdin
List existing TrueCloud Backup tasks:
python3 create_task.py list-tasks
@@ -44,39 +45,49 @@ List existing TrueCloud Backup tasks:
import argparse
import calendar
import getpass
import json
import os
import subprocess
import sys
import time
__version__ = "0.2.0"
__version__ = "0.4.1"
_PATCH_DIR = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
_STATUS_FILE = os.path.join(_PATCH_DIR, "hook_status.json")
def midclt_call(method, *args):
"""Call a middleware method locally via `midclt`, the supported JSON-RPC transport
that replaces the deprecated /api/v2.0 REST API (removed in TrueNAS 26.04). Must run
on the TrueNAS host. Each arg is JSON-encoded (a dict for create; none for queries).
Exits with a clear message on failure."""
cmd = ["midclt", "call", method] + [json.dumps(a) for a in args]
"""Call a middleware method on the local host.
Uses `truenas_api_client` -- the library that backs `midclt` itself -- rather
than shelling out to `midclt`.
This is a SECURITY requirement, not a style choice. `midclt call <method>
<json>` puts its arguments in the process's **argv**, and `cloud_backup.create`
carries the restic repository password. argv is world-readable via `ps`, so
shelling out would expose the key to the entire backup repo to every local
user for the duration of the call. Going through the client library keeps it
in this process's memory.
"""
try:
proc = subprocess.run(cmd, capture_output=True, text=True, timeout=120)
except FileNotFoundError:
print("ERROR: `midclt` not found — run this script ON the TrueNAS host.",
file=sys.stderr)
from truenas_api_client import Client
except ImportError:
print(
"ERROR: `truenas_api_client` not importable — run this script ON the\n"
" TrueNAS host. (It ships with midclt.)",
file=sys.stderr,
)
sys.exit(1)
except subprocess.SubprocessError as exc:
print(f"ERROR: midclt call failed: {exc}", file=sys.stderr)
try:
with Client() as client:
return client.call(method, *args)
except Exception as exc: # noqa: BLE001 - surface any middleware error verbatim
# Never echo `args` here: for cloud_backup.create it contains the password.
print(f"ERROR: {method}: {exc}", file=sys.stderr)
sys.exit(1)
if proc.returncode != 0:
print(f"ERROR: midclt {method}: {(proc.stderr or proc.stdout).strip()}",
file=sys.stderr)
sys.exit(1)
out = proc.stdout.strip()
return json.loads(out) if out else None
# ── Sub-commands ──────────────────────────────────────────────────────────────
@@ -84,8 +95,11 @@ def midclt_call(method, *args):
def _middlewared_start_epoch():
"""Epoch timestamp of the running middlewared main process, or None."""
try:
# Partial path (S607) is fine here: this runs as root on TrueNAS, so an
# attacker who can poison PATH already has root. Hard-coding a path would
# be less portable (/bin vs /usr/bin) for no security gain.
pid = int(subprocess.run(
["systemctl", "show", "--property=MainPID", "--value", "middlewared"],
["systemctl", "show", "--property=MainPID", "--value", "middlewared"], # noqa: S607
capture_output=True, text=True, timeout=10, check=True,
).stdout.strip())
if pid <= 0:
@@ -213,6 +227,37 @@ def cmd_list_tasks(_args):
print(f"{t['id']:>4} {enabled:<8} {ptype:<14} {t.get('description', '')}")
def _resolve_password(args):
"""Get the restic repo password without writing it to the user's shell history.
That password is the key to the whole backup repository. `--password <secret>`
persists it in ~/.bash_history and exposes it in `ps` for the lifetime of the
shell command, so it is accepted but warned about; stdin and an interactive
prompt are the safe paths.
"""
if args.password_stdin:
if args.password:
print("ERROR: use either --password or --password-stdin, not both.",
file=sys.stderr)
sys.exit(1)
password = sys.stdin.readline().rstrip("\n")
elif args.password:
print(
"WARNING: --password puts the restic repository password in your shell\n"
" history. Prefer: echo -n 'pw' | ... --password-stdin",
file=sys.stderr,
)
password = args.password
else:
password = getpass.getpass("Restic repository password: ")
if not password:
print("ERROR: the restic repository password must not be empty.",
file=sys.stderr)
sys.exit(1)
return password
def cmd_create(args):
parts = args.schedule.split()
if len(parts) != 5:
@@ -223,6 +268,8 @@ def cmd_create(args):
sys.exit(1)
minute, hour, dom, month, dow = parts
password = _resolve_password(args)
body = {
"description": args.name,
"path": args.path,
@@ -231,7 +278,7 @@ def cmd_create(args):
"bucket": args.bucket,
"folder": args.folder,
},
"password": args.password,
"password": password,
"keep_last": args.keep_last,
"transfer_setting": args.transfer_setting,
"schedule": {
@@ -297,8 +344,11 @@ def main():
help="Bucket (S3) or container (B2) name")
c.add_argument("--folder", default="",
help="Path within the bucket (default: root)")
c.add_argument("--password", required=True,
help="Restic repository encryption password (choose a strong one)")
c.add_argument("--password", default=None,
help="Restic repository password. UNSAFE: it lands in your shell "
"history. Prefer --password-stdin, or omit both and be prompted.")
c.add_argument("--password-stdin", action="store_true",
help="Read the restic repository password from stdin (recommended)")
c.add_argument("--keep-last", type=int, default=14, metavar="N",
help="Snapshots to retain after each run (default: 14)")
c.add_argument("--schedule", default="0 2 * * *",
+137
View File
@@ -0,0 +1,137 @@
#!/usr/bin/env python3
"""Apply and revert truecloud-patch's blocks in middlewared's modules.
Every patch this project makes to a middlewared module is an appended block that
begins with the MARKER line. That makes patching idempotent (truncate at the
marker, re-append) and reverting exact (truncate at the marker, stop).
This is the single implementation of that. It used to live in two places --
apply.sh's heredoc and an inline heredoc in uninstall.sh -- and the uninstall copy
was the untested one.
python3 mw_patch.py revert-all # remove every block + the nested module
python3 mw_patch.py revert-nested # remove only the nested module's blocks
"""
from __future__ import annotations
import os
import sys
MARKER = "\n# TRUECLOUD_PATCH"
#: Modules the providers module (B2/S3) patches.
PROVIDER_RELPATHS = [
("rclone", "remote", "b2.py"),
("plugins", "cloud_backup", "restic.py"),
]
#: Modules the nested-snapshot module patches. Order matters on revert -- see
#: revert(): the loadable module goes first.
NESTED_RELPATHS = [
("plugins", "cloud", "crud.py"),
("plugins", "cloud_backup", "sync.py"),
("plugins", "cloud", "snapshot.py"),
]
#: The importable module the nested blocks depend on.
NESTED_MODULE = ("plugins", "cloud", "_truecloud_nested.py")
def patch_file(path, block):
"""Append `block`, replacing any block we appended before. Idempotent."""
with open(path, encoding="utf-8") as fh:
content = fh.read()
idx = content.find(MARKER)
base = content[:idx] if idx != -1 else content
with open(path, "w", encoding="utf-8") as fh:
fh.write(base.rstrip("\n") + "\n" + block)
def unpatch_file(path):
"""Strip our appended block, restoring the stock file. True if it was patched."""
try:
with open(path, encoding="utf-8") as fh:
content = fh.read()
except OSError:
return False
idx = content.find(MARKER)
if idx == -1:
return False
try:
with open(path, "w", encoding="utf-8") as fh:
fh.write(content[:idx].rstrip("\n") + "\n")
except OSError:
return False
return True
def revert(mw_dir, relpaths, module_relpath=None):
"""Remove our blocks from `relpaths`, and the module at `module_relpath`.
The module is deleted FIRST. Every injected block is guarded by
`if _tc_nested is not None`, so once the module is gone the blocks all no-op
even if a later unpatch fails -- the stock guard comes back regardless.
Returns the names of what was actually reverted.
"""
reverted = []
if module_relpath:
try:
os.unlink(os.path.join(mw_dir, *module_relpath))
reverted.append(module_relpath[-1])
except OSError:
pass
for rel in relpaths:
if unpatch_file(os.path.join(mw_dir, *rel)):
reverted.append(rel[-1])
return reverted
def revert_nested(mw_dir):
"""Undo the nested-snapshot patch only. Leaves the providers patch alone.
restic.py also carries a block, but it belongs to the providers module --
reverting it would silently break B2 backups.
"""
return revert(mw_dir, NESTED_RELPATHS, NESTED_MODULE)
def revert_all(mw_dir):
"""Undo every patch this project applies."""
return revert(mw_dir, NESTED_RELPATHS + PROVIDER_RELPATHS, NESTED_MODULE)
def find_middlewared_dir():
"""Directory of the installed `middlewared` package, or None."""
try:
import middlewared
except ImportError:
return None
return os.path.dirname(os.path.abspath(middlewared.__file__))
def main(argv):
if len(argv) < 2 or argv[1] not in ("revert-all", "revert-nested"):
print(__doc__, file=sys.stderr)
return 2
mw_dir = find_middlewared_dir()
if mw_dir is None:
print(" middlewared not importable — nothing to revert.")
return 0
fn = revert_all if argv[1] == "revert-all" else revert_nested
reverted = fn(mw_dir)
if reverted:
print(" Reverted: " + ", ".join(reverted))
else:
print(" Nothing to revert (overlay already removed, or never patched).")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))
+37 -4
View File
@@ -9,8 +9,8 @@ in the minified JS in one of two forms depending on TrueNAS / Angular version:
TrueNAS 24.x (static inline array):
"filterByProviders",["STORJ_IX"]
TrueNAS 25.x+ (Angular pureFunction binding):
"filterByProviders",pe(115,Rn,i.CloudSyncProviderName.Storj)
TrueNAS 25.x+ (Angular pureFunction binding, inside a chained property call):
c(2,"filterByProviders",pe(115,Rn,i.CloudSyncProviderName.Storj))("required",!0)
Both are replaced so the dropdown includes S3 and B2. The file is backed up
before modification so uninstall.sh can restore it.
@@ -33,11 +33,21 @@ WEBUI_CANDIDATES = [
# Patterns tried in order; the first match wins.
# Each entry is (compiled_regex, replacement_string).
_PATTERNS = [
# TrueNAS 25.x+: Angular emits a pureFunction call instead of a literal array.
# TrueNAS 25.x+: Angular emits a pureFunction call instead of a literal array,
# inside a CHAINED property binding — so the call is followed by two closing
# parens, one for pe(...) and one for the property(...) it sits in:
#
# c(2,"filterByProviders",pe(115,Rn,i.CloudSyncProviderName.Storj))("required",!0)
# ^^
# The pattern consumes both and re-emits one, leaving the paren balance
# unchanged. Getting that wrong is a syntax error in the bundle and the whole
# web UI goes blank — see tests/test_patch_ui.py.
#
# The minified names (pe / slot index / Rn / i) change across builds;
# CloudSyncProviderName.Storj is stable because it is a TypeScript enum name.
(re.compile(r'("filterByProviders",)\w+\(\d+,\w+,\w+\.CloudSyncProviderName\.Storj\)\)'),
r'\1["STORJ_IX","S3","B2"])'),
# TrueNAS 24.x and earlier: static inline array.
(re.compile(r'("filterByProviders",)\["STORJ_IX"\]'),
r'\1["STORJ_IX","S3","B2"]'),
@@ -55,6 +65,11 @@ def _match_pattern(content):
return None, None
def _paren_delta(s):
"""Net parenthesis balance. Patching must not change it — see main()."""
return s.count("(") - s.count(")")
def find_bundle():
"""
Search WEBUI_CANDIDATES for the JS chunk containing the filterByProviders
@@ -125,6 +140,24 @@ def main():
)
return
# Never write JS whose parentheses we have unbalanced. A pattern that eats one
# paren too many is a syntax error in the bundle and the entire TrueNAS web UI
# goes blank -- and because MARKER is then present, every later run reports
# "already patched" and skips, so the patch cannot heal itself. Recovery means
# hand-restoring the .pre-truecloud-patch backup.
#
# This is not hypothetical: it shipped once. Refuse instead.
if _paren_delta(patched) != _paren_delta(content):
print(
"[truecloud-patch] ERROR: the replacement would unbalance the bundle's "
"parentheses — refusing to write.\n"
"[truecloud-patch] The UI is UNCHANGED and still works. This means the "
"pattern no longer fits this TrueNAS build.\n"
"[truecloud-patch] File an issue at "
"https://github.com/sudolulo/truenas-truecloud-patch"
)
return
tmp = path + ".tmp"
try:
with open(tmp, "w", encoding="utf-8") as fh:
+27 -1
View File
@@ -279,7 +279,11 @@ def current_mounts_under(root, mounts_file="/proc/self/mounts"):
def _run(cmd):
return subprocess.run(cmd, capture_output=True, text=True, check=False)
# List form, never shell=True: `cmd` is built from our own mount plan, so ZFS
# dataset names cannot inject. Runs as root by definition (it mounts).
return subprocess.run( # noqa: S603
cmd, capture_output=True, text=True, check=False
)
def apply_plan(mounts, runner=_run, isdir=os.path.isdir):
@@ -371,6 +375,28 @@ async def delete_snapshot_tree(middleware, snapshot, logger=None):
"""
dataset = snapshot.partition("@")[0]
# Fast path: ONE recursive delete removes the parent and every child that
# `zfs snapshot -r` created (252 on a real pool). Deleting them individually
# also works, but it is neither cheap nor atomic -- a run killed part-way
# through 252 sequential deletes leaves exactly the orphans this function
# exists to prevent.
try:
await middleware.call("zfs.snapshot.delete", snapshot, {"recursive": True})
return
except Exception as e: # noqa: BLE001 - fall through to the explicit sweep
# Usually just "parent already gone" (stock's finally won the race once our
# mounts were released), which the sweep below handles. Log it rather than
# swallow it: if the real cause is something else, this is the only place
# it is visible -- the sweep would report a different, downstream failure.
if logger:
logger.debug(
"truecloud-patch: recursive delete of %s failed (%r); sweeping "
"the tree by name instead", snapshot, e,
)
# The parent may already be gone -- stock's `finally` can win the race once
# our mounts are released -- which fails the recursive delete while the
# children survive. Sweep them by name.
try:
snaps = await middleware.call(
"zfs.snapshot.query", [["name", "^", dataset]], {"select": ["name"]}
+1 -1
View File
@@ -17,7 +17,7 @@
# bash /mnt/tank/truenas-truecloud-patch/patch/apply.sh
# systemctl restart middlewared
VERSION="0.3.0"
VERSION="0.4.1"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
+28 -1
View File
@@ -263,10 +263,37 @@ class TestOptIn:
def test_patching_is_skipped_entirely_when_disabled(self):
# The guard-relaxing crud.py patch must be inside the enabled branch.
src = heredoc_source()
gate = src.index("if not nested_enabled:")
gate = src.index("if not nested_needed:")
crud = src.index("patch_file(crud_py, CRUD_BLOCK)")
assert gate < crud, "crud.py patch must sit inside the opt-in branch"
def test_disabling_REVERTS_the_patch_rather_than_merely_skipping_it(self):
"""Skipping is not disabling.
The overlay persists for the whole boot, so a patch applied by an earlier
run this boot is still on disk — and middlewared re-imports it on the
restart install.sh performs. Without an active revert,
`--disable-nested-snapshots` reports "disabled" while the feature keeps
running until the next reboot.
"""
src = heredoc_source()
# The implementation lives in patch/mw_patch.py (see test_mw_patch.py);
# apply.sh must import and actually call it.
assert "from mw_patch import patch_file, revert_nested" in src
gate = src.index("if not nested_needed:")
revert = src.index("reverted = revert_nested(")
patch = src.index("patch_file(crud_py, CRUD_BLOCK)")
assert gate < revert < patch, "revert belongs in the not-needed branch"
def test_import_failure_skips_the_patch_rather_than_crashing(self):
# apply.sh runs at PREINIT. If mw_patch.py cannot be imported it must
# degrade to "middlewared starts stock", never take the boot down.
src = heredoc_source()
i = src.index("from mw_patch import")
tail = src[i:i + 400]
assert "except ImportError" in tail
assert "skipping backend patch" in tail
def test_guard_is_relaxed_only_after_traversal_is_installed():
# Ordering in apply.sh is a safety property: copy module -> patch snapshot.py
+107
View File
@@ -0,0 +1,107 @@
"""Tests for create_task.py, focused on the restic repository password.
That password is the encryption key for the entire cloud backup repository. It
used to travel through `midclt call cloud_backup.create '<json>'` -- i.e. through
the subprocess's **argv**, which is world-readable via `ps` -- and `--password`
wrote it into the user's shell history forever.
"""
import importlib.util
import io
import os
import sys
import types
import pytest
SPEC = importlib.util.spec_from_file_location(
"create_task",
os.path.join(os.path.dirname(__file__), "..", "patch", "create_task.py"),
)
def load():
mod = importlib.util.module_from_spec(SPEC)
SPEC.loader.exec_module(mod)
return mod
class Args:
def __init__(self, password=None, password_stdin=False):
self.password = password
self.password_stdin = password_stdin
class TestPasswordNeverReachesArgv:
"""The whole reason this module talks to the client library."""
def test_midclt_call_spawns_no_subprocess(self):
import inspect
src = inspect.getsource(load().midclt_call)
assert "subprocess" not in src, (
"shelling out to `midclt` puts cloud_backup.create's JSON -- including "
"the restic repo password -- into argv, which any local user can read "
"with ps"
)
assert "truenas_api_client" in src
def test_errors_never_echo_the_call_arguments(self):
# A failed cloud_backup.create must not print the body back at the user;
# it contains the password.
import inspect
src = inspect.getsource(load().midclt_call)
assert "{args}" not in src
assert "args!r" not in src
class TestResolvePassword:
def test_reads_from_stdin(self, monkeypatch):
mod = load()
monkeypatch.setattr(sys, "stdin", io.StringIO("s3cret\n"))
assert mod._resolve_password(Args(password_stdin=True)) == "s3cret"
def test_strips_only_the_trailing_newline(self, monkeypatch):
# A password may legitimately contain spaces; only the line ending goes.
mod = load()
monkeypatch.setattr(sys, "stdin", io.StringIO(" pass word \n"))
assert mod._resolve_password(Args(password_stdin=True)) == " pass word "
def test_cli_password_still_works_but_warns(self, monkeypatch, capsys):
mod = load()
pw = mod._resolve_password(Args(password="cli-secret"))
assert pw == "cli-secret"
assert "shell" in capsys.readouterr().err.lower(), "must warn about history"
def test_prompts_when_neither_flag_given(self, monkeypatch):
mod = load()
monkeypatch.setattr(
mod, "getpass", types.SimpleNamespace(getpass=lambda _p: "prompted")
)
assert mod._resolve_password(Args()) == "prompted"
def test_rejects_both_flags(self, monkeypatch):
mod = load()
monkeypatch.setattr(sys, "stdin", io.StringIO("x\n"))
with pytest.raises(SystemExit):
mod._resolve_password(Args(password="a", password_stdin=True))
def test_rejects_an_empty_password(self, monkeypatch):
# An empty restic password would silently create an unencrypted-ish repo.
mod = load()
monkeypatch.setattr(sys, "stdin", io.StringIO("\n"))
with pytest.raises(SystemExit):
mod._resolve_password(Args(password_stdin=True))
class TestVersion:
def test_version_is_not_stale(self):
# __version__ sat at 0.2.0 through three releases because the drift check
# only looked at VERSION= in shell scripts. It covers this file now.
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "tools"))
from release_notes import normalise, script_versions
repo = os.path.join(os.path.dirname(__file__), "..")
versions = {normalise(v) for v in script_versions(repo).values()}
assert len(versions) == 1, f"version drift: {sorted(versions)}"
+160
View File
@@ -0,0 +1,160 @@
"""Tests for mw_patch — the single implementation of apply/revert.
apply.sh and uninstall.sh both go through this. It used to be duplicated in an
untested shell heredoc, which is exactly how the two could have drifted apart:
apply.sh reverting one set of files and uninstall.sh another.
"""
import os
import sys
import pytest
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "patch"))
from mw_patch import ( # noqa: E402
MARKER,
NESTED_MODULE,
NESTED_RELPATHS,
PROVIDER_RELPATHS,
patch_file,
revert_all,
revert_nested,
unpatch_file,
)
STOCK = "import os\n\n\ndef stock():\n return 1\n"
BLOCK = "\n# TRUECLOUD_PATCH\ninjected = 1\n"
def build_mw(root):
"""A fake middlewared tree with every file this project touches."""
mw = os.path.join(root, "middlewared")
for rel in NESTED_RELPATHS + PROVIDER_RELPATHS:
path = os.path.join(mw, *rel)
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "w", encoding="utf-8") as fh:
fh.write(STOCK)
with open(os.path.join(mw, *NESTED_MODULE), "w", encoding="utf-8") as fh:
fh.write("# module\n")
return mw
def read(mw, rel):
with open(os.path.join(mw, *rel), encoding="utf-8") as fh:
return fh.read()
class TestPatchFile:
def test_appends_the_block(self, tmp_path):
p = tmp_path / "m.py"
p.write_text(STOCK)
patch_file(str(p), BLOCK)
assert MARKER in p.read_text()
assert p.read_text().startswith("import os")
def test_is_idempotent(self, tmp_path):
# apply.sh runs on EVERY boot. Without truncate-then-append, repeated runs
# would stack duplicate copies of the block into a middlewared module.
p = tmp_path / "m.py"
p.write_text(STOCK)
for _ in range(5):
patch_file(str(p), BLOCK)
assert p.read_text().count("# TRUECLOUD_PATCH") == 1
assert p.read_text().count("injected = 1") == 1
def test_round_trips_back_to_stock(self, tmp_path):
p = tmp_path / "m.py"
p.write_text(STOCK)
patch_file(str(p), BLOCK)
assert unpatch_file(str(p)) is True
assert p.read_text() == STOCK
class TestUnpatchFile:
def test_returns_false_on_an_unpatched_file(self, tmp_path):
p = tmp_path / "m.py"
p.write_text(STOCK)
assert unpatch_file(str(p)) is False
assert p.read_text() == STOCK
def test_returns_false_on_a_missing_file(self, tmp_path):
assert unpatch_file(str(tmp_path / "nope.py")) is False
class TestRevertNested:
def test_reverts_only_the_nested_files(self, tmp_path):
mw = build_mw(str(tmp_path))
for rel in NESTED_RELPATHS + PROVIDER_RELPATHS:
patch_file(os.path.join(mw, *rel), BLOCK)
reverted = revert_nested(mw)
for rel in NESTED_RELPATHS:
assert read(mw, rel) == STOCK, f"{rel[-1]} should be stock"
assert rel[-1] in reverted
def test_never_touches_the_providers_patch(self, tmp_path):
# restic.py carries a TRUECLOUD_PATCH block too, but it belongs to the
# providers module. Reverting it would silently break B2 backups.
mw = build_mw(str(tmp_path))
for rel in NESTED_RELPATHS + PROVIDER_RELPATHS:
patch_file(os.path.join(mw, *rel), BLOCK)
revert_nested(mw)
for rel in PROVIDER_RELPATHS:
assert MARKER in read(mw, rel), f"{rel[-1]} must keep its providers block"
def test_removes_the_module(self, tmp_path):
mw = build_mw(str(tmp_path))
assert os.path.exists(os.path.join(mw, *NESTED_MODULE))
reverted = revert_nested(mw)
assert not os.path.exists(os.path.join(mw, *NESTED_MODULE))
assert "_truecloud_nested.py" in reverted
def test_module_is_removed_before_the_files_are_unpatched(self, tmp_path):
# Every injected block is guarded by `if _tc_nested is not None`, so once
# the module is gone they all no-op — the stock guard is restored even if
# a later unpatch fails.
mw = build_mw(str(tmp_path))
for rel in NESTED_RELPATHS:
patch_file(os.path.join(mw, *rel), BLOCK)
reverted = revert_nested(mw)
assert reverted[0] == "_truecloud_nested.py"
def test_is_idempotent(self, tmp_path):
mw = build_mw(str(tmp_path))
for rel in NESTED_RELPATHS:
patch_file(os.path.join(mw, *rel), BLOCK)
revert_nested(mw)
assert revert_nested(mw) == []
class TestRevertAll:
def test_reverts_providers_and_nested(self, tmp_path):
mw = build_mw(str(tmp_path))
for rel in NESTED_RELPATHS + PROVIDER_RELPATHS:
patch_file(os.path.join(mw, *rel), BLOCK)
revert_all(mw)
for rel in NESTED_RELPATHS + PROVIDER_RELPATHS:
assert read(mw, rel) == STOCK, f"{rel[-1]} should be stock"
assert not os.path.exists(os.path.join(mw, *NESTED_MODULE))
def test_is_a_noop_on_a_stock_tree(self, tmp_path):
mw = build_mw(str(tmp_path))
os.unlink(os.path.join(mw, *NESTED_MODULE))
assert revert_all(mw) == []
class TestTargetsAreDisjoint:
def test_no_file_is_in_both_module_lists(self):
# If restic.py ever appeared in NESTED_RELPATHS, revert_nested would break
# B2 backups.
assert not set(NESTED_RELPATHS) & set(PROVIDER_RELPATHS)
@pytest.mark.parametrize("rel", PROVIDER_RELPATHS)
def test_provider_targets_are_not_nested_targets(self, rel):
assert rel not in NESTED_RELPATHS
+171
View File
@@ -0,0 +1,171 @@
"""Tests for the Angular bundle patch.
This is the one part of the patch that edits *minified third-party JavaScript* by
regex, so it is the easiest place to silently produce a broken bundle: a pattern
that matches nothing leaves the dropdown Storj-only, and a pattern that matches
sloppily can unbalance the parentheses and take the whole web UI down.
Nothing checked it until now. The snippets below are verbatim from a real
TrueNAS 25.x bundle (chunk-*.js, pre-patch).
"""
import os
import re
import sys
import pytest
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "patch"))
from patch_ui import MARKER, _PATTERNS, _match_pattern # noqa: E402
# Verbatim from /usr/share/truenas/webui/chunk-FX2QXNQU.js on TrueNAS 25.10.
# Angular emits the binding as a chained ɵɵproperty(...)(...) call, so the
# pureFunction call is followed by TWO closing parens: one for pe(...), one for
# property(...).
REAL_25X = (
'c(2,"filterByProviders",pe(115,Rn,i.CloudSyncProviderName.Storj))'
'("required",!0),r(3'
)
# TrueNAS 24.x and earlier emitted a literal array.
REAL_24X = 'c(2,"filterByProviders",["STORJ_IX"])("required",!0),r(3'
def apply_patch(content):
"""Run the same match-and-substitute main() does."""
find, replace = _match_pattern(content)
assert find is not None, "no pattern matched"
patched, count = find.subn(replace, content)
return patched, count
def paren_delta(s):
"""Net paren balance. The snippets are fragments of a minified file, so they
are not balanced on their own -- what must hold is that patching does not
CHANGE the balance. Consuming one paren too many is a syntax error in the
bundle, and the whole TrueNAS web UI goes blank."""
return s.count("(") - s.count(")")
@pytest.mark.parametrize("source", [REAL_25X, REAL_24X], ids=["25.x", "24.x"])
class TestAgainstRealBundles:
def test_matches_exactly_once(self, source):
# main() refuses to write unless count == 1 — more than one match would
# mean the pattern is too loose to trust against a minified bundle.
_patched, count = apply_patch(source)
assert count == 1
def test_result_contains_all_three_providers(self, source):
patched, _ = apply_patch(source)
assert MARKER in patched
assert '"filterByProviders",["STORJ_IX","S3","B2"]' in patched
def test_patch_does_not_change_paren_balance(self, source):
# Consuming one paren too many (or too few) is a syntax error in the
# bundle and the entire TrueNAS web UI goes blank. This is the invariant
# the 25.x pattern has to get right: it eats `pe(...)` which sits inside
# a chained property(...)(...) call.
patched, _ = apply_patch(source)
assert paren_delta(patched) == paren_delta(source)
def test_surrounding_code_is_untouched(self, source):
patched, _ = apply_patch(source)
assert patched.startswith("c(2,")
assert patched.endswith('("required",!0),r(3')
def test_patch_is_idempotent(self, source):
# apply.sh re-runs every boot; MARKER short-circuits an already-patched
# file, but the pattern must also not match its own output.
patched, _ = apply_patch(source)
find, _replace = _match_pattern(patched)
if find is not None:
# Only the 24.x literal-array pattern may still "match" — and only if
# it would produce the same text. Anything else means double-patching.
again, _ = apply_patch(patched)
assert again == patched, "re-patching must be a no-op"
def test_storj_only_bundle_is_recognised():
assert _match_pattern(REAL_25X)[0] is not None
def test_unrelated_javascript_is_never_touched():
# A pattern loose enough to hit unrelated code would corrupt the bundle.
for noise in (
'c(2,"filterByProviders",pe(115,Rn,i.SomethingElse.Storj))',
'c(2,"otherBinding",pe(115,Rn,i.CloudSyncProviderName.Storj))',
'"filterByProviders"',
):
find, _ = _match_pattern(noise)
assert find is None, f"pattern must not match: {noise}"
def test_every_pattern_is_anchored_to_filterbyproviders():
# Guards against a future pattern broad enough to rewrite arbitrary JS.
for find, _replace in _PATTERNS:
assert "filterByProviders" in find.pattern
def test_patterns_compile_and_replacements_reference_group_one():
for find, replace in _PATTERNS:
assert isinstance(find, re.Pattern)
assert r"\1" in replace, "replacement must preserve the binding name"
class TestCorruptionGuard:
"""A bad pattern must never reach the bundle.
This is not hypothetical. Commit 47cdf72 shipped a pattern that consumed one
closing paren and emitted one, netting an extra `)`:
c(2,"filterByProviders",["STORJ_IX","S3","B2"]))("required",!0)
^^ syntax error
The web UI went blank. And because MARKER was then present in the file, every
subsequent run reported "already patched" and skipped — so the patch could not
heal itself, and the bundle had to be hand-restored from the backup.
"""
# Verbatim from 47cdf72.
BROKEN = (
re.compile(r'("filterByProviders",)\w+\(\d+,\w+,\w+\.CloudSyncProviderName\.Storj\)'),
r'\1["STORJ_IX","S3","B2"])',
)
def test_the_regression_that_blanked_the_ui_is_detectable(self):
find, replace = self.BROKEN
patched, count = find.subn(replace, REAL_25X)
assert count == 1, "it did match — that is why it got written"
assert paren_delta(patched) != paren_delta(REAL_25X), (
"the paren balance changes; this is the signal main() now refuses on"
)
def test_main_refuses_to_write_an_unbalanced_bundle(self, monkeypatch, tmp_path, capsys):
import patch_ui
bundle = tmp_path / "chunk-TEST.js"
bundle.write_text(REAL_25X, encoding="utf-8")
monkeypatch.setattr(patch_ui, "WEBUI_CANDIDATES", [str(tmp_path)])
monkeypatch.setattr(patch_ui, "_PATTERNS", [self.BROKEN])
patch_ui.main()
out = capsys.readouterr().out
assert "refusing to write" in out
# The bundle must be byte-for-byte untouched — a broken UI is far worse
# than an unpatched one.
assert bundle.read_text(encoding="utf-8") == REAL_25X
def test_a_good_pattern_still_writes(self, monkeypatch, tmp_path):
import patch_ui
bundle = tmp_path / "chunk-TEST.js"
bundle.write_text(REAL_25X, encoding="utf-8")
monkeypatch.setattr(patch_ui, "WEBUI_CANDIDATES", [str(tmp_path)])
patch_ui.main()
assert MARKER in bundle.read_text(encoding="utf-8")
assert (tmp_path / "chunk-TEST.js.pre-truecloud-patch").exists()
+122
View File
@@ -0,0 +1,122 @@
"""Tests for the release automation.
The release workflow refuses to publish unless these hold, so a bad tag fails
loudly in CI instead of shipping a release whose notes are empty, wrong, or whose
scripts announce a different version than the tag.
That last one is not hypothetical: VERSION= drifted to three different values
across install.sh / uninstall.sh / recover.sh / apply.sh and nothing noticed.
"""
import os
import sys
import pytest
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "tools"))
from release_notes import ( # noqa: E402
changelog_versions,
check,
extract_notes,
normalise,
script_versions,
)
REPO = os.path.join(os.path.dirname(__file__), "..")
SAMPLE = """\
# Changelog
## v0.3.0 — 2026-07-13
### Added
- the new thing
## v0.2.1 — 2026-07-09
### Fixed
- the old thing
## v0.2.0 — 2026-07-08
- first
"""
class TestExtractNotes:
def test_returns_only_that_versions_body(self):
body = extract_notes(SAMPLE, "v0.3.0")
assert "the new thing" in body
assert "the old thing" not in body
# The version heading itself is dropped (GitHub renders its own title),
# but sub-headings like "### Added" must survive.
assert not body.startswith("## v")
assert body.startswith("### Added")
def test_stops_at_the_next_version_heading(self):
body = extract_notes(SAMPLE, "v0.2.1")
assert "the old thing" in body
assert "first" not in body
def test_last_section_runs_to_end_of_file(self):
assert "first" in extract_notes(SAMPLE, "v0.2.0")
def test_accepts_the_tag_with_or_without_the_v(self):
assert extract_notes(SAMPLE, "0.3.0") == extract_notes(SAMPLE, "v0.3.0")
def test_unknown_version_raises_rather_than_returning_empty(self):
# An empty release body is worse than a failed release.
with pytest.raises(KeyError, match="no section"):
extract_notes(SAMPLE, "v9.9.9")
class TestChangelogVersions:
def test_lists_versions_newest_first(self):
assert changelog_versions(SAMPLE) == ["0.3.0", "0.2.1", "0.2.0"]
class TestAgainstTheRealRepo:
"""These run against the actual files, so drift breaks the build."""
def test_every_script_declares_a_version(self):
from release_notes import VERSIONED_FILES
found = script_versions(REPO)
missing = [f for f in VERSIONED_FILES if f not in found]
assert not missing, f"no VERSION= in: {missing}"
def test_all_scripts_agree_on_the_version(self):
versions = {normalise(v) for v in script_versions(REPO).values()}
assert len(versions) == 1, f"scripts disagree on version: {sorted(versions)}"
def test_the_current_version_has_a_changelog_section(self):
version = next(iter({normalise(v) for v in script_versions(REPO).values()}))
with open(os.path.join(REPO, "CHANGELOG.md"), encoding="utf-8") as fh:
body = extract_notes(fh.read(), version)
assert body, f"CHANGELOG.md has no content for v{version}"
def test_the_current_version_is_the_newest_changelog_entry(self):
version = next(iter({normalise(v) for v in script_versions(REPO).values()}))
with open(os.path.join(REPO, "CHANGELOG.md"), encoding="utf-8") as fh:
newest = changelog_versions(fh.read())[0]
assert newest == version, (
f"scripts say v{version} but the newest CHANGELOG entry is v{newest}"
)
def test_check_passes_for_the_current_version(self):
version = next(iter({normalise(v) for v in script_versions(REPO).values()}))
assert check(version, REPO) == []
class TestCheckCatchesMistakes:
def test_reports_a_tag_that_no_script_matches(self):
problems = check("v9.9.9", REPO)
assert problems
assert any("declares VERSION" in p for p in problems)
def test_reports_a_missing_changelog_section(self):
problems = check("v9.9.9", REPO)
assert any("no section" in p for p in problems)
+26 -8
View File
@@ -209,9 +209,16 @@ class FakeMiddleware:
if method == "zfs.snapshot.query":
return [{"name": n} for n in self.snapshots]
if method == "zfs.snapshot.delete":
if args[0] not in self.snapshots:
name = args[0]
opts = args[1] if len(args) > 1 else {}
if name not in self.snapshots:
raise RuntimeError("does not exist")
self.snapshots.remove(args[0])
if opts.get("recursive"):
# Real `zfs destroy -r` takes the parent and every child snapshot.
for n in snapshot_tree_names(name, list(self.snapshots)):
self.snapshots.remove(n)
else:
self.snapshots.remove(name)
return True
raise AssertionError(f"unexpected call {method}")
@@ -233,24 +240,35 @@ class TestDeleteSnapshotTree:
asyncio.run(delete_snapshot_tree(mw, "Tap@snap"))
assert mw.snapshots == []
def test_survives_query_failure_by_deleting_at_least_the_parent(self):
def test_uses_a_single_recursive_delete_not_252_individual_ones(self):
# 252 sequential deletes are slow AND not atomic: a run killed part-way
# through leaves exactly the orphans this function exists to prevent.
mw = FakeMiddleware(["Tap@snap", "Tap/apps@snap", "Tap/apps/lidarr@snap"])
asyncio.run(delete_snapshot_tree(mw, "Tap@snap"))
assert mw.snapshots == []
deletes = [a for m, a in mw.calls if m == "zfs.snapshot.delete"]
assert len(deletes) == 1, "should be ONE recursive call, not one per snapshot"
assert deletes[0][1] == {"recursive": True}
assert not [m for m, _a in mw.calls if m == "zfs.snapshot.query"], (
"no enumeration needed on the fast path"
)
def test_survives_recursive_and_query_failure_by_deleting_the_parent(self):
class Broken(FakeMiddleware):
async def call(self, method, *args):
if method == "zfs.snapshot.query":
raise RuntimeError("boom")
if method == "zfs.snapshot.delete" and len(args) > 1:
raise RuntimeError("recursive delete unavailable")
return await super().call(method, *args)
mw = Broken(["Tap@snap"])
asyncio.run(delete_snapshot_tree(mw, "Tap@snap"))
assert mw.snapshots == []
def test_attempts_no_delete_when_the_tree_is_already_gone(self):
# A successful query returning nothing means there is nothing to do.
# Falling back to the parent here would log a spurious "does not exist"
# warning on every clean run.
def test_leaves_unrelated_snapshots_alone_when_the_tree_is_gone(self):
mw = FakeMiddleware(["Tap@unrelated"])
asyncio.run(delete_snapshot_tree(mw, "Tap@snap"))
assert [m for m, _a in mw.calls if m == "zfs.snapshot.delete"] == []
assert mw.snapshots == ["Tap@unrelated"]
+165
View File
@@ -0,0 +1,165 @@
#!/usr/bin/env python3
"""Extract one version's section from CHANGELOG.md, and check version consistency.
Used by .github/workflows/release.yml so a release's body is always the changelog
entry -- there is no second place to write release notes, and therefore no second
place for them to be wrong.
python3 tools/release_notes.py notes v0.3.0 # -> the section body
python3 tools/release_notes.py version # -> version per the scripts
python3 tools/release_notes.py check v0.3.0 # -> exit 1 on any mismatch
"""
from __future__ import annotations
import os
import re
import sys
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CHANGELOG = os.path.join(ROOT, "CHANGELOG.md")
# Everything that announces a version must agree with everything else. They drifted
# to three different values once (0.0.4 / 0.2.1) before anything checked them --
# and create_task.py's __version__ then sat at 0.2.0 through three more releases,
# because the first version of this check only looked at VERSION= in shell scripts.
VERSIONED_FILES = [
"install.sh",
"uninstall.sh",
"recover.sh",
os.path.join("patch", "apply.sh"),
os.path.join("patch", "create_task.py"), # exposes `--version` to users
"update.sh",
]
# `VERSION="x"` (shell) or `__version__ = "x"` (python).
_VERSION_RE = re.compile(r'^(?:VERSION=|__version__\s*=\s*)"([^"]+)"', re.M)
_HEADING_RE = re.compile(r"^##\s+v?(\d+\.\d+\.\d+[^\s]*)", re.M)
def normalise(v: str) -> str:
return v.strip().lstrip("v")
def script_versions(root: str = ROOT) -> dict[str, str]:
"""VERSION= as declared by each script."""
found = {}
for rel in VERSIONED_FILES:
path = os.path.join(root, rel)
try:
with open(path, encoding="utf-8") as fh:
m = _VERSION_RE.search(fh.read())
except OSError:
continue
if m:
found[rel] = m.group(1)
return found
def changelog_versions(text: str) -> list[str]:
"""Versions with a section in the changelog, newest first."""
return [normalise(v) for v in _HEADING_RE.findall(text)]
def extract_notes(text: str, version: str) -> str:
"""The body of one version's section, without its heading.
Raises KeyError if the version has no section -- a release with an empty or
wrong body is worse than a failed release.
"""
want = normalise(version)
lines = text.splitlines()
start = None
for i, line in enumerate(lines):
m = _HEADING_RE.match(line)
if m and normalise(m.group(1)) == want:
start = i + 1
break
if start is None:
raise KeyError(f"CHANGELOG.md has no section for v{want}")
end = len(lines)
for i in range(start, len(lines)):
if _HEADING_RE.match(lines[i]):
end = i
break
return "\n".join(lines[start:end]).strip()
def check(version: str, root: str = ROOT) -> list[str]:
"""Every reason this version is not releasable. Empty list means it is."""
want = normalise(version)
problems = []
versions = script_versions(root)
for rel, got in sorted(versions.items()):
if normalise(got) != want:
problems.append(f"{rel} declares VERSION={got!r}, tag is v{want}")
missing = [r for r in VERSIONED_FILES if r not in versions]
for rel in missing:
problems.append(f"{rel} has no VERSION= line")
try:
with open(os.path.join(root, "CHANGELOG.md"), encoding="utf-8") as fh:
text = fh.read()
except OSError as e:
problems.append(f"cannot read CHANGELOG.md: {e}")
return problems
try:
body = extract_notes(text, want)
except KeyError as e:
problems.append(str(e))
else:
if not body:
problems.append(f"CHANGELOG.md section for v{want} is empty")
return problems
def main(argv):
if len(argv) < 2:
print(__doc__, file=sys.stderr)
return 2
cmd = argv[1]
if cmd == "version":
versions = set(map(normalise, script_versions().values()))
if len(versions) != 1:
print(f"scripts disagree on version: {sorted(versions)}", file=sys.stderr)
return 1
print(versions.pop())
return 0
if len(argv) < 3:
print(f"usage: {argv[0]} {cmd} <version>", file=sys.stderr)
return 2
version = argv[2]
if cmd == "notes":
# An explicit path lets update.sh show the notes from the CHANGELOG of the
# version it is about to install (`git show <tag>:CHANGELOG.md`), not the
# one already checked out.
path = argv[3] if len(argv) > 3 else CHANGELOG
with open(path, encoding="utf-8") as fh:
print(extract_notes(fh.read(), version))
return 0
if cmd == "check":
problems = check(version)
for p in problems:
print(f"::error::{p}")
if problems:
return 1
print(f"v{normalise(version)} is consistent across scripts and CHANGELOG")
return 0
print(f"unknown command: {cmd}", file=sys.stderr)
return 2
if __name__ == "__main__":
sys.exit(main(sys.argv))
+15 -1
View File
@@ -3,7 +3,7 @@
set -euo pipefail
VERSION="0.3.0"
VERSION="0.4.1"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
_HOOK_COMMENT='TrueCloud provider patch (S3/B2)'
@@ -91,6 +91,20 @@ if [ "$_ov_found" -eq 0 ]; then
fi
echo ""
# ── Revert file-level patches ─────────────────────────────────────────────────
# Unmounting the overlay is what normally reverts everything — the lower layer is
# the untouched /usr. But apply.sh only mounts an overlay when the directory is
# read-only; on a writable /usr it patches the real files in place. Uninstall
# would then remove the boot hook and report success while leaving every patch
# applied. Strip our appended blocks explicitly.
# Same implementation apply.sh uses (patch/mw_patch.py) — a second shell copy of
# this would be the untested one.
echo "Reverting any file-level patches ..."
python3 "$PATCH_DIR/patch/mw_patch.py" revert-all || \
echo " WARNING: could not revert file-level patches."
echo ""
# ── Unmount nested-snapshot staging trees ─────────────────────────────────────
# These bind mounts pin their ZFS snapshots, so they must go before anything
# tries to destroy those snapshots. Deepest first.
+293
View File
@@ -0,0 +1,293 @@
#!/bin/bash
# update.sh — fetch a newer release of truecloud-patch and apply it.
#
# ── RUN THIS BY HAND. NEVER FROM CRON OR A SYSTEMD TIMER. ─────────────────────
#
# This patch injects Python into middlewared and re-applies itself at every boot.
# An unattended pull would let any bad upstream commit reach your box with no
# human in the loop, and take effect on the next reboot. That is not theoretical:
# v0.0.4 shipped a boot-time bug that took every app on the box down.
#
# The manual step IS the safety gate. Keep it.
#
# By default this updates to the newest RELEASE TAG, not to main. main can be
# mid-refactor; a tag is the tested artifact. Use --main only if you know why.
#
# bash update.sh # to the newest release, with a confirmation
# bash update.sh --check # show what would happen; change nothing
# bash update.sh --rollback # undo the last update
set -euo pipefail
VERSION="0.4.1"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
_PREV_FILE="$PATCH_DIR/.update_previous"
_target=""
_use_main=0
_assume_yes=0
_check_only=0
_rollback=0
usage() {
cat <<USAGE
Usage: bash update.sh [options]
Options:
--to <ref> Update to a specific tag or commit (default: newest release tag)
--main Update to origin/main — UNRELEASED code, no guarantees
--check Show what an update would do and exit; changes nothing
--rollback Return to the revision recorded before the last update
--yes, -y Skip the confirmation prompt
-h, --help Show this help
Updating preserves your nested-snapshot opt-in setting either way.
USAGE
}
# An UNTRACKED file that the target tracks makes `git checkout` abort. The dirty-
# tree check deliberately ignores untracked files, so this slips past it and the
# checkout then dies mid-operation. Not hypothetical: a hand-copied
# patch/wait_restart.sh blocked a pull on a real box exactly this way.
#
# Used by BOTH the update and the rollback path -- rolling back moves the tree too,
# and would hit the identical failure.
_abort_if_untracked_blockers() {
local ref="$1" blocking
# Set intersection of {untracked, not ignored} and {tracked by the target}. Two
# git calls, not one `ls-files --error-unmatch` per file in the target tree.
# --exclude-standard is deliberate: git silently overwrites *ignored* files on
# checkout, so those are not blockers — only untracked-and-not-ignored ones are.
blocking="$(comm -12 \
<(git ls-files --others --exclude-standard | sort) \
<(git ls-tree -r --name-only "$ref" | sort) \
| sed 's/^/ /')"
[ -n "$blocking" ] || return 0
echo "ERROR: these untracked files would be overwritten:" >&2
printf '%s\n\n' "$blocking" >&2
echo " They exist here but git does not track them — most likely hand-copied" >&2
echo " or scp'd in. Move or delete them, then re-run." >&2
# "Delete update.sh, then re-run update.sh" is impossible. If the script itself
# is a blocker, it was hand-copied in to bootstrap; the honest answer is to
# bootstrap with git instead, which installs it properly.
case "$blocking" in
*update.sh*)
echo "" >&2
echo " update.sh itself is untracked here — you copied it in to bootstrap." >&2
echo " Do that with git instead, once; it installs update.sh properly:" >&2
echo "" >&2
echo " rm -f $PATCH_DIR/update.sh" >&2
echo " git -C $PATCH_DIR checkout $ref" >&2
echo " bash $PATCH_DIR/install.sh" >&2
echo "" >&2
echo " Every later update is then just: bash update.sh" >&2
;;
esac
exit 1
}
while [ $# -gt 0 ]; do
case "$1" in
--to)
if [ -z "${2:-}" ]; then
echo "ERROR: --to needs a tag, branch, or commit." >&2
exit 1
fi
_target="$2"; shift ;;
--main) _use_main=1 ;;
--check) _check_only=1 ;;
--rollback) _rollback=1 ;;
--yes|-y) _assume_yes=1 ;;
-h|--help) usage; exit 0 ;;
*) echo "ERROR: unknown option: $1" >&2; echo "" >&2; usage >&2; exit 1 ;;
esac
shift
done
echo "=== TrueNAS TrueCloud Provider Patch — Update (v${VERSION}) ==="
echo ""
# ── Preflight ─────────────────────────────────────────────────────────────────
if [ "$(id -u)" -ne 0 ]; then
echo "ERROR: must be run as root (install.sh needs it)." >&2
exit 1
fi
cd "$PATCH_DIR"
if ! git rev-parse --git-dir >/dev/null 2>&1; then
echo "ERROR: $PATCH_DIR is not a git clone — nothing to update." >&2
echo " Re-clone from https://github.com/sudolulo/truenas-truecloud-patch" >&2
exit 1
fi
# Past `sudo git pull`s can leave root-owned objects in .git that then break any
# non-root git command. We run as root, so we would only make that worse.
_owner="$(stat -c '%U' "$PATCH_DIR")"
if [ -n "$_owner" ] && [ "$_owner" != "root" ]; then
chown -R "$_owner" "$PATCH_DIR/.git" 2>/dev/null || true
fi
# A dirty tree means someone edited or scp'd files in place; merging over that
# silently loses their changes, or conflicts halfway through.
if [ -n "$(git status --porcelain --untracked-files=no)" ]; then
echo "ERROR: the working tree has uncommitted changes:" >&2
git status --short --untracked-files=no >&2
echo "" >&2
echo " Refusing to update over them. Commit, stash, or discard them first:" >&2
echo " git -C $PATCH_DIR checkout -- ." >&2
exit 1
fi
# ── Rollback ──────────────────────────────────────────────────────────────────
if [ "$_rollback" -eq 1 ]; then
if [ ! -f "$_PREV_FILE" ]; then
echo "ERROR: no previous revision recorded — nothing to roll back to." >&2
exit 1
fi
_prev="$(cat "$_PREV_FILE")"
if ! git rev-parse --verify --quiet "${_prev}^{commit}" >/dev/null; then
echo "ERROR: recorded revision '$_prev' is not a valid commit." >&2
echo " The history may have been rewritten. Pick a target explicitly:" >&2
echo " bash update.sh --to <tag>" >&2
exit 1
fi
_abort_if_untracked_blockers "$_prev"
echo "Rolling back to $_prev ..."
git checkout -q --detach "$_prev"
echo "Reverted. Re-applying ..."
echo ""
bash "$PATCH_DIR/install.sh"
exit 0
fi
# ── Work out where we are and where we are going ──────────────────────────────
echo "Fetching ..."
git fetch --quiet --tags --prune origin
_current="$(git rev-parse HEAD)"
_current_desc="$(git describe --tags --always 2>/dev/null || echo "$_current")"
if [ -n "$_target" ]; then
:
elif [ "$_use_main" -eq 1 ]; then
_target="origin/main"
else
# Newest release tag by VERSION order, not by tag date. Date order is only
# correct while tags are created in ascending version order; it breaks the
# moment a hotfix is tagged out of band (a v0.3.6 released after v0.4.0 would
# sort as "newest" by date and silently downgrade the box).
#
# Filter to PLAIN vX.Y.Z: git's version sort ranks `v0.5.0-rc1` ABOVE `v0.5.0`
# (verified), so without this a release candidate would be installed as though
# it were the newest release. The release workflow deliberately supports
# rc/beta/alpha tags, so they will exist.
_target="$(git tag -l 'v*' --sort=-version:refname \
| grep -E '^v[0-9]+\.[0-9]+\.[0-9]+$' | head -1)"
if [ -z "$_target" ]; then
echo "ERROR: no release tags found; use --main to track unreleased code." >&2
exit 1
fi
fi
if ! _target_sha="$(git rev-parse --verify --quiet "${_target}^{commit}")"; then
echo "ERROR: '$_target' is not a valid tag, branch, or commit." >&2
exit 1
fi
echo " current: $_current_desc"
echo " target: $_target ($(git rev-parse --short "$_target_sha"))"
echo ""
if [ "$_current" = "$_target_sha" ]; then
echo "Already up to date. Nothing to do."
exit 0
fi
_abort_if_untracked_blockers "$_target_sha"
# ── Show what is coming ───────────────────────────────────────────────────────
echo "Commits you do not have yet:"
git log --oneline --no-decorate "$_current..$_target_sha" | sed 's/^/ /' || true
echo ""
# Reuse tools/release_notes.py rather than re-implementing the extractor here —
# a second copy would be the untested one. Read the CHANGELOG *of the target*, so
# the notes describe what you are about to install.
if [ -f "$PATCH_DIR/tools/release_notes.py" ] && [ "$_use_main" -eq 0 ] \
&& [ -z "${_target##v*}" ]; then
_cl="$(mktemp)"
if git show "$_target_sha:CHANGELOG.md" > "$_cl" 2>/dev/null && [ -s "$_cl" ]; then
echo "Release notes for $_target:"
python3 "$PATCH_DIR/tools/release_notes.py" notes "$_target" "$_cl" \
2>/dev/null | sed 's/^/ /' || echo " (no notes for $_target)"
echo ""
fi
rm -f "$_cl"
fi
if [ "$_use_main" -eq 1 ]; then
echo "NOTE: --main tracks UNRELEASED code. It has passed CI, but it is not a"
echo " tested release, and apply.sh runs at every boot."
echo ""
fi
if [ "$_check_only" -eq 1 ]; then
echo "--check given; nothing changed."
exit 0
fi
# ── Confirm ───────────────────────────────────────────────────────────────────
if [ "$_assume_yes" -eq 0 ]; then
printf "Apply this update and restart middlewared? [y/N] "
read -r _answer </dev/tty || _answer=""
case "$_answer" in
y|Y|yes|YES) ;;
*) echo "Aborted. Nothing changed."; exit 0 ;;
esac
echo ""
fi
# ── Apply ─────────────────────────────────────────────────────────────────────
# Record where we were BEFORE moving, so --rollback works even if install.sh dies.
echo "$_current" > "$_PREV_FILE"
echo "Checking out $_target ..."
git checkout -q --detach "$_target_sha"
echo " now at $(git describe --tags --always)"
echo ""
echo "Applying (this preserves your nested-snapshot setting) ..."
echo ""
if ! bash "$PATCH_DIR/install.sh"; then
echo ""
echo "ERROR: install.sh failed after updating." >&2
echo " Roll back with: bash $PATCH_DIR/update.sh --rollback" >&2
echo " Or disable the patch entirely: bash $PATCH_DIR/recover.sh" >&2
exit 1
fi
echo ""
echo "=== Update complete ==="
echo " $_current_desc -> $(git describe --tags --always)"
echo ""
if ! git symbolic-ref -q HEAD >/dev/null; then
echo "NOTE: the checkout is now pinned to a release tag (detached HEAD), which is"
echo " what you want for a deployment. Plain \`git pull\` will not work here —"
echo " use \`bash update.sh\` from now on."
echo ""
fi
echo "If anything looks wrong:"
echo " bash $PATCH_DIR/update.sh --rollback # back to $_current_desc"
echo " bash $PATCH_DIR/recover.sh # kill switch + restart"