Compare commits

...
16 Commits
Author SHA1 Message Date
flan 14ef25e3c6 Run CI as a single job so concurrent jobs cannot race
CI / ci (shell + python 3.11-3.13) (push) Successful in 46s
Two CI failures shared one cause: four jobs starting together on the
self-hosted runner.

act caches each action as one shared clone under /root/.cache/act/<hash> and
re-pulls it per job, so concurrent jobs fight over that directory and the
loser dies with "lstat .../<file>: no such file or directory" before any test
runs — a different victim each push. And the runner force-pulls its base image
per job, so four jobs meant four anonymous Docker Hub pulls per push; that hit
429 Too Many Requests and every job started failing before it began, including
the shell job nothing had touched.

Installing uv without an action only shrank the surface, since every job still
used actions/checkout. Concurrency is the ingredient, so this removes it: one
job cannot race itself whatever actions it uses, and one job is one pull. The
version sweep moves inside the job and still runs every version after one
fails, preserving what fail-fast: false bought.
2026-08-26 05:34:11 +00:00
flan 02ba653127 Install uv without an action so the matrix jobs stop racing
CI / python 3.12 (push) Failing after 1s
CI / shell (shellcheck + syntax) (push) Failing after 1s
CI / python 3.11 (push) Failing after 1s
CI / python 3.13 (push) Failing after 0s
act caches each action as one shared git clone under /root/.cache/act/<hash>
and re-pulls it per job. The three python matrix jobs start within the same
second on the self-hosted runner, race on that directory, and the loser dies
with "lstat /root/.cache/act/<hash>/.npmrc: no such file or directory" before
any test runs — a red main with zero suite output and a different victim each
push (3.12, then 3.11).

The action was only fetching a binary; the matrix interpreter is selected per
command by uvx --python. A run: step has no action-cache entry and cannot
race, and keeps the jobs parallel — serialising the matrix would cost 3x the
wall clock and still leave actions/checkout shared across four jobs. The uv
version is pinned under the same rule as ruff.

The accompanying test parses uses: directives rather than the raw text, so the
comment can still name the action it avoids, and uses re instead of PyYAML
because CI runs uvx pytest, whose environment holds pytest and nothing else.
2026-08-26 05:28:26 +00:00
flan 5f8d42f2cf Preserve patched_at across the post-restart re-mount
TrueNAS compatibility / compat (push) Successful in 15s
Release / release (push) Successful in 15s
CI / python 3.11 (push) Failing after 7s
CI / python 3.12 (push) Successful in 17s
CI / python 3.13 (push) Successful in 20s
CI / shell (shellcheck + syntax) (push) Successful in 9s
create_task.py verify decides "loaded" by comparing middlewared's start time
against hook_status.json's patched_at. The post-restart re-mount restores the
same patch the boot pass already applied, but it runs after the restart — so
letting apply.sh re-stamp made patched_at newer than the process that had
correctly imported the patch, and verify reported FAIL while everything was
working. That would have fired on every boot where docker's nvidia sysext
merge detaches the overlay.

Found on hardware while exercising the candidate: a new lying status
introduced by the fix for a lying status. The snapshot lives in /run, not the
repo — a leftover file there would leave the tree dirty, which update.sh
refuses to run over.
2026-08-26 05:20:23 +00:00
flan 3915f92dec Do not restart middlewared a second time on a post-restart miss
TrueNAS compatibility / compat (push) Successful in 12s
Release / release (push) Successful in 17s
CI / shell (shellcheck + syntax) (push) Successful in 8s
CI / python 3.12 (push) Successful in 21s
CI / python 3.11 (push) Successful in 21s
CI / python 3.13 (push) Successful in 24s
try-restart returns as soon as middlewared is READY, and middlewared then
brings docker up — docker.configure_nvidia merges the stock nvidia sysext
over /usr at that point, detaching the overlay after the patched modules
have already been imported. Checking the disk there reports "missing" on a
perfectly healthy system, and the retry that followed would restart a
correctly-patched middlewared straight back into the same race, then log
ERROR when nothing was wrong.

The disk is the right oracle before the restart and the wrong one after it.
After the restart the overlay is re-mounted so the next restart finds patched
files, and whether this middlewared actually holds the patch is left to the
in-process alert, which is the only thing that can answer it exactly.
2026-08-26 04:57:33 +00:00
flan d1db3a3f60 release v0.8.0
CI / shell (shellcheck + syntax) (push) Successful in 12s
CI / python 3.11 (push) Successful in 20s
CI / python 3.13 (push) Successful in 21s
TrueNAS compatibility / compat (push) Successful in 17s
CI / python 3.12 (push) Successful in 39s
Release / release (push) Successful in 20s
2026-08-26 04:37:50 +00:00
flan 0fc994f676 Re-apply and verify the patch before the deferred restart
CI / shell (shellcheck + syntax) (push) Successful in 16s
CI / python 3.12 (push) Successful in 54s
CI / python 3.11 (push) Successful in 55s
CI / python 3.13 (push) Successful in 44s
Patching at PREINIT and restarting minutes later is only sound while the
patched files are still on the live path when middlewared re-imports them,
and PREINIT cannot guarantee that. The overlay sits inside /usr, so anything
that remounts that hierarchy detaches it — a systemd-sysext merge/refresh
from another PREINIT hook, or middlewared's own docker.configure_nvidia at
runtime. Init scripts run sequentially in id order, so a hook registered
after this one always wins, and reordering them would not help because
docker.configure_nvidia fires long after PREINIT is done.

Observed on 25.10.6: the overlay was mounted at 16:41:56, a sysext refresh
unmerged and remerged /usr four seconds later, and the deferred restart at
16:47:24 loaded stock modules. Every B2 cloud_backup task then failed with
NotImplementedError for nineteen hours across four scheduled runs while
apply.log and hook_status.json both reported the patch active.

wait_restart.sh now re-applies immediately before restarting — after boot has
settled, which is also after every sysext merge and docker nvidia
configuration — verifies the marker is on the live path, restarts, and
verifies again, retrying once. It is no longer exec'd, so something can run
after the restart to find out what it loaded. apply.sh records the resolved
middlewared directory in .mw_dir for that check, and honours TRUECLOUD_REAPPLY
so the re-apply pass does not schedule a second restart.

_ensure_writable treated "one of our overlays is listed here" as "already
done", but it only reaches that check when the directory is not writable, and
a live overlay of ours always is — a shadowed overlay was indistinguishable
from a healthy one. It is now detached and re-mounted, reusing the upperdir so
files patched earlier in the boot survive, with a fresh workdir and a retry on
a private one, since overlayfs refuses a workdir a detached mount still holds.

Add a CRITICAL hourly alert for the case none of this can prevent: the patch
being on disk but not in the running process. apply.log can only report the
first. The alert asks the second question from inside middlewared, where the
patch's own stamps make it exact, and checks both halves since either can go
missing alone. It stays quiet when the kill switch is set or the providers
module has been retired as native, and is not muted by update_alerts_disabled.

wait_restart.sh also logs to apply.log now: journald retention on a busy box
is easily shorter than the interval between reboots, and the boot that caused
this had already rotated away by the time it was investigated.
2026-08-26 04:37:04 +00:00
flan 520b2d3735 Merge pull request 'docs: refresh the TrueNAS compatibility matrix' (#12) from bot/compat-matrix into main
CI / shell (shellcheck + syntax) (push) Successful in 9s
CI / python 3.12 (push) Successful in 2m3s
CI / python 3.13 (push) Successful in 2m4s
CI / python 3.11 (push) Successful in 2m7s
Reviewed-on: #12
2026-08-09 17:41:51 -04:00
truecloud-patch bot 4371797f94 docs: refresh the TrueNAS compatibility matrix 2026-08-09 06:18:05 +00:00
flan 45c89fd001 Add CI status badge to README
CI / shell (shellcheck + syntax) (push) Successful in 8s
CI / python 3.11 (push) Successful in 18s
CI / python 3.12 (push) Successful in 16s
CI / python 3.13 (push) Successful in 18s
2026-08-04 02:45:42 +00:00
flan 670dbd25f6 Remove README badge wall; move development disclosure to NOTICE
CI / python 3.13 (push) Successful in 17s
CI / shell (shellcheck + syntax) (push) Successful in 7s
CI / python 3.11 (push) Successful in 30s
CI / python 3.12 (push) Successful in 15s
2026-08-03 23:43:00 +00:00
flan 827c4358fd Skip the chmod-based unreadable-sidecar tests when running as root
CI / shell (shellcheck + syntax) (push) Successful in 8s
CI / python 3.11 (push) Successful in 20s
CI / python 3.12 (push) Successful in 13s
CI / python 3.13 (push) Successful in 17s
The Gitea runner image executes jobs as root, and chmod(0) cannot make a file
unreadable for root (CAP_DAC_OVERRIDE) — the two tests failed there while
passing on GitHub's non-root runner. Skipping as root keeps the scenario
exercised everywhere it is constructible.
2026-08-03 20:28:37 +00:00
flan ea00bd0685 Repoint forge references from git.onetick.ninja to git.arch.fyi
CI / shell (shellcheck + syntax) (push) Successful in 10s
CI / python 3.11 (push) Failing after 19s
CI / python 3.12 (push) Failing after 25s
CI / python 3.13 (push) Failing after 17s
2026-08-03 20:03:16 +00:00
flan 5a3d4288a2 Fix Gitea-runner python matrix: uv-managed interpreters, pin ruff
CI / shell (shellcheck + syntax) (push) Successful in 7s
CI / python 3.13 (push) Failing after 16s
CI / python 3.11 (push) Failing after 16s
CI / python 3.12 (push) Failing after 19s
Swap README badges to the public GitHub mirror (workflows and releases).
2026-08-03 19:47:29 +00:00
flan b364a17735 Add project badges
CI / python 3.11 (push) Failing after 14s
CI / shell (shellcheck + syntax) (push) Successful in 9s
CI / python 3.13 (push) Failing after 11s
CI / python 3.12 (push) Failing after 16s
2026-08-03 19:28:34 +00:00
flan 3e1de8ffd1 Add sponsor badges to README
CI / python 3.12 (push) Failing after 11s
CI / python 3.13 (push) Failing after 16s
CI / shell (shellcheck + syntax) (push) Successful in 7s
CI / python 3.11 (push) Failing after 10s
2026-08-03 19:09:25 +00:00
flan a7cbb3a994 Add donation links (GitHub Sponsors, Ko-fi)
CI / shell (shellcheck + syntax) (push) Successful in 7s
CI / python 3.11 (push) Failing after 15s
CI / python 3.12 (push) Failing after 25s
CI / python 3.13 (push) Failing after 13s
2026-08-03 17:24:39 +00:00
22 changed files with 1050 additions and 56 deletions
+2
View File
@@ -0,0 +1,2 @@
github: sudolulo
ko_fi: sudolulo
+70 -23
View File
@@ -9,9 +9,31 @@ on:
permissions:
contents: read
# ONE job, deliberately. This was four (shell + a 3-way python matrix) and they
# started within the same second on the self-hosted Gitea runner, which is what
# made CI unreliable in two separate ways:
#
# 1. The act action-cache race. `act` caches each ACTION as a single shared git
# clone under /root/.cache/act/<hash> and re-pulls it per job, so concurrent
# jobs using the same action fight over that directory and the loser dies
# with `lstat /root/.cache/act/<hash>/<file>: no such file or directory` --
# a red `main` with zero suite output, and a different victim each push
# (3.12 on one, 3.11 on the next). Dropping one action only shrank the
# surface: every job still used actions/checkout. Concurrency is the actual
# ingredient, so removing it removes the whole class -- a single job cannot
# race itself, no matter which actions it uses.
#
# 2. Docker Hub 429s. The runner force-pulls its base image per job, so four
# jobs meant four anonymous pulls per push. A few pushes and re-runs in an
# afternoon exhausted the anonymous limit and every job failed before it
# started -- including the shell job, which nothing had touched. One job is
# one pull.
#
# The cost is wall-clock parallelism, and this repo does not need it: the suite
# is ~1.5s, so container start and interpreter downloads dominate either way.
jobs:
shell:
name: shell (shellcheck + syntax)
ci:
name: ci (shell + python 3.11-3.13)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
@@ -31,32 +53,57 @@ jobs:
env:
SHELLCHECK_OPTS: -S warning -e SC1091
python:
name: python ${{ matrix.python }}
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
# TrueNAS SCALE middleware runs 3.11+; keep the patch importable across
# the versions it may be injected into.
python: ["3.11", "3.12", "3.13"]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python }}
- name: install dev deps
run: python -m pip install --upgrade pip pytest ruff
# uv-managed interpreters instead of actions/setup-python: the prebuilt-CPython
# download path setup-python relies on does not work on the self-hosted Gitea
# runner (all three matrix jobs failed at setup there while passing on GitHub);
# uv works identically on both.
#
# Installed by a plain `run:` step rather than astral-sh/setup-uv: one fewer
# action is one fewer thing to go wrong, and the action was only ever
# fetching a binary -- the interpreter is chosen per command by `uvx
# --python`, never by the action.
#
# Pinned for the same reason ruff is pinned below: an unpinned uv means any
# upstream release can turn main red with no code change here.
- name: install uv
env:
UV_VERSION: "0.11.21"
run: |
curl -LsSf "https://astral.sh/uv/${UV_VERSION}/install.sh" | sh
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
- name: ruff
run: ruff check patch tests tools
# Pinned: an unpinned ruff means any upstream release can turn main red
# with no code change.
run: uvx ruff@0.16.1 check patch tests tools
# TrueNAS SCALE middleware runs 3.11+; keep the patch importable across the
# versions it may be injected into. Every version runs even after one
# fails -- that is what `fail-fast: false` bought when this was a matrix,
# and losing it would mean a 3.11 break hides whether 3.12 and 3.13 are
# fine, which is exactly the information you want at that moment.
- name: pytest
run: pytest tests -v
env:
PYTHONS: "3.11 3.12 3.13"
run: |
fail=0
for v in $PYTHONS; do
echo "::group::pytest on python $v"
uvx --python "$v" pytest tests -v \
|| { echo "::error::suite failed on python $v"; fail=1; }
echo "::endgroup::"
done
exit $fail
- name: verify injected middleware blocks compile
# Belt-and-braces: the *_BLOCK strings are appended into live middlewared
# modules. A syntax error there would break the box at boot.
run: pytest tests/test_apply_blocks.py -v
env:
PYTHONS: "3.11 3.12 3.13"
run: |
fail=0
for v in $PYTHONS; do
uvx --python "$v" pytest tests/test_apply_blocks.py -v \
|| { echo "::error::injected blocks failed to compile on python $v"; fail=1; }
done
exit $fail
+1 -1
View File
@@ -124,7 +124,7 @@ jobs:
echo "--- release body ---"
cat /tmp/notes.md
# This repo is canonically hosted on Gitea (git.onetick.ninja) and mirrored to
# This repo is canonically hosted on Gitea (git.arch.fyi) and mirrored to
# GitHub, and BOTH run this workflow -- Gitea reads .github/workflows too. So
# the publish step has to work on whichever forge it lands on. Everything
# above is forge-agnostic; only the "create a release" API differs.
+4
View File
@@ -3,6 +3,10 @@
/apply.log.1
/apply.log.2
/hook_status.json
# The resolved middlewared directory, recorded by apply.sh so wait_restart.sh can
# check whether the patched modules are still on the live path without
# re-deriving site-packages.
/.mw_dir
/disabled
/nested_snapshots_enabled
# Written by apply.sh when a module's assumptions no longer fit the installed
+114 -1
View File
@@ -6,9 +6,21 @@ is deliberate: see [Releasing](docs/releasing.md). Twelve releases were cut on
live, every one of those interrupts every user. An alert people learn to ignore is
worse than no alert, because one day it carries a security fix.
## Unreleased
## v0.8.0 — 2026-08-26
### Changed
- **CI's python matrix is green on the self-hosted Gitea runner again.** The real
failure was that the Gitea runner image executes jobs as root, and the two
unreadable-sidecar tests build their scenario with `chmod(0)` — which cannot make
a file unreadable for root (`CAP_DAC_OVERRIDE`). Those two tests now skip as root
with that reason; GitHub's non-root runner still exercises them. The matrix also
moved to uv-managed interpreters (one toolchain across both runners) and ruff is
pinned to 0.16.1 so an upstream ruff release can't turn `main` red without a code
change.
- **README badges point at the public GitHub mirror** (workflow status and
releases) instead of the private forge. The release badge had also been reading
the stale Gitea v0.6.1 release instead of the current v0.7.0 on GitHub.
- **`master` is now labelled `27-dev`, because it is not the next release.** iX
branches each major onto its own `release/` line and master rolls straight on to the
one after — on 2026-07-14 every recent commit on master targeted `27.0.0-BETA.1`
@@ -27,6 +39,107 @@ worse than no alert, because one day it carries a security fix.
### Fixed
- **CI turned `main` red on two of three pushes without running a single test,
then stopped running at all.** Two failures, one cause: four concurrent jobs on
one self-hosted runner.
`act`, the engine behind the Gitea runner, caches each *action* as one shared
git clone under `/root/.cache/act/<hash>` and re-pulls it per job. Jobs
starting in the same second fight over that directory, and the loser dies with
`lstat /root/.cache/act/<hash>/.npmrc: no such file or directory` — before any
suite output exists, with a different victim each push (3.12 on one, 3.11 on
the next). Separately, the runner force-pulls its base image per job, so four
jobs meant four anonymous Docker Hub pulls per push; a few pushes and re-runs
in one afternoon hit `429 Too Many Requests` and *every* job began failing
before it started — including the shell job, which nothing had touched.
CI is now a single job. Dropping one action (uv is installed by a `run:` step
rather than `astral-sh/setup-uv`, which was only ever fetching a binary) only
shrank the surface, because every job still used `actions/checkout`.
Concurrency is the actual ingredient, so removing it removes the whole class:
one job cannot race itself whatever actions it uses, and one job is one pull.
The Python sweep moved inside that job and still runs every version after one
fails — that is what `fail-fast: false` bought, and losing it would mean a 3.11
break hides whether 3.12 and 3.13 are fine. The cost is wall-clock
parallelism, which this repo does not need: the suite is ~1.5s, so container
start and interpreter downloads dominate either way.
A red gate that is usually noise is worse than no gate, because the one time
it means something, nobody looks.
- **The patch survived being applied and then silently stopped existing, because
something else remounted `/usr` four seconds later.** On a box running
TrueNAS 25.10.6 the boot of 2026-08-19 went: 16:41:56 `apply.sh` mounts its
overlay on `/usr/lib/python3/dist-packages`, patches `b2.py`/`restic.py`, logs
every step `OK`; **16:42:00** a second PREINIT hook runs `systemd-sysext
refresh` over `/usr` — `Unmerged '/usr'.` / `Merged extensions into '/usr'.` —
and our overlay, which lives *inside* that hierarchy, is torn off with it;
16:47:24 our own deferred restart fires exactly as designed and middlewared
imports the **stock** modules. Every B2 TrueCloud Backup task then failed with
`NotImplementedError` from stock `rclone/base.py` for nineteen hours, across
four scheduled runs, while `apply.log` and `hook_status.json` both said the
patch was active.
Nothing in the patch was wrong, which is the point: applying at PREINIT and
restarting later is only sound if the patched files are still on the live path
when middlewared re-imports them, and **that is not something PREINIT can
guarantee**. Init scripts run sequentially in id order, so any hook registered
after ours always wins. Worse, hook ordering cannot fix it either —
middlewared's own `docker.configure_nvidia` merges a sysext over `/usr` at
*runtime*, long after every PREINIT hook is finished.
So the deferred restart no longer trusts the PREINIT pass. `wait_restart.sh`
now re-applies immediately before it restarts middlewared — after boot has
settled, which is also after every sysext merge and docker nvidia
configuration — and verifies the marker is genuinely on the live path before
restarting. It is no longer `exec systemctl try-restart middlewared`, because
something has to run afterwards.
What runs afterwards deliberately does **not** restart again. `try-restart`
returns as soon as middlewared is READY, and middlewared then brings docker up
— `docker.configure_nvidia` merges the stock nvidia sysext over `/usr` at that
point, detaching the overlay *after* the patched modules have already been
imported. A disk check there reports "missing" on a perfectly healthy system,
and restarting on that signal would restart a correctly-patched middlewared
straight back into the same race. So the overlay is re-mounted for the benefit
of the next restart, and the question of whether *this* middlewared actually
holds the patch is left to the one thing that can answer it exactly — the
in-process alert below. That re-mount preserves `hook_status.json`'s
`patched_at`: `create_task.py verify` decides "loaded" by comparing
middlewared's start time against that stamp, so a re-apply running *after* the
restart would have made the stamp newer than the process which correctly
imported the patch, and `verify` would have reported FAIL forever on every
boot where the sysext merge detaches the overlay. Caught on hardware while
validating the candidate — a new lying status introduced by the fix for a
lying status.
Two supporting fixes fell out of the same failure. `_ensure_writable` treated
"one of our overlays is listed on this directory" as "already done" — but it
only ever reaches that check when the directory is **not** writable, and a live
overlay of ours always is. A shadowed overlay was therefore indistinguishable
from a healthy one; it is now detached and re-mounted, reusing the same
upperdir so everything patched earlier in the boot reappears intact, with a
fresh workdir because overlayfs refuses one left behind by a detached mount.
- **middlewared now says so when it is running stock.** The gap that let this
cost nineteen hours was not the remount, it was that nothing could tell the
difference between "patched on disk" and "patched in the running process".
`apply.log` can only ever report the first. A new CRITICAL alert asks the
second question from inside middlewared, hourly, where it is exact: the patch
stamps the objects it replaces, so a missing stamp means this interpreter
imported stock code. It checks both halves — `restic.py`'s `_truecloud_patched`
marker and whether `B2RcloneRemote.get_restic_config` is still the base class's
— since either can go missing alone. It stays quiet when the kill switch is
set or the providers module has been retired as native, and it is deliberately
**not** silenced by `update_alerts_disabled`: that mutes release notifications,
not a broken backup path.
Boot-time diagnosis also no longer depends on the journal. `wait_restart.sh`
logged only to the journal, and journald retention on a busy box is easily
shorter than the interval between reboots — the 2026-08-19 boot had already
rotated away by the time it was investigated. It now writes to `apply.log`
alongside everything else.
- **The next maintenance release was never checked, and it is the one that reaches
users.** Shipped versions were discovered from `TS-*` tags and unreleased ones from
`release/*` branches carrying `-BETA`/`-RC`. A branched-but-untagged *maintenance*
+4
View File
@@ -0,0 +1,4 @@
truenas-truecloud-patch
Parts of this project were written with AI assistance (Claude); all of it is
reviewed and tested before release.
+19 -4
View File
@@ -1,5 +1,7 @@
# truenas-truecloud-patch
[![CI](https://git.arch.fyi/flan/truenas-truecloud-patch/actions/workflows/ci.yml/badge.svg)](https://git.arch.fyi/flan/truenas-truecloud-patch/actions)
Extends TrueNAS SCALE's **TrueCloud Backup** to:
- back up to **Backblaze B2 and any S3-compatible provider**, not just Storj;
@@ -59,9 +61,9 @@ If something is wrong, the reason is in `apply.log` — start at
| --- | --- | --- | --- |
| 24.10.2.4 | ok | ok | — |
| 25.04.2.6 | ok | ok | — |
| 25.10.4 | ok | ok | v0.7.0: 3 live tasks — 191-dataset nested backup of /mnt/Tap, a 215-filesystem/2-zvol backup of /mnt/Tank/backups, and a non-nested one; 0 orphans, 0 leaked mounts, byte-identical restore; the collector also reclaimed a real orphan the pool had been carrying |
| 25.10.5 | ok | ok | — |
| 24.10.2.5 _(unreleased)_ | ok | ok | — |
| 25.10.5 _(unreleased)_ | ok | ok | — |
| 25.10.6 _(unreleased)_ | ok | ok | — |
| 26.0.0-BETA.3 _(unreleased)_ | ok | ok | — |
| master _(27-dev)_ | **BROKEN** | **BROKEN** | — |
@@ -154,6 +156,17 @@ alert, which both take the newest plain `vX.Y.Z` tag. That is what lets debuggin
happen in `-rc` tags instead of in your notification bell — see
[Releasing](docs/releasing.md).
**A second alert reports the patch not being loaded**, and this one you cannot
turn off with `--no-update-alerts` — it is CRITICAL, hourly, and it means B2/S3
backup tasks are about to fail. Being patched *on disk* and being patched *in the
running middlewared* are different facts, and only middlewared can answer the
second one: the patch stamps the objects it replaces, so a missing stamp means
the process imported stock code. It fires if something detaches the patch overlay
(a `systemd-sysext` merge over `/usr`, for instance) and the self-healing re-apply
in the deferred restart could not put it back. `bash install.sh` clears it. It
stays quiet when the kill switch is set, or when the providers module has been
retired because TrueNAS went native.
The changelog is read from whichever forge `origin` points at, derived from the
remote rather than hard-coded. That is not cosmetic: when the changelog cannot be
read, the alert deliberately fires **anyway** rather than risk hiding a security
@@ -236,5 +249,7 @@ and restores the original UI bundle from backup.
- Filing a TrueNAS bug? **Remove the patch first** and reproduce on a stock system.
- Provided as-is, no warranty. See LICENSE.
Parts of this project were written with AI assistance (Claude); all of it is reviewed
and tested before release. Bugs are mine.
## Support
If truenas-truecloud-patch is useful to you, consider supporting development via
[GitHub Sponsors](https://github.com/sponsors/sudolulo) or [Ko-fi](https://ko-fi.com/sudolulo).
+31 -7
View File
@@ -74,22 +74,46 @@ Two different things must survive two different events:
and creates a transient systemd unit (`truecloud-mw-restart`, via
`systemd-run --no-block`) running `patch/wait_restart.sh` — detached so it
cannot disrupt the remainder of the boot sequence.
5. **Once boot has settled, middlewared restarts once** and imports the
patched modules from the overlay. `wait_restart.sh` holds the restart until
the systemd boot job queue has drained (so in-flight `ix-*` units like
`ix-reporting` finish first) *and* middlewared's docker/apps startup has
reached a terminal state — plain unit ordering cannot see either, and
restarting middlewared while they run kills apps and dashboard reporting
5. **Once boot has settled, the patch is re-applied and middlewared restarts
once**, importing the patched modules from the overlay. `wait_restart.sh`
holds the restart until the systemd boot job queue has drained (so in-flight
`ix-*` units like `ix-reporting` finish first) *and* middlewared's docker/apps
startup has reached a terminal state — plain unit ordering cannot see either,
and restarting middlewared while they run kills apps and dashboard reporting
for the whole boot. S3/B2 backup support is then active until the next
reboot, when the cycle repeats.
The **re-apply** in that sentence is load-bearing, not a safety blanket. The
overlay from step 3 sits *inside* `/usr`, so anything that remounts that
hierarchy detaches it, and two ordinary things do exactly that after our hook
has finished: another PREINIT script running `systemd-sysext merge`/`refresh`
over `/usr` (an out-of-tree nvidia driver, say), and middlewared's own
`docker.configure_nvidia` when it brings docker up. Init scripts run
sequentially in id order, so a hook registered after ours always wins — and
ordering them differently would still not help, because `docker.configure_nvidia`
fires at runtime. `wait_restart.sh` therefore re-runs `apply.sh` at the point
where boot has settled and every such remount is behind it, re-mounting the
overlay if it was torn off (same upper layer, so files patched in step 3
reappear intact), then verifies the patch is really on the live path, restarts,
and verifies again — retrying once if it was lost in between.
This is the failure that made it necessary: on 2026-08-19 the overlay was
mounted at 16:41:56 and a sysext refresh unmerged and remerged `/usr` four
seconds later. The restart at 16:47:24 loaded stock modules, and every B2
backup failed for nineteen hours while `apply.log` said `OK` — because
`apply.log` can only report what was written to disk, never what the restart
imported. That second question is now asked from inside middlewared by an
hourly CRITICAL alert (see [Update alerts](../README.md#update-alerts)).
What you will observe: one middlewared restart shortly after every boot (a
brief web UI/API blip; running services are unaffected). Between steps 3
and 5 there is a short window — typically well under a minute — where the UI
already shows S3/B2 (the JS bundle is read from disk per request) but the
backend is still stock. A backup job that fires inside that window fails once
with `NotImplementedError` and succeeds on its next run; see
[Troubleshooting](recovery.md) if it persists beyond boot.
[Troubleshooting](recovery.md) if it persists beyond boot. If the backend is
still stock an hour after boot, middlewared raises the "installed but NOT
loaded" alert rather than leaving you to notice via a failed backup.
Manual runs of `bash patch/apply.sh` never trigger the restart — that only
happens in boot context. `install.sh` and `recover.sh` perform their own
+31 -4
View File
@@ -120,8 +120,9 @@ If a module shows `[FAIL]`:
The traceback ends in `rclone/base.py` → `raise NotImplementedError` and
contains no `_tc_` frames: the running middlewared is executing stock code.
Either the deferred restart never fired, or the patch never landed on disk
this boot. Diagnose in this order:
Either the deferred restart never fired, the patch never landed on disk this
boot, or it landed and was then torn off before the restart. Diagnose in this
order:
```bash
# Did apply.sh run this boot, at which version, and did it schedule the restart?
@@ -130,11 +131,21 @@ tail -40 /mnt/tank/truenas-truecloud-patch/apply.log
# Full check — compares the running process against the patch timestamp
python3 /mnt/tank/truenas-truecloud-patch/patch/create_task.py verify
# Did the deferred restart unit run, fail, or never get created?
systemctl status truecloud-mw-restart.service
# What the deferred restart did -- re-apply, restart, and what it verified.
# apply.log is the durable record; journald retention on a busy box is often
# shorter than the gap between reboots, so the journal may have nothing left.
grep wait_restart /mnt/tank/truenas-truecloud-patch/apply.log | tail -20
journalctl -u truecloud-mw-restart.service --no-pager | tail -20
# Did something remount /usr and detach the patch overlay?
systemd-sysext status
findmnt -o TARGET,SOURCE /usr/lib/python3/dist-packages
```
`systemctl status truecloud-mw-restart.service` reporting *"could not be
found"* is **normal** — the unit is transient and is collected once it exits.
It is not evidence that the restart was skipped.
- `verify` reports the process started **before** the patch → the restart
didn't happen. `systemctl restart middlewared` fixes it immediately; the
journal output above tells you why it was missed.
@@ -145,6 +156,22 @@ journalctl -u truecloud-mw-restart.service --no-pager | tail -20
- `apply.log` header shows `[v0.0.3]` or older → update:
`git pull && bash install.sh` (v0.0.4 fixed patches not loading after
reboot).
- `apply.log` says the patch applied, but `findmnt` shows no `truecloud-mw`
overlay on the dist-packages path → something remounted `/usr` after our
PREINIT hook and detached it. `systemd-sysext status` names the culprit if it
is a sysext (the `SINCE` column will sit a few seconds *after* the `apply.log`
timestamp). Releases from 2026-08-26 on re-apply and verify immediately before
the restart, so this should self-heal; if you are seeing it, update first.
**TrueNAS raises "truecloud-patch is installed but NOT loaded"**
The definitive symptom, and it does not depend on a backup failing first: the
running middlewared has stock cloud_backup modules even though the patch is
installed and its providers module is meant to be active. `bash install.sh`
re-applies and restarts. The alert clears within the hour. It is silent when the
kill switch is set or the providers module has been retired as native, and it is
deliberately not muted by `update_alerts_disabled` — that silences release
notifications, not a broken backup path.
**Apply log** (check after each reboot or install):
```bash
+1 -1
View File
@@ -25,7 +25,7 @@ are worth calling out, because nothing else would catch what they catch:
newest CHANGELOG entry. `VERSION=` had silently drifted to three different
values across the scripts before anything checked.
The project is hosted on **Gitea** (`git.onetick.ninja/flan/truenas-truecloud-patch`)
The project is hosted on **Gitea** (`git.arch.fyi/flan/truenas-truecloud-patch`)
and mirrored to GitHub. Both run the same workflows — Gitea reads
`.github/workflows/` too — so a change is checked twice, on two independent runners.
+1 -1
View File
@@ -18,7 +18,7 @@
set -euo pipefail
VERSION="0.7.0"
VERSION="0.8.0"
# The directory containing install.sh is the permanent install location.
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
+104 -2
View File
@@ -21,6 +21,7 @@ the way a `git fetch` from middlewared (running as root) would.
import datetime
import importlib.util
import json
import logging
import os
import re
@@ -47,8 +48,8 @@ _VERSION_RE = re.compile(r'^VERSION="([^"]+)"', re.M)
#: owner/repo out of any of:
#: git@github.com:sudolulo/repo.git
#: https://github.com/sudolulo/repo.git
#: ssh://git@git.onetick.ninja:55214/flan/repo.git
#: https://git.onetick.ninja/flan/repo.git
#: ssh://git@git.arch.fyi:55214/flan/repo.git
#: https://git.arch.fyi/flan/repo.git
#: The SSH port is deliberately not captured: it is not the web port.
_REMOTE_RE = re.compile(
r"^(?:\w+://)?(?:[^@/]+@)?([^:/]+)(?::\d+)?[:/]([^/]+)/([^/]+?)(?:\.git)?/?$"
@@ -77,6 +78,107 @@ class TrueCloudPatchSecurityUpdateAlertClass(AlertClass):
)
class TrueCloudPatchNotLoadedAlertClass(AlertClass):
category = AlertCategory.SYSTEM
level = AlertLevel.CRITICAL
title = "truecloud-patch is installed but NOT loaded"
text = (
"truecloud-patch patched middlewared on disk, but this middlewared is "
"running the STOCK cloud_backup modules -- B2 and S3 TrueCloud Backup "
"tasks will fail with NotImplementedError. Something remounted /usr "
"after the patch was applied (a systemd-sysext merge, or "
"docker.configure_nvidia), detaching the patch overlay. Re-apply with: "
"bash %(dir)s/install.sh"
)
class TrueCloudPatchNotLoadedAlertSource(ThreadedAlertSource):
"""Does the middlewared running this check actually have the patch in it?
This is the one question apply.log cannot answer. apply.sh reports what it
wrote to disk; whether the restart that followed imported those files is a
separate fact, and on 2026-08-19 the two disagreed silently for nineteen
hours while every B2 backup task failed. Asking from inside the process is
exact -- the patch stamps the objects it replaces, so a missing stamp means
this interpreter imported stock code.
Deliberately NOT silenced by the update-alert marker: that mutes release
notifications, not a broken backup path. Only the patch's own kill switch
(the `disabled` file, meaning the operator turned the patch off) stops it.
"""
schedule = IntervalSchedule(datetime.timedelta(hours=1))
run_on_backup_node = False
def check_sync(self):
try:
return self._check()
except Exception:
# An alert source must never take middlewared down with it.
logger.debug("truecloud-patch loaded check failed", exc_info=True)
return None
# -- internals ------------------------------------------------------------
def _check(self):
if os.path.exists(os.path.join(PATCH_DIR, "disabled")):
return None
# Only the providers module puts B2/S3 on the restic path. If it was
# never applied here, or TrueNAS went native and it was retired, then
# "not loaded" is the correct state and not a fault.
status = self._hook_status()
if not status:
return None
providers = status.get("patches", {}).get("providers", {})
if not providers.get("active"):
return None
if self._providers_loaded():
return None
return Alert(
TrueCloudPatchNotLoadedAlertClass,
{"dir": PATCH_DIR},
key=None,
)
def _hook_status(self):
try:
with open(os.path.join(PATCH_DIR, "hook_status.json")) as f:
return json.load(f)
except (OSError, ValueError):
return None
def _providers_loaded(self):
"""True when THIS interpreter holds the patched provider objects.
Two independent stamps, because the two halves are written separately
and either can be missing on its own:
* restic.py -- apply.sh sets `_truecloud_patched` on the wrapper it
installs over `get_restic_config`.
* b2.py -- apply.sh binds a B2-specific `get_restic_config` onto
`B2RcloneRemote`. Comparing it against the base implementation is
exact and survives renames of the patch's own helper.
"""
try:
from middlewared.plugins.cloud_backup.restic import get_restic_config
except Exception:
return False
if not getattr(get_restic_config, "_truecloud_patched", False):
return False
try:
from middlewared.rclone.base import BaseRcloneRemote
from middlewared.rclone.remote.b2 import B2RcloneRemote
except Exception:
return False
base = getattr(BaseRcloneRemote, "get_restic_config", None)
b2 = getattr(B2RcloneRemote, "get_restic_config", None)
return b2 is not None and b2 is not base
class TrueCloudPatchUpdateAlertSource(ThreadedAlertSource):
schedule = IntervalSchedule(datetime.timedelta(hours=24))
run_on_backup_node = False
+52 -7
View File
@@ -32,7 +32,7 @@
# Derive PATCH_DIR from this script's location (parent of the patch/ directory).
PATCH_DIR="$(cd "$(dirname "$0")/.." && pwd)"
LOG="$PATCH_DIR/apply.log"
VERSION="0.7.0"
VERSION="0.8.0"
# Rotate log at 512 KB to avoid unbounded growth on a system volume.
# Keep two prior generations (.1 and .2) so the last three boots are always available.
@@ -64,17 +64,45 @@ _ensure_writable() {
rm -f "$dir/.truecloud-probe"
return 0
fi
# Already our overlay on this exact directory from an earlier run this boot?
# Not writable, so any overlay of ours listed on this directory is a
# SHADOWED leftover rather than a working mount: something remounted the
# hierarchy above it -- a systemd-sysext merge/refresh over /usr, or
# middlewared's own docker.configure_nvidia -- and buried it. A live
# overlay of ours is always writable, so this must never be treated as
# "already done"; doing so is what let a buried overlay pass for a healthy
# one and left the backend patch on disk but never loaded.
if mount | grep -qF "truecloud-${tag} on ${dir} "; then
return 0
echo "NOTICE: a previous truecloud-${tag} overlay on $dir is shadowed --"
echo "NOTICE: the hierarchy above it was remounted. Detaching and re-mounting."
umount -l "$dir" 2>/dev/null
fi
# Keep the SAME upperdir across re-mounts: it holds everything patched
# earlier this boot, so re-mounting restores those files intact instead of
# re-deriving them. The workdir is scratch and must be empty, so it is
# recreated -- a stale one left behind by a detached mount fails the mount.
local upper="/run/truecloud-${tag}-upper" work="/run/truecloud-${tag}-work"
mkdir -p "$upper" "$work"
mkdir -p "$upper"
rm -rf "$work" 2>/dev/null
mkdir -p "$work"
if mount -t overlay "truecloud-${tag}" \
-o "lowerdir=$dir,upperdir=$upper,workdir=$work" "$dir" 2>/dev/null; then
echo "OK: Mounted writable overlay on $dir"
return 0
fi
# A lazily-detached overlay releases its workdir only once its last user is
# gone, and overlayfs refuses a workdir that is still in use. That would turn
# the re-mount this function exists to perform into a hard failure, so retry
# once on a private workdir. It is scratch in /run (tmpfs) and goes away at
# the next boot; the upperdir, which holds the patched files, is unchanged.
work="/run/truecloud-${tag}-work.$$"
rm -rf "$work" 2>/dev/null
mkdir -p "$work"
if mount -t overlay "truecloud-${tag}" \
-o "lowerdir=$dir,upperdir=$upper,workdir=$work" "$dir" 2>/dev/null; then
echo "OK: Mounted writable overlay on $dir (fresh workdir)"
return 0
fi
rmdir "$work" 2>/dev/null
echo "WARNING: overlay mount failed on $dir — backend patch will be skipped."
return 1
}
@@ -196,6 +224,13 @@ _tc_native_nested=$(printf '%s' "$_tc_info" | sed -n '2p')
SITE_PKG=$(printf '%s' "$_tc_info" | sed -n '3p')
_MW_DIR=$(printf '%s' "$_tc_info" | sed -n '4p')
# Record the resolved middlewared directory so patch/wait_restart.sh can check,
# without re-deriving any of this, whether the patched modules are still on the
# live filesystem path at the moment it restarts middlewared.
if [ -n "$_MW_DIR" ]; then
printf '%s\n' "$_MW_DIR" > "$PATCH_DIR/.mw_dir" 2>/dev/null
fi
# Nested support is opt-in; if it was never enabled, it cannot be the reason to
# keep the patch alive.
if [ -f "$PATCH_DIR/nested_snapshots_enabled" ]; then
@@ -999,8 +1034,13 @@ fi
# runs (install.sh, recovery) never trigger a restart.
#
# The unit runs wait_restart.sh, which blocks until boot has actually
# settled (systemd job queue drained, docker/apps state terminal) before
# restarting. systemd ordering alone (After=multi-user.target, ≤ v0.0.4)
# settled (systemd job queue drained, docker/apps state terminal), then
# RE-APPLIES this script before restarting. The re-apply is not belt-and-
# braces: our overlay lives inside /usr, and a systemd-sysext merge or
# middlewared's docker.configure_nvidia remounts /usr *after* PREINIT and
# detaches it, so what we patch here can be gone by restart time (seen
# 2026-08-19). wait_restart.sh re-mounts and re-verifies at the moment it
# matters. systemd ordering alone (After=multi-user.target, ≤ v0.0.4)
# fired while ix-reporting and the docker/apps startup were still in flight
# and killed both — apps and dashboard stats stayed down until the next
# boot. No Type=oneshot: a oneshot's start job would hold the boot queue
@@ -1015,7 +1055,12 @@ echo "--- deferred restart ---"
#
# "No module active at all" cannot reach here: that is the kill-switch branch
# above, which exits.
if ! grep -aq middlewared "/proc/$PPID/cmdline" 2>/dev/null; then
if [ "${TRUECLOUD_REAPPLY:-0}" = "1" ]; then
# Invoked by patch/wait_restart.sh as its pre-restart re-apply pass. That
# unit already exists to do the restart and verifies the result, so
# scheduling another one here would be a loop.
echo "Re-apply pass from wait_restart.sh — that unit owns the restart."
elif ! grep -aq middlewared "/proc/$PPID/cmdline" 2>/dev/null; then
echo "Manual run (parent is not middlewared) — no restart scheduled."
elif [ "$_backend_ok" != "1" ]; then
echo "Nothing landed on disk — no restart scheduled (nothing new to load)."
+1 -1
View File
@@ -52,7 +52,7 @@ import subprocess
import sys
import time
__version__ = "0.7.0"
__version__ = "0.8.0"
_PATCH_DIR = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
_STATUS_FILE = os.path.join(_PATCH_DIR, "hook_status.json")
+114 -1
View File
@@ -27,6 +27,50 @@
# --wait` below waits for that same queue to drain — the unit would deadlock
# on itself until the timeout. apply.sh schedules this with the default
# service type, whose start job completes at fork.
#
# THE RE-APPLY PASS (added 2026-08-26). Applying the patch at PREINIT and
# restarting later is only sound if the patched files are still on the live
# path at the moment middlewared re-imports them. They may not be: our patch
# lives in an overlay mounted *inside* /usr, and anything that remounts the
# hierarchy above it detaches or buries that overlay. Two things on a normal
# TrueNAS box do exactly that, both AFTER our PREINIT hook has run:
#
# - `systemd-sysext merge/refresh` over /usr (an nvidia sysext, for
# instance) — `Unmerged '/usr'` then `Merged extensions into '/usr'`;
# - middlewared's own `docker.configure_nvidia`, which merges the stock
# nvidia sysext over /usr when it brings docker up.
#
# PREINIT scripts run sequentially in id order, so a hook registered after
# ours always wins the race, silently. Observed 2026-08-19: our overlay was
# mounted at 16:41:56 and a sysext refresh tore /usr down four seconds later;
# the restart at 16:47:24 then loaded stock modules and every B2 cloud_backup
# job failed for the next nineteen hours while apply.log said "OK".
#
# Ordering the hooks cannot fix this — docker.configure_nvidia re-merges at
# runtime, long after every PREINIT hook is done. So instead of trusting the
# PREINIT pass, re-apply immediately before the restart (apply.sh is
# idempotent and re-mounts a lost overlay, keeping the same upperdir so
# already-patched files survive), verify the marker is really on the live
# path, and verify again afterwards.
PATCH_DIR="$(cd "$(dirname "$0")/.." && pwd)"
LOG="$PATCH_DIR/apply.log"
_log() { echo "[wait_restart] $*" >> "$LOG" 2>/dev/null; }
# Is the providers patch visible on the live filesystem path -- i.e. would a
# middlewared starting right now import it? Reads the marker apply.sh leaves
# in restic.py. Returns 0 when patched, 1 when stock, 2 when we cannot tell
# (no recorded middlewared dir yet, or the file is gone).
_patch_visible() {
local mw_dir restic_py
mw_dir=$(cat "$PATCH_DIR/.mw_dir" 2>/dev/null)
[ -n "$mw_dir" ] || return 2
restic_py="$mw_dir/plugins/cloud_backup/restic.py"
[ -f "$restic_py" ] || return 2
grep -q "TRUECLOUD_PATCH" "$restic_py" 2>/dev/null && return 0
return 1
}
# 1. systemd layer: wait for the boot job queue to drain. This covers every
# ix-* oneshot still activating, including ix-reporting's in-flight midclt
@@ -39,6 +83,9 @@ timeout 900 systemctl is-system-running --wait > /dev/null 2>&1
# transitional states (PENDING/INITIALIZING/STOPPING/MIGRATING — see
# middlewared/plugins/docker/state_utils.py). An empty answer means
# midclt could not respond at all; keep waiting. Cap at 10 minutes.
# This also covers docker.configure_nvidia, the runtime /usr re-merge:
# waiting for docker to reach a terminal state means the merge that would
# bury our overlay has already happened by the time we re-apply below.
for _ in $(seq 1 120); do
_status=$(midclt call docker.status 2>/dev/null \
| grep -oE '"status": "[A-Z_]+"' | cut -d'"' -f4)
@@ -52,4 +99,70 @@ done
# queryable state (smb.configure and friends). Bounded insurance.
sleep 30
exec systemctl try-restart middlewared
# 4. Re-apply pass. Boot has settled, so every sysext merge and docker nvidia
# configuration that could bury our overlay is behind us. Re-running
# apply.sh is cheap and idempotent: it re-mounts the overlay if it was
# detached (same upperdir, so files patched at PREINIT reappear intact)
# and re-patches anything that reverted to stock.
_patch_visible
case $? in
0) _log "providers patch still visible on the live path before restart" ;;
1) _log "PATCH LOST since PREINIT (something remounted /usr) — re-applying" ;;
*) _log "cannot confirm patch state before restart — re-applying anyway" ;;
esac
TRUECLOUD_REAPPLY=1 /bin/bash "$PATCH_DIR/patch/apply.sh"
if ! _patch_visible; then
_log "WARNING: patch is STILL not on the live path after the re-apply pass;"
_log "WARNING: restarting anyway, but middlewared will load stock modules."
fi
# 5. The restart itself.
systemctl try-restart middlewared
# 6. Record what the restart landed on -- but do NOT restart again on a miss.
#
# `try-restart` returns as soon as middlewared is READY; it then brings docker
# up asynchronously, and `docker.configure_nvidia` merges the stock nvidia
# sysext over /usr at that point. That detaches our overlay AFTER the new
# middlewared has already imported the patched modules -- so a disk check here
# can report "missing" on a perfectly healthy system. Restarting on that signal
# would restart a correctly-patched middlewared and then hit the same race
# again, so the disk is deliberately not treated as a verdict after the restart.
#
# The authoritative answer is whether the running process holds the patch, and
# only middlewared can answer that. The alert source installed by apply.sh
# checks exactly that, in-process and hourly, and is what reports a genuine
# miss. What is still worth doing here is putting the overlay back, so the next
# middlewared restart -- whenever and whyever it happens -- finds patched files.
if _patch_visible; then
_log "OK: providers patch present on the live path across the restart"
else
_log "overlay detached again after the restart (expected when docker's"
_log "nvidia sysext merge follows it) -- re-mounting for the next restart."
_log "Whether THIS middlewared loaded the patch is answered in-process by"
_log "the 'installed but NOT loaded' alert, not by this check."
# Preserve hook_status.json's patched_at across this re-mount.
#
# create_task.py verify decides "loaded" by comparing middlewared's start
# time against patched_at. This re-apply restores the SAME patch the boot
# pass already applied, but it runs *after* the restart -- so letting it
# re-stamp would make patched_at newer than the process that correctly
# imported the patch, and verify would report FAIL forever, on every boot
# where docker's sysext merge detaches the overlay. That is precisely the
# lying-status failure this release exists to remove, so do not introduce a
# new one. The snapshot lives in /run, never in the repo: a leftover file
# there would leave the tree dirty and update.sh refuses to run over that.
_saved_status=/run/truecloud-hook_status.pre
cp -p "$PATCH_DIR/hook_status.json" "$_saved_status" 2>/dev/null
TRUECLOUD_REAPPLY=1 /bin/bash "$PATCH_DIR/patch/apply.sh"
if [ -f "$_saved_status" ]; then
mv -f "$_saved_status" "$PATCH_DIR/hook_status.json" 2>/dev/null
fi
fi
_log "=== deferred restart complete ==="
+1 -1
View File
@@ -17,7 +17,7 @@
# bash /mnt/tank/truenas-truecloud-patch/patch/apply.sh
# systemctl restart middlewared
VERSION="0.7.0"
VERSION="0.8.0"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
+210
View File
@@ -0,0 +1,210 @@
"""Behavioural tests for the "installed but NOT loaded" alert.
apply.log can only report what was written to disk. Whether the middlewared that
restarted afterwards actually imported those files is a different fact, and when
the two disagree nothing else notices: on 2026-08-19 every B2 backup failed for
nineteen hours while the log said OK. This alert is the only thing that closes
that gap, so it is tested against real objects rather than by reading source.
The middlewared package does not exist off-box, so the modules the alert source
imports are stubbed here.
"""
import importlib.util
import json
import os
import sys
import types
import pytest
ALERT_SRC = os.path.join(os.path.dirname(__file__), "..", "patch", "alert_source.py")
class _StubAlertClass:
pass
class _StubThreadedAlertSource:
pass
class _StubAlert:
def __init__(self, klass, args=None, key=None):
self.klass = klass
self.args = args
self.key = key
def _module(name):
mod = types.ModuleType(name)
sys.modules[name] = mod
return mod
@pytest.fixture
def alert_source(monkeypatch, tmp_path):
"""Load patch/alert_source.py against stubbed middlewared modules."""
for name in list(sys.modules):
if name == "middlewared" or name.startswith("middlewared."):
monkeypatch.delitem(sys.modules, name, raising=False)
_module("middlewared")
_module("middlewared.alert")
base = _module("middlewared.alert.base")
base.Alert = _StubAlert
base.AlertClass = _StubAlertClass
base.ThreadedAlertSource = _StubThreadedAlertSource
base.AlertCategory = types.SimpleNamespace(SYSTEM="SYSTEM")
base.AlertLevel = types.SimpleNamespace(
INFO="INFO", WARNING="WARNING", CRITICAL="CRITICAL"
)
schedule = _module("middlewared.alert.schedule")
schedule.IntervalSchedule = lambda delta: ("interval", delta)
spec = importlib.util.spec_from_file_location("_tc_alert_source", ALERT_SRC)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
mod.PATCH_DIR = str(tmp_path)
return mod
def _write_status(tmp_path, providers_active=True):
payload = {
"patched_at": "2026-08-26T00:00:00Z",
"patches": {
"providers": {"ok": True, "active": providers_active, "detail": "x"},
"nested_snapshots": {"ok": True, "active": True, "detail": "x"},
},
}
(tmp_path / "hook_status.json").write_text(json.dumps(payload))
def _install_provider_modules(monkeypatch, *, restic_patched, b2_patched):
"""Stub the two modules the alert inspects, in the requested state."""
plugins = _module("middlewared.plugins")
_module("middlewared.plugins.cloud_backup")
restic = _module("middlewared.plugins.cloud_backup.restic")
def get_restic_config(task):
return None
if restic_patched:
get_restic_config._truecloud_patched = True
restic.get_restic_config = get_restic_config
rclone_base = _module("middlewared.rclone.base")
_module("middlewared.rclone")
_module("middlewared.rclone.remote")
b2_mod = _module("middlewared.rclone.remote.b2")
class BaseRcloneRemote:
def get_restic_config(self, task):
raise NotImplementedError
class B2RcloneRemote(BaseRcloneRemote):
pass
if b2_patched:
B2RcloneRemote.get_restic_config = staticmethod(lambda task: ("url", {}))
rclone_base.BaseRcloneRemote = BaseRcloneRemote
b2_mod.B2RcloneRemote = B2RcloneRemote
b2_mod.BaseRcloneRemote = BaseRcloneRemote
plugins.__path__ = []
for name in (
"middlewared.plugins",
"middlewared.plugins.cloud_backup",
"middlewared.plugins.cloud_backup.restic",
"middlewared.rclone",
"middlewared.rclone.base",
"middlewared.rclone.remote",
"middlewared.rclone.remote.b2",
):
monkeypatch.setitem(sys.modules, name, sys.modules[name])
def _source(alert_source):
cls = alert_source.TrueCloudPatchNotLoadedAlertSource
return cls.__new__(cls)
def test_no_alert_when_patch_is_loaded(alert_source, monkeypatch, tmp_path):
_write_status(tmp_path)
_install_provider_modules(monkeypatch, restic_patched=True, b2_patched=True)
assert _source(alert_source)._check() is None
def test_alert_when_middlewared_loaded_stock_modules(alert_source, monkeypatch, tmp_path):
"""The exact 2026-08-19 state: patched on disk, stock in the process."""
_write_status(tmp_path)
_install_provider_modules(monkeypatch, restic_patched=False, b2_patched=False)
alert = _source(alert_source)._check()
assert alert is not None
assert alert.klass is alert_source.TrueCloudPatchNotLoadedAlertClass
def test_alert_when_only_b2_half_is_missing(alert_source, monkeypatch, tmp_path):
# b2.py is the half that supplies B2's get_restic_config. restic.py alone
# being patched still means every B2 task raises NotImplementedError.
_write_status(tmp_path)
_install_provider_modules(monkeypatch, restic_patched=True, b2_patched=False)
assert _source(alert_source)._check() is not None
def test_alert_when_only_restic_half_is_missing(alert_source, monkeypatch, tmp_path):
_write_status(tmp_path)
_install_provider_modules(monkeypatch, restic_patched=False, b2_patched=True)
assert _source(alert_source)._check() is not None
def test_silent_when_the_kill_switch_is_set(alert_source, monkeypatch, tmp_path):
# The operator turned the patch off on purpose; stock is the intended state.
_write_status(tmp_path)
(tmp_path / "disabled").write_text("")
_install_provider_modules(monkeypatch, restic_patched=False, b2_patched=False)
assert _source(alert_source)._check() is None
def test_silent_when_providers_module_is_retired(alert_source, monkeypatch, tmp_path):
# TrueNAS went native for B2: not loading our providers patch is correct.
_write_status(tmp_path, providers_active=False)
_install_provider_modules(monkeypatch, restic_patched=False, b2_patched=False)
assert _source(alert_source)._check() is None
def test_silent_when_the_patch_was_never_applied_here(alert_source, monkeypatch, tmp_path):
# No hook_status.json at all -- nothing claims a patch, so nothing is broken.
_install_provider_modules(monkeypatch, restic_patched=False, b2_patched=False)
assert _source(alert_source)._check() is None
def test_update_alert_silencer_does_not_mute_a_broken_backup_path(
alert_source, monkeypatch, tmp_path
):
# update_alerts_disabled mutes release notifications. It must not hide the
# fact that TrueCloud backups are silently running stock.
_write_status(tmp_path)
(tmp_path / "update_alerts_disabled").write_text("")
_install_provider_modules(monkeypatch, restic_patched=False, b2_patched=False)
assert _source(alert_source)._check() is not None
def test_check_sync_never_raises(alert_source, monkeypatch, tmp_path):
"""An alert source that raises is polled forever inside middlewared."""
_write_status(tmp_path)
def boom(self):
raise RuntimeError("provider import exploded")
monkeypatch.setattr(
alert_source.TrueCloudPatchNotLoadedAlertSource, "_check", boom, raising=True
)
assert _source(alert_source).check_sync() is None
def test_alert_is_critical_and_names_the_recovery_command(alert_source):
klass = alert_source.TrueCloudPatchNotLoadedAlertClass
assert klass.level == "CRITICAL"
assert "install.sh" in klass.text
+6
View File
@@ -1914,6 +1914,9 @@ class TestTheSnapshotRecordIsNeverLostToAnUnreadableFile:
only record of a tree it had just failed to read.
"""
@pytest.mark.skipif(os.geteuid() == 0, reason=
"chmod(0) cannot make a file unreadable for root (CAP_DAC_OVERRIDE); "
"the Gitea runner image executes jobs as root")
def test_an_unreadable_sidecar_raises_rather_than_reading_as_empty(self, tmp_path):
sc = tmp_path / "cloud_backup-5.snapshot"
sc.write_text("Tap@snap\n")
@@ -1927,6 +1930,9 @@ class TestTheSnapshotRecordIsNeverLostToAnUnreadableFile:
def test_a_missing_sidecar_is_simply_empty(self, tmp_path):
assert tn._read_sidecar(str(tmp_path / "nope")) == []
@pytest.mark.skipif(os.geteuid() == 0, reason=
"chmod(0) cannot make a file unreadable for root (CAP_DAC_OVERRIDE); "
"the Gitea runner image executes jobs as root")
def test_cleanup_all_still_UNMOUNTS_when_a_sidecar_cannot_be_read(self, tmp_path):
# cleanup_all is what recover.sh and uninstall.sh call, i.e. it runs precisely
# when the box is already stuck. Its job is to get the mounts off. Aborting on
+199
View File
@@ -0,0 +1,199 @@
"""The deferred restart must re-apply the patch before it restarts middlewared.
Patching at PREINIT and restarting minutes later is only sound while the patched
files are still on the live path when middlewared re-imports them. They may not
be: the patch lives in an overlay mounted inside /usr, and anything that
remounts that hierarchy detaches it. On 2026-08-19 a systemd-sysext refresh over
/usr ran four seconds after apply.sh mounted its overlay; the deferred restart
then loaded stock modules and every B2 cloud_backup job failed for nineteen
hours while apply.log reported "OK".
These tests pin the ordering that makes that non-recoverable failure impossible:
re-apply, verify, restart, verify again.
"""
import os
import re
import subprocess
import pytest
HERE = os.path.dirname(__file__)
WAIT_RESTART = os.path.join(HERE, "..", "patch", "wait_restart.sh")
APPLY_SH = os.path.join(HERE, "..", "patch", "apply.sh")
def wait_restart_source():
with open(WAIT_RESTART, encoding="utf-8") as fh:
return fh.read()
def apply_source():
with open(APPLY_SH, encoding="utf-8") as fh:
return fh.read()
def test_wait_restart_is_executable():
# apply.sh schedules it as `/bin/bash <script>`, but install.sh ships exec
# bits and a mode-only diff once blocked update.sh outright (v0.6.0).
assert os.access(WAIT_RESTART, os.X_OK)
def test_wait_restart_is_syntactically_valid():
subprocess.run(["bash", "-n", WAIT_RESTART], check=True)
def test_reapply_runs_before_the_restart():
src = wait_restart_source()
reapply = src.index("TRUECLOUD_REAPPLY=1")
restart = src.index("systemctl try-restart middlewared")
assert reapply < restart, "the re-apply pass must precede the restart"
def test_restart_is_not_exec_so_verification_can_follow():
# Up to v0.7.0 the script ended in `exec systemctl try-restart middlewared`,
# which replaces the shell -- nothing could run afterwards. The post-restart
# verification only exists if the restart is a plain call.
src = wait_restart_source()
assert not re.search(r"^\s*exec\s+systemctl", src, re.M)
def test_middlewared_is_restarted_exactly_once():
"""No restart loop.
`try-restart` returns at READY; middlewared then brings docker up, and
docker.configure_nvidia merges the nvidia sysext over /usr right about then
-- detaching the overlay AFTER the patched modules are already imported. A
disk check after the restart therefore false-negatives on a healthy system,
and restarting on that signal would restart a correctly-patched middlewared
straight back into the same race.
"""
src = wait_restart_source()
assert src.count("systemctl try-restart middlewared") == 1
def test_patch_is_verified_after_the_restart():
src = wait_restart_source()
restart = src.index("systemctl try-restart middlewared")
assert "_patch_visible" in src[restart:], (
"the script must check what the restart actually loaded"
)
def test_verification_reads_the_marker_apply_sh_writes():
# _patch_visible greps restic.py for TRUECLOUD_PATCH; apply.sh must still be
# the thing that puts it there, or the check silently always fails.
assert "TRUECLOUD_PATCH" in wait_restart_source()
assert "TRUECLOUD_PATCH" in apply_source()
def test_verification_uses_the_recorded_middlewared_dir():
# wait_restart.sh must not re-derive site-packages; apply.sh records it.
assert ".mw_dir" in wait_restart_source()
assert ".mw_dir" in apply_source()
def test_apply_sh_records_the_middlewared_dir():
src = apply_source()
assert re.search(r'>\s*"\$PATCH_DIR/\.mw_dir"', src), (
"apply.sh must write the resolved middlewared dir for wait_restart.sh"
)
def test_reapply_pass_does_not_schedule_another_restart():
# wait_restart.sh owns the restart. If the re-apply pass scheduled its own
# transient unit, each boot would spawn restarts recursively.
src = apply_source()
guard = src.index('if [ "${TRUECLOUD_REAPPLY:-0}" = "1" ]; then')
systemd_run = src.index("systemd-run --no-block")
assert guard < systemd_run, (
"the TRUECLOUD_REAPPLY branch must short-circuit before systemd-run"
)
def test_shadowed_overlay_is_remounted_not_accepted():
"""A buried overlay must never pass for a healthy one.
_ensure_writable reaches its mount-table check only when the directory is
NOT writable -- and a live overlay of ours is always writable. So a
truecloud mount listed at that point is shadowed, and returning 0 there is
exactly how a detached overlay used to masquerade as applied.
"""
src = apply_source()
start = src.index("_ensure_writable()")
end = src.index("\n}", start)
body = src[start:end]
check = body.index('mount | grep -qF "truecloud-${tag} on ${dir} "')
following = body[check:]
# The old code did `return 0` immediately inside this branch.
branch_end = following.index("fi")
assert "return 0" not in following[:branch_end]
assert "umount -l" in following[:branch_end]
def test_workdir_is_recreated_before_mounting():
# overlayfs refuses a workdir left behind by a detached mount, so a stale
# one would turn every re-mount attempt into "overlay mount failed".
src = apply_source()
start = src.index("_ensure_writable()")
end = src.index("\n}", start)
body = src[start:end]
assert re.search(r'rm -rf "\$work"', body)
def test_upperdir_is_preserved_across_remounts():
# The upperdir holds everything patched earlier this boot; reusing it is
# what lets a re-mount restore those files instead of re-deriving them.
src = apply_source()
start = src.index("_ensure_writable()")
end = src.index("\n}", start)
body = src[start:end]
assert 'rm -rf "$upper"' not in body
@pytest.mark.parametrize("state", ["0", "1", "2"])
def test_patch_visible_returns_three_distinct_states(state):
# patched / stock / cannot-tell must stay distinguishable: "cannot tell"
# has to re-apply rather than assume the patch is fine.
src = wait_restart_source()
assert f"return {state}" in src or f") return {state}" in src
def test_mount_retries_on_a_private_workdir():
"""A lazily-detached overlay can still pin the shared workdir.
overlayfs refuses a workdir that is in use, so without a retry the re-mount
this whole fix depends on would fail exactly when it is most needed.
"""
src = apply_source()
start = src.index("_ensure_writable()")
end = src.index("\n}", start)
body = src[start:end]
assert body.count("mount -t overlay") == 2, "expected a retry mount"
assert 'work="/run/truecloud-${tag}-work.$$"' in body
def test_post_restart_remount_preserves_the_patched_at_stamp():
"""create_task.py verify compares middlewared's start time to patched_at.
The post-restart re-mount restores the same patch the boot pass applied, so
letting apply.sh re-stamp would make patched_at newer than the process that
correctly imported it -- verify would then report FAIL forever on every boot
where docker's sysext merge detaches the overlay.
"""
src = wait_restart_source()
tail = src[src.index("systemctl try-restart middlewared"):]
assert "hook_status.json" in tail
save = tail.index("/run/truecloud-hook_status.pre")
reapply = tail.index("TRUECLOUD_REAPPLY=1")
restore = tail.rindex("hook_status.json")
assert save < reapply < restore, "snapshot must bracket the re-apply"
def test_status_snapshot_is_not_written_into_the_repo():
"""A leftover file in the repo dir leaves the tree dirty, and update.sh
refuses to run over a dirty tree -- that once made the patch un-updatable."""
src = wait_restart_source()
assert "/run/truecloud-hook_status.pre" in src
assert '"$PATCH_DIR/hook_status.json.pre' not in src
+78
View File
@@ -172,3 +172,81 @@ class TestCompatCannotSilentlyPass:
with open(os.path.join(WORKFLOWS, "compat.yml"), encoding="utf-8") as fh:
src = fh.read()
assert "steps.check.outputs.shipped_broken != '0'" in src
class TestActionCacheRace:
"""CI must not run concurrent jobs on the self-hosted runner.
`act` caches each ACTION as one shared clone under /root/.cache/act/<hash>
and re-pulls it per job, so jobs starting together fight over that directory
and the loser dies with `lstat .../<file>: no such file or directory` before
any test runs -- a red `main` with zero suite output and a different victim
each push. The runner also force-pulls its base image per job, so job count
is also Docker Hub pull count, and four-per-push exhausted the anonymous
limit in an afternoon. Both problems have the same cure: one job.
"""
def _ci(self):
with open(os.path.join(WORKFLOWS, "ci.yml"), encoding="utf-8") as fh:
return fh.read()
def test_ci_runs_as_exactly_one_job(self):
"""The fix is the absence of concurrency, not the absence of one action.
Dropping astral-sh/setup-uv only shrank the surface -- every job still
used actions/checkout. A single job cannot race itself whatever actions
it uses, which is why this, and not the action count, is the invariant.
"""
ci = self._ci()
# Scope to the jobs: block -- `on:` has two-space keys of its own
# (push/pull_request/workflow_dispatch) that look identical otherwise.
body = ci[ci.index("\njobs:"):]
jobs = re.findall(r"^ (\w[\w-]*):$", body, re.M)
assert len(jobs) == 1, (
f"ci.yml defines {len(jobs)} jobs ({jobs}); concurrent jobs on the "
"self-hosted runner race on act's shared action cache and multiply "
"Docker Hub pulls. Keep CI to one job."
)
def test_no_matrix_reintroduces_parallel_jobs(self):
ci = self._ci()
assert "strategy:" not in ci and "matrix:" not in ci, (
"a matrix fans out into concurrent jobs again -- sweep versions "
"inside one job instead"
)
def test_every_python_version_still_runs_after_one_fails(self):
"""`fail-fast: false` is what the loop has to preserve.
A 3.11 break must not hide whether 3.12 and 3.13 are fine; that is
precisely the information you want at that moment.
"""
ci = self._ci()
assert 'PYTHONS: "3.11 3.12 3.13"' in ci
assert ci.count("fail=1") >= 2, "the sweeps must collect failures, not exit early"
def test_uv_is_installed_without_an_action(self):
"""Checks `uses:` directives, not prose.
The comment in ci.yml names the action it deliberately avoids, and that
explanation is the most useful thing in the file -- a test that greps the
raw text would forbid documenting the very lesson it enforces. Parsed
with a regex rather than PyYAML on purpose: CI runs `uvx pytest`, whose
environment holds pytest and nothing else, so a third-party import here
fails on the runner while passing locally.
"""
ci = self._ci()
used = re.findall(r"^\s*-?\s*uses:\s*(\S+)", ci, re.M)
assert not [u for u in used if "setup-uv" in u], (
"the action was only fetching a binary; a run: step does the same "
"with one less moving part"
)
assert "astral.sh/uv/" in ci
def test_the_uv_version_is_pinned(self):
ci = self._ci()
assert re.search(r'UV_VERSION:\s*"\d+\.\d+\.\d+"', ci), (
"an unpinned uv lets any upstream release turn main red with no "
"code change here -- the same rule ruff is pinned under"
)
assert "https://astral.sh/uv/${UV_VERSION}/install.sh" in ci
+6 -1
View File
@@ -3,7 +3,7 @@
set -euo pipefail
VERSION="0.7.0"
VERSION="0.8.0"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
_HOOK_COMMENT='TrueCloud provider patch (S3/B2)'
@@ -124,6 +124,11 @@ if [ -f "$PATCH_DIR/nested_snapshots_enabled" ]; then
rm -f "$PATCH_DIR/nested_snapshots_enabled"
echo " Removed nested-snapshot opt-in marker."
fi
# Runtime breadcrumb recorded by apply.sh for wait_restart.sh. Harmless, but a
# stale path left in an uninstalled tree is exactly the sort of thing that reads
# as state later.
rm -f "$PATCH_DIR/.mw_dir"
echo ""
if [ "$_restore_failed" -eq 1 ]; then
+1 -1
View File
@@ -19,7 +19,7 @@
set -euo pipefail
VERSION="0.7.0"
VERSION="0.8.0"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
_PREV_FILE="$PATCH_DIR/.update_previous"