Re-apply and verify the patch before the deferred restart
Patching at PREINIT and restarting minutes later is only sound while the patched files are still on the live path when middlewared re-imports them, and PREINIT cannot guarantee that. The overlay sits inside /usr, so anything that remounts that hierarchy detaches it — a systemd-sysext merge/refresh from another PREINIT hook, or middlewared's own docker.configure_nvidia at runtime. Init scripts run sequentially in id order, so a hook registered after this one always wins, and reordering them would not help because docker.configure_nvidia fires long after PREINIT is done. Observed on 25.10.6: the overlay was mounted at 16:41:56, a sysext refresh unmerged and remerged /usr four seconds later, and the deferred restart at 16:47:24 loaded stock modules. Every B2 cloud_backup task then failed with NotImplementedError for nineteen hours across four scheduled runs while apply.log and hook_status.json both reported the patch active. wait_restart.sh now re-applies immediately before restarting — after boot has settled, which is also after every sysext merge and docker nvidia configuration — verifies the marker is on the live path, restarts, and verifies again, retrying once. It is no longer exec'd, so something can run after the restart to find out what it loaded. apply.sh records the resolved middlewared directory in .mw_dir for that check, and honours TRUECLOUD_REAPPLY so the re-apply pass does not schedule a second restart. _ensure_writable treated "one of our overlays is listed here" as "already done", but it only reaches that check when the directory is not writable, and a live overlay of ours always is — a shadowed overlay was indistinguishable from a healthy one. It is now detached and re-mounted, reusing the upperdir so files patched earlier in the boot survive, with a fresh workdir and a retry on a private one, since overlayfs refuses a workdir a detached mount still holds. Add a CRITICAL hourly alert for the case none of this can prevent: the patch being on disk but not in the running process. apply.log can only report the first. The alert asks the second question from inside middlewared, where the patch's own stamps make it exact, and checks both halves since either can go missing alone. It stays quiet when the kill switch is set or the providers module has been retired as native, and is not muted by update_alerts_disabled. wait_restart.sh also logs to apply.log now: journald retention on a busy box is easily shorter than the interval between reboots, and the boot that caused this had already rotated away by the time it was investigated.
This commit is contained in:
@@ -21,6 +21,7 @@ the way a `git fetch` from middlewared (running as root) would.
|
||||
|
||||
import datetime
|
||||
import importlib.util
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
@@ -77,6 +78,107 @@ class TrueCloudPatchSecurityUpdateAlertClass(AlertClass):
|
||||
)
|
||||
|
||||
|
||||
class TrueCloudPatchNotLoadedAlertClass(AlertClass):
|
||||
category = AlertCategory.SYSTEM
|
||||
level = AlertLevel.CRITICAL
|
||||
title = "truecloud-patch is installed but NOT loaded"
|
||||
text = (
|
||||
"truecloud-patch patched middlewared on disk, but this middlewared is "
|
||||
"running the STOCK cloud_backup modules -- B2 and S3 TrueCloud Backup "
|
||||
"tasks will fail with NotImplementedError. Something remounted /usr "
|
||||
"after the patch was applied (a systemd-sysext merge, or "
|
||||
"docker.configure_nvidia), detaching the patch overlay. Re-apply with: "
|
||||
"bash %(dir)s/install.sh"
|
||||
)
|
||||
|
||||
|
||||
class TrueCloudPatchNotLoadedAlertSource(ThreadedAlertSource):
|
||||
"""Does the middlewared running this check actually have the patch in it?
|
||||
|
||||
This is the one question apply.log cannot answer. apply.sh reports what it
|
||||
wrote to disk; whether the restart that followed imported those files is a
|
||||
separate fact, and on 2026-08-19 the two disagreed silently for nineteen
|
||||
hours while every B2 backup task failed. Asking from inside the process is
|
||||
exact -- the patch stamps the objects it replaces, so a missing stamp means
|
||||
this interpreter imported stock code.
|
||||
|
||||
Deliberately NOT silenced by the update-alert marker: that mutes release
|
||||
notifications, not a broken backup path. Only the patch's own kill switch
|
||||
(the `disabled` file, meaning the operator turned the patch off) stops it.
|
||||
"""
|
||||
|
||||
schedule = IntervalSchedule(datetime.timedelta(hours=1))
|
||||
run_on_backup_node = False
|
||||
|
||||
def check_sync(self):
|
||||
try:
|
||||
return self._check()
|
||||
except Exception:
|
||||
# An alert source must never take middlewared down with it.
|
||||
logger.debug("truecloud-patch loaded check failed", exc_info=True)
|
||||
return None
|
||||
|
||||
# -- internals ------------------------------------------------------------
|
||||
|
||||
def _check(self):
|
||||
if os.path.exists(os.path.join(PATCH_DIR, "disabled")):
|
||||
return None
|
||||
|
||||
# Only the providers module puts B2/S3 on the restic path. If it was
|
||||
# never applied here, or TrueNAS went native and it was retired, then
|
||||
# "not loaded" is the correct state and not a fault.
|
||||
status = self._hook_status()
|
||||
if not status:
|
||||
return None
|
||||
providers = status.get("patches", {}).get("providers", {})
|
||||
if not providers.get("active"):
|
||||
return None
|
||||
|
||||
if self._providers_loaded():
|
||||
return None
|
||||
|
||||
return Alert(
|
||||
TrueCloudPatchNotLoadedAlertClass,
|
||||
{"dir": PATCH_DIR},
|
||||
key=None,
|
||||
)
|
||||
|
||||
def _hook_status(self):
|
||||
try:
|
||||
with open(os.path.join(PATCH_DIR, "hook_status.json")) as f:
|
||||
return json.load(f)
|
||||
except (OSError, ValueError):
|
||||
return None
|
||||
|
||||
def _providers_loaded(self):
|
||||
"""True when THIS interpreter holds the patched provider objects.
|
||||
|
||||
Two independent stamps, because the two halves are written separately
|
||||
and either can be missing on its own:
|
||||
|
||||
* restic.py -- apply.sh sets `_truecloud_patched` on the wrapper it
|
||||
installs over `get_restic_config`.
|
||||
* b2.py -- apply.sh binds a B2-specific `get_restic_config` onto
|
||||
`B2RcloneRemote`. Comparing it against the base implementation is
|
||||
exact and survives renames of the patch's own helper.
|
||||
"""
|
||||
try:
|
||||
from middlewared.plugins.cloud_backup.restic import get_restic_config
|
||||
except Exception:
|
||||
return False
|
||||
if not getattr(get_restic_config, "_truecloud_patched", False):
|
||||
return False
|
||||
|
||||
try:
|
||||
from middlewared.rclone.base import BaseRcloneRemote
|
||||
from middlewared.rclone.remote.b2 import B2RcloneRemote
|
||||
except Exception:
|
||||
return False
|
||||
base = getattr(BaseRcloneRemote, "get_restic_config", None)
|
||||
b2 = getattr(B2RcloneRemote, "get_restic_config", None)
|
||||
return b2 is not None and b2 is not base
|
||||
|
||||
|
||||
class TrueCloudPatchUpdateAlertSource(ThreadedAlertSource):
|
||||
schedule = IntervalSchedule(datetime.timedelta(hours=24))
|
||||
run_on_backup_node = False
|
||||
|
||||
+51
-6
@@ -64,17 +64,45 @@ _ensure_writable() {
|
||||
rm -f "$dir/.truecloud-probe"
|
||||
return 0
|
||||
fi
|
||||
# Already our overlay on this exact directory from an earlier run this boot?
|
||||
# Not writable, so any overlay of ours listed on this directory is a
|
||||
# SHADOWED leftover rather than a working mount: something remounted the
|
||||
# hierarchy above it -- a systemd-sysext merge/refresh over /usr, or
|
||||
# middlewared's own docker.configure_nvidia -- and buried it. A live
|
||||
# overlay of ours is always writable, so this must never be treated as
|
||||
# "already done"; doing so is what let a buried overlay pass for a healthy
|
||||
# one and left the backend patch on disk but never loaded.
|
||||
if mount | grep -qF "truecloud-${tag} on ${dir} "; then
|
||||
return 0
|
||||
echo "NOTICE: a previous truecloud-${tag} overlay on $dir is shadowed --"
|
||||
echo "NOTICE: the hierarchy above it was remounted. Detaching and re-mounting."
|
||||
umount -l "$dir" 2>/dev/null
|
||||
fi
|
||||
# Keep the SAME upperdir across re-mounts: it holds everything patched
|
||||
# earlier this boot, so re-mounting restores those files intact instead of
|
||||
# re-deriving them. The workdir is scratch and must be empty, so it is
|
||||
# recreated -- a stale one left behind by a detached mount fails the mount.
|
||||
local upper="/run/truecloud-${tag}-upper" work="/run/truecloud-${tag}-work"
|
||||
mkdir -p "$upper" "$work"
|
||||
mkdir -p "$upper"
|
||||
rm -rf "$work" 2>/dev/null
|
||||
mkdir -p "$work"
|
||||
if mount -t overlay "truecloud-${tag}" \
|
||||
-o "lowerdir=$dir,upperdir=$upper,workdir=$work" "$dir" 2>/dev/null; then
|
||||
echo "OK: Mounted writable overlay on $dir"
|
||||
return 0
|
||||
fi
|
||||
# A lazily-detached overlay releases its workdir only once its last user is
|
||||
# gone, and overlayfs refuses a workdir that is still in use. That would turn
|
||||
# the re-mount this function exists to perform into a hard failure, so retry
|
||||
# once on a private workdir. It is scratch in /run (tmpfs) and goes away at
|
||||
# the next boot; the upperdir, which holds the patched files, is unchanged.
|
||||
work="/run/truecloud-${tag}-work.$$"
|
||||
rm -rf "$work" 2>/dev/null
|
||||
mkdir -p "$work"
|
||||
if mount -t overlay "truecloud-${tag}" \
|
||||
-o "lowerdir=$dir,upperdir=$upper,workdir=$work" "$dir" 2>/dev/null; then
|
||||
echo "OK: Mounted writable overlay on $dir (fresh workdir)"
|
||||
return 0
|
||||
fi
|
||||
rmdir "$work" 2>/dev/null
|
||||
echo "WARNING: overlay mount failed on $dir — backend patch will be skipped."
|
||||
return 1
|
||||
}
|
||||
@@ -196,6 +224,13 @@ _tc_native_nested=$(printf '%s' "$_tc_info" | sed -n '2p')
|
||||
SITE_PKG=$(printf '%s' "$_tc_info" | sed -n '3p')
|
||||
_MW_DIR=$(printf '%s' "$_tc_info" | sed -n '4p')
|
||||
|
||||
# Record the resolved middlewared directory so patch/wait_restart.sh can check,
|
||||
# without re-deriving any of this, whether the patched modules are still on the
|
||||
# live filesystem path at the moment it restarts middlewared.
|
||||
if [ -n "$_MW_DIR" ]; then
|
||||
printf '%s\n' "$_MW_DIR" > "$PATCH_DIR/.mw_dir" 2>/dev/null
|
||||
fi
|
||||
|
||||
# Nested support is opt-in; if it was never enabled, it cannot be the reason to
|
||||
# keep the patch alive.
|
||||
if [ -f "$PATCH_DIR/nested_snapshots_enabled" ]; then
|
||||
@@ -999,8 +1034,13 @@ fi
|
||||
# runs (install.sh, recovery) never trigger a restart.
|
||||
#
|
||||
# The unit runs wait_restart.sh, which blocks until boot has actually
|
||||
# settled (systemd job queue drained, docker/apps state terminal) before
|
||||
# restarting. systemd ordering alone (After=multi-user.target, ≤ v0.0.4)
|
||||
# settled (systemd job queue drained, docker/apps state terminal), then
|
||||
# RE-APPLIES this script before restarting. The re-apply is not belt-and-
|
||||
# braces: our overlay lives inside /usr, and a systemd-sysext merge or
|
||||
# middlewared's docker.configure_nvidia remounts /usr *after* PREINIT and
|
||||
# detaches it, so what we patch here can be gone by restart time (seen
|
||||
# 2026-08-19). wait_restart.sh re-mounts and re-verifies at the moment it
|
||||
# matters. systemd ordering alone (After=multi-user.target, ≤ v0.0.4)
|
||||
# fired while ix-reporting and the docker/apps startup were still in flight
|
||||
# and killed both — apps and dashboard stats stayed down until the next
|
||||
# boot. No Type=oneshot: a oneshot's start job would hold the boot queue
|
||||
@@ -1015,7 +1055,12 @@ echo "--- deferred restart ---"
|
||||
#
|
||||
# "No module active at all" cannot reach here: that is the kill-switch branch
|
||||
# above, which exits.
|
||||
if ! grep -aq middlewared "/proc/$PPID/cmdline" 2>/dev/null; then
|
||||
if [ "${TRUECLOUD_REAPPLY:-0}" = "1" ]; then
|
||||
# Invoked by patch/wait_restart.sh as its pre-restart re-apply pass. That
|
||||
# unit already exists to do the restart and verifies the result, so
|
||||
# scheduling another one here would be a loop.
|
||||
echo "Re-apply pass from wait_restart.sh — that unit owns the restart."
|
||||
elif ! grep -aq middlewared "/proc/$PPID/cmdline" 2>/dev/null; then
|
||||
echo "Manual run (parent is not middlewared) — no restart scheduled."
|
||||
elif [ "$_backend_ok" != "1" ]; then
|
||||
echo "Nothing landed on disk — no restart scheduled (nothing new to load)."
|
||||
|
||||
+89
-1
@@ -27,6 +27,50 @@
|
||||
# --wait` below waits for that same queue to drain — the unit would deadlock
|
||||
# on itself until the timeout. apply.sh schedules this with the default
|
||||
# service type, whose start job completes at fork.
|
||||
#
|
||||
# THE RE-APPLY PASS (added 2026-08-26). Applying the patch at PREINIT and
|
||||
# restarting later is only sound if the patched files are still on the live
|
||||
# path at the moment middlewared re-imports them. They may not be: our patch
|
||||
# lives in an overlay mounted *inside* /usr, and anything that remounts the
|
||||
# hierarchy above it detaches or buries that overlay. Two things on a normal
|
||||
# TrueNAS box do exactly that, both AFTER our PREINIT hook has run:
|
||||
#
|
||||
# - `systemd-sysext merge/refresh` over /usr (an nvidia sysext, for
|
||||
# instance) — `Unmerged '/usr'` then `Merged extensions into '/usr'`;
|
||||
# - middlewared's own `docker.configure_nvidia`, which merges the stock
|
||||
# nvidia sysext over /usr when it brings docker up.
|
||||
#
|
||||
# PREINIT scripts run sequentially in id order, so a hook registered after
|
||||
# ours always wins the race, silently. Observed 2026-08-19: our overlay was
|
||||
# mounted at 16:41:56 and a sysext refresh tore /usr down four seconds later;
|
||||
# the restart at 16:47:24 then loaded stock modules and every B2 cloud_backup
|
||||
# job failed for the next nineteen hours while apply.log said "OK".
|
||||
#
|
||||
# Ordering the hooks cannot fix this — docker.configure_nvidia re-merges at
|
||||
# runtime, long after every PREINIT hook is done. So instead of trusting the
|
||||
# PREINIT pass, re-apply immediately before the restart (apply.sh is
|
||||
# idempotent and re-mounts a lost overlay, keeping the same upperdir so
|
||||
# already-patched files survive), verify the marker is really on the live
|
||||
# path, and verify again afterwards.
|
||||
|
||||
PATCH_DIR="$(cd "$(dirname "$0")/.." && pwd)"
|
||||
LOG="$PATCH_DIR/apply.log"
|
||||
|
||||
_log() { echo "[wait_restart] $*" >> "$LOG" 2>/dev/null; }
|
||||
|
||||
# Is the providers patch visible on the live filesystem path -- i.e. would a
|
||||
# middlewared starting right now import it? Reads the marker apply.sh leaves
|
||||
# in restic.py. Returns 0 when patched, 1 when stock, 2 when we cannot tell
|
||||
# (no recorded middlewared dir yet, or the file is gone).
|
||||
_patch_visible() {
|
||||
local mw_dir restic_py
|
||||
mw_dir=$(cat "$PATCH_DIR/.mw_dir" 2>/dev/null)
|
||||
[ -n "$mw_dir" ] || return 2
|
||||
restic_py="$mw_dir/plugins/cloud_backup/restic.py"
|
||||
[ -f "$restic_py" ] || return 2
|
||||
grep -q "TRUECLOUD_PATCH" "$restic_py" 2>/dev/null && return 0
|
||||
return 1
|
||||
}
|
||||
|
||||
# 1. systemd layer: wait for the boot job queue to drain. This covers every
|
||||
# ix-* oneshot still activating, including ix-reporting's in-flight midclt
|
||||
@@ -39,6 +83,9 @@ timeout 900 systemctl is-system-running --wait > /dev/null 2>&1
|
||||
# transitional states (PENDING/INITIALIZING/STOPPING/MIGRATING — see
|
||||
# middlewared/plugins/docker/state_utils.py). An empty answer means
|
||||
# midclt could not respond at all; keep waiting. Cap at 10 minutes.
|
||||
# This also covers docker.configure_nvidia, the runtime /usr re-merge:
|
||||
# waiting for docker to reach a terminal state means the merge that would
|
||||
# bury our overlay has already happened by the time we re-apply below.
|
||||
for _ in $(seq 1 120); do
|
||||
_status=$(midclt call docker.status 2>/dev/null \
|
||||
| grep -oE '"status": "[A-Z_]+"' | cut -d'"' -f4)
|
||||
@@ -52,4 +99,45 @@ done
|
||||
# queryable state (smb.configure and friends). Bounded insurance.
|
||||
sleep 30
|
||||
|
||||
exec systemctl try-restart middlewared
|
||||
# 4. Re-apply pass. Boot has settled, so every sysext merge and docker nvidia
|
||||
# configuration that could bury our overlay is behind us. Re-running
|
||||
# apply.sh is cheap and idempotent: it re-mounts the overlay if it was
|
||||
# detached (same upperdir, so files patched at PREINIT reappear intact)
|
||||
# and re-patches anything that reverted to stock.
|
||||
_patch_visible
|
||||
case $? in
|
||||
0) _log "providers patch still visible on the live path before restart" ;;
|
||||
1) _log "PATCH LOST since PREINIT (something remounted /usr) — re-applying" ;;
|
||||
*) _log "cannot confirm patch state before restart — re-applying anyway" ;;
|
||||
esac
|
||||
|
||||
TRUECLOUD_REAPPLY=1 /bin/bash "$PATCH_DIR/patch/apply.sh"
|
||||
|
||||
if ! _patch_visible; then
|
||||
_log "WARNING: patch is STILL not on the live path after the re-apply pass;"
|
||||
_log "WARNING: restarting anyway, but middlewared will load stock modules."
|
||||
fi
|
||||
|
||||
# 5. The restart itself.
|
||||
systemctl try-restart middlewared
|
||||
|
||||
# 6. Verify what the restart actually loaded, and retry once if the patch was
|
||||
# torn off in the window between the re-apply and the restart. A silent
|
||||
# "on disk but never loaded" is the exact failure this whole script exists
|
||||
# to prevent, so it must never pass unreported.
|
||||
if _patch_visible; then
|
||||
_log "OK: providers patch present on the live path across the restart"
|
||||
else
|
||||
_log "patch missing again after the restart — one more re-apply and restart"
|
||||
TRUECLOUD_REAPPLY=1 /bin/bash "$PATCH_DIR/patch/apply.sh"
|
||||
systemctl try-restart middlewared
|
||||
if _patch_visible; then
|
||||
_log "OK: providers patch loaded after the second attempt"
|
||||
else
|
||||
_log "ERROR: the patch could not be kept on the live path. TrueNAS is"
|
||||
_log "ERROR: running STOCK cloud_backup — B2/S3 backup tasks will fail."
|
||||
_log "ERROR: middlewared raises the 'not loaded' alert for this."
|
||||
fi
|
||||
fi
|
||||
|
||||
_log "=== deferred restart complete ==="
|
||||
|
||||
Reference in New Issue
Block a user