Compare commits
2
Commits
v0.8.0-rc1
...
v0.8.0-rc3
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5f8d42f2cf | ||
|
|
3915f92dec |
+21
-4
@@ -63,10 +63,27 @@ worse than no alert, because one day it carries a security fix.
|
||||
So the deferred restart no longer trusts the PREINIT pass. `wait_restart.sh`
|
||||
now re-applies immediately before it restarts middlewared — after boot has
|
||||
settled, which is also after every sysext merge and docker nvidia
|
||||
configuration — verifies the marker is genuinely on the live path, restarts,
|
||||
and verifies again, retrying once if the patch was torn off in between. It is
|
||||
no longer `exec systemctl try-restart middlewared`, because something has to
|
||||
run afterwards to find out what that restart actually loaded.
|
||||
configuration — and verifies the marker is genuinely on the live path before
|
||||
restarting. It is no longer `exec systemctl try-restart middlewared`, because
|
||||
something has to run afterwards.
|
||||
|
||||
What runs afterwards deliberately does **not** restart again. `try-restart`
|
||||
returns as soon as middlewared is READY, and middlewared then brings docker up
|
||||
— `docker.configure_nvidia` merges the stock nvidia sysext over `/usr` at that
|
||||
point, detaching the overlay *after* the patched modules have already been
|
||||
imported. A disk check there reports "missing" on a perfectly healthy system,
|
||||
and restarting on that signal would restart a correctly-patched middlewared
|
||||
straight back into the same race. So the overlay is re-mounted for the benefit
|
||||
of the next restart, and the question of whether *this* middlewared actually
|
||||
holds the patch is left to the one thing that can answer it exactly — the
|
||||
in-process alert below. That re-mount preserves `hook_status.json`'s
|
||||
`patched_at`: `create_task.py verify` decides "loaded" by comparing
|
||||
middlewared's start time against that stamp, so a re-apply running *after* the
|
||||
restart would have made the stamp newer than the process which correctly
|
||||
imported the patch, and `verify` would have reported FAIL forever on every
|
||||
boot where the sysext merge detaches the overlay. Caught on hardware while
|
||||
validating the candidate — a new lying status introduced by the fix for a
|
||||
lying status.
|
||||
|
||||
Two supporting fixes fell out of the same failure. `_ensure_writable` treated
|
||||
"one of our overlays is listed on this directory" as "already done" — but it
|
||||
|
||||
+37
-12
@@ -121,22 +121,47 @@ fi
|
||||
# 5. The restart itself.
|
||||
systemctl try-restart middlewared
|
||||
|
||||
# 6. Verify what the restart actually loaded, and retry once if the patch was
|
||||
# torn off in the window between the re-apply and the restart. A silent
|
||||
# "on disk but never loaded" is the exact failure this whole script exists
|
||||
# to prevent, so it must never pass unreported.
|
||||
# 6. Record what the restart landed on -- but do NOT restart again on a miss.
|
||||
#
|
||||
# `try-restart` returns as soon as middlewared is READY; it then brings docker
|
||||
# up asynchronously, and `docker.configure_nvidia` merges the stock nvidia
|
||||
# sysext over /usr at that point. That detaches our overlay AFTER the new
|
||||
# middlewared has already imported the patched modules -- so a disk check here
|
||||
# can report "missing" on a perfectly healthy system. Restarting on that signal
|
||||
# would restart a correctly-patched middlewared and then hit the same race
|
||||
# again, so the disk is deliberately not treated as a verdict after the restart.
|
||||
#
|
||||
# The authoritative answer is whether the running process holds the patch, and
|
||||
# only middlewared can answer that. The alert source installed by apply.sh
|
||||
# checks exactly that, in-process and hourly, and is what reports a genuine
|
||||
# miss. What is still worth doing here is putting the overlay back, so the next
|
||||
# middlewared restart -- whenever and whyever it happens -- finds patched files.
|
||||
if _patch_visible; then
|
||||
_log "OK: providers patch present on the live path across the restart"
|
||||
else
|
||||
_log "patch missing again after the restart — one more re-apply and restart"
|
||||
_log "overlay detached again after the restart (expected when docker's"
|
||||
_log "nvidia sysext merge follows it) -- re-mounting for the next restart."
|
||||
_log "Whether THIS middlewared loaded the patch is answered in-process by"
|
||||
_log "the 'installed but NOT loaded' alert, not by this check."
|
||||
|
||||
# Preserve hook_status.json's patched_at across this re-mount.
|
||||
#
|
||||
# create_task.py verify decides "loaded" by comparing middlewared's start
|
||||
# time against patched_at. This re-apply restores the SAME patch the boot
|
||||
# pass already applied, but it runs *after* the restart -- so letting it
|
||||
# re-stamp would make patched_at newer than the process that correctly
|
||||
# imported the patch, and verify would report FAIL forever, on every boot
|
||||
# where docker's sysext merge detaches the overlay. That is precisely the
|
||||
# lying-status failure this release exists to remove, so do not introduce a
|
||||
# new one. The snapshot lives in /run, never in the repo: a leftover file
|
||||
# there would leave the tree dirty and update.sh refuses to run over that.
|
||||
_saved_status=/run/truecloud-hook_status.pre
|
||||
cp -p "$PATCH_DIR/hook_status.json" "$_saved_status" 2>/dev/null
|
||||
|
||||
TRUECLOUD_REAPPLY=1 /bin/bash "$PATCH_DIR/patch/apply.sh"
|
||||
systemctl try-restart middlewared
|
||||
if _patch_visible; then
|
||||
_log "OK: providers patch loaded after the second attempt"
|
||||
else
|
||||
_log "ERROR: the patch could not be kept on the live path. TrueNAS is"
|
||||
_log "ERROR: running STOCK cloud_backup — B2/S3 backup tasks will fail."
|
||||
_log "ERROR: middlewared raises the 'not loaded' alert for this."
|
||||
|
||||
if [ -f "$_saved_status" ]; then
|
||||
mv -f "$_saved_status" "$PATCH_DIR/hook_status.json" 2>/dev/null
|
||||
fi
|
||||
fi
|
||||
|
||||
|
||||
@@ -58,6 +58,20 @@ def test_restart_is_not_exec_so_verification_can_follow():
|
||||
assert not re.search(r"^\s*exec\s+systemctl", src, re.M)
|
||||
|
||||
|
||||
def test_middlewared_is_restarted_exactly_once():
|
||||
"""No restart loop.
|
||||
|
||||
`try-restart` returns at READY; middlewared then brings docker up, and
|
||||
docker.configure_nvidia merges the nvidia sysext over /usr right about then
|
||||
-- detaching the overlay AFTER the patched modules are already imported. A
|
||||
disk check after the restart therefore false-negatives on a healthy system,
|
||||
and restarting on that signal would restart a correctly-patched middlewared
|
||||
straight back into the same race.
|
||||
"""
|
||||
src = wait_restart_source()
|
||||
assert src.count("systemctl try-restart middlewared") == 1
|
||||
|
||||
|
||||
def test_patch_is_verified_after_the_restart():
|
||||
src = wait_restart_source()
|
||||
restart = src.index("systemctl try-restart middlewared")
|
||||
@@ -158,3 +172,28 @@ def test_mount_retries_on_a_private_workdir():
|
||||
body = src[start:end]
|
||||
assert body.count("mount -t overlay") == 2, "expected a retry mount"
|
||||
assert 'work="/run/truecloud-${tag}-work.$$"' in body
|
||||
|
||||
|
||||
def test_post_restart_remount_preserves_the_patched_at_stamp():
|
||||
"""create_task.py verify compares middlewared's start time to patched_at.
|
||||
|
||||
The post-restart re-mount restores the same patch the boot pass applied, so
|
||||
letting apply.sh re-stamp would make patched_at newer than the process that
|
||||
correctly imported it -- verify would then report FAIL forever on every boot
|
||||
where docker's sysext merge detaches the overlay.
|
||||
"""
|
||||
src = wait_restart_source()
|
||||
tail = src[src.index("systemctl try-restart middlewared"):]
|
||||
assert "hook_status.json" in tail
|
||||
save = tail.index("/run/truecloud-hook_status.pre")
|
||||
reapply = tail.index("TRUECLOUD_REAPPLY=1")
|
||||
restore = tail.rindex("hook_status.json")
|
||||
assert save < reapply < restore, "snapshot must bracket the re-apply"
|
||||
|
||||
|
||||
def test_status_snapshot_is_not_written_into_the_repo():
|
||||
"""A leftover file in the repo dir leaves the tree dirty, and update.sh
|
||||
refuses to run over a dirty tree -- that once made the patch un-updatable."""
|
||||
src = wait_restart_source()
|
||||
assert "/run/truecloud-hook_status.pre" in src
|
||||
assert '"$PATCH_DIR/hook_status.json.pre' not in src
|
||||
|
||||
Reference in New Issue
Block a user