diff --git a/CHANGELOG.md b/CHANGELOG.md index 047818e..ecfe1a8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -63,10 +63,20 @@ worse than no alert, because one day it carries a security fix. So the deferred restart no longer trusts the PREINIT pass. `wait_restart.sh` now re-applies immediately before it restarts middlewared — after boot has settled, which is also after every sysext merge and docker nvidia - configuration — verifies the marker is genuinely on the live path, restarts, - and verifies again, retrying once if the patch was torn off in between. It is - no longer `exec systemctl try-restart middlewared`, because something has to - run afterwards to find out what that restart actually loaded. + configuration — and verifies the marker is genuinely on the live path before + restarting. It is no longer `exec systemctl try-restart middlewared`, because + something has to run afterwards. + + What runs afterwards deliberately does **not** restart again. `try-restart` + returns as soon as middlewared is READY, and middlewared then brings docker up + — `docker.configure_nvidia` merges the stock nvidia sysext over `/usr` at that + point, detaching the overlay *after* the patched modules have already been + imported. A disk check there reports "missing" on a perfectly healthy system, + and restarting on that signal would restart a correctly-patched middlewared + straight back into the same race. So the overlay is re-mounted for the benefit + of the next restart, and the question of whether *this* middlewared actually + holds the patch is left to the one thing that can answer it exactly — the + in-process alert below. Two supporting fixes fell out of the same failure. `_ensure_writable` treated "one of our overlays is listed on this directory" as "already done" — but it diff --git a/patch/wait_restart.sh b/patch/wait_restart.sh index 2347998..8ffb729 100755 --- a/patch/wait_restart.sh +++ b/patch/wait_restart.sh @@ -121,23 +121,29 @@ fi # 5. The restart itself. systemctl try-restart middlewared -# 6. Verify what the restart actually loaded, and retry once if the patch was -# torn off in the window between the re-apply and the restart. A silent -# "on disk but never loaded" is the exact failure this whole script exists -# to prevent, so it must never pass unreported. +# 6. Record what the restart landed on -- but do NOT restart again on a miss. +# +# `try-restart` returns as soon as middlewared is READY; it then brings docker +# up asynchronously, and `docker.configure_nvidia` merges the stock nvidia +# sysext over /usr at that point. That detaches our overlay AFTER the new +# middlewared has already imported the patched modules -- so a disk check here +# can report "missing" on a perfectly healthy system. Restarting on that signal +# would restart a correctly-patched middlewared and then hit the same race +# again, so the disk is deliberately not treated as a verdict after the restart. +# +# The authoritative answer is whether the running process holds the patch, and +# only middlewared can answer that. The alert source installed by apply.sh +# checks exactly that, in-process and hourly, and is what reports a genuine +# miss. What is still worth doing here is putting the overlay back, so the next +# middlewared restart -- whenever and whyever it happens -- finds patched files. if _patch_visible; then _log "OK: providers patch present on the live path across the restart" else - _log "patch missing again after the restart — one more re-apply and restart" + _log "overlay detached again after the restart (expected when docker's" + _log "nvidia sysext merge follows it) -- re-mounting for the next restart." + _log "Whether THIS middlewared loaded the patch is answered in-process by" + _log "the 'installed but NOT loaded' alert, not by this check." TRUECLOUD_REAPPLY=1 /bin/bash "$PATCH_DIR/patch/apply.sh" - systemctl try-restart middlewared - if _patch_visible; then - _log "OK: providers patch loaded after the second attempt" - else - _log "ERROR: the patch could not be kept on the live path. TrueNAS is" - _log "ERROR: running STOCK cloud_backup — B2/S3 backup tasks will fail." - _log "ERROR: middlewared raises the 'not loaded' alert for this." - fi fi _log "=== deferred restart complete ===" diff --git a/tests/test_wait_restart.py b/tests/test_wait_restart.py index 5eefc16..7b0fd8b 100644 --- a/tests/test_wait_restart.py +++ b/tests/test_wait_restart.py @@ -58,6 +58,20 @@ def test_restart_is_not_exec_so_verification_can_follow(): assert not re.search(r"^\s*exec\s+systemctl", src, re.M) +def test_middlewared_is_restarted_exactly_once(): + """No restart loop. + + `try-restart` returns at READY; middlewared then brings docker up, and + docker.configure_nvidia merges the nvidia sysext over /usr right about then + -- detaching the overlay AFTER the patched modules are already imported. A + disk check after the restart therefore false-negatives on a healthy system, + and restarting on that signal would restart a correctly-patched middlewared + straight back into the same race. + """ + src = wait_restart_source() + assert src.count("systemctl try-restart middlewared") == 1 + + def test_patch_is_verified_after_the_restart(): src = wait_restart_source() restart = src.index("systemctl try-restart middlewared")