Files
truenas-truecloud-patch/recover.sh
T
flan a572eb2164 Support ZFS snapshots on datasets with child datasets
TrueCloud Backup's "Take Snapshot" option is rejected on any path containing
child datasets:

  This option is only available for datasets that have no further nesting

That excludes every pool running Apps, where each app is its own dataset and
often has config/pgdata children. Without the option the backup reads live
files, so databases are captured mid-write and an app that continuously
rewrites its files can stall a run as restic chases a moving target.

The stock guard is correct and must not simply be removed. create_snapshot()
already takes a recursive ZFS snapshot, but points the backup tool at the
parent dataset's .zfs/snapshot/, and ZFS does not expose child datasets there:

  /mnt/Tap/.zfs/snapshot/<snap>/apps/                -> 0 entries
  /mnt/Tap/apps/lidarr/config/.zfs/snapshot/<snap>/  -> the real data

Deleting the check would make restic walk a near-empty tree, report success,
and upload almost nothing.

Implement the missing traversal instead. After the recursive snapshot is taken,
each descendant dataset's own .zfs/snapshot/<snap> is bind-mounted into a
staging tree mirroring the original layout, and the backup tool is pointed at
the staging root. The guard is relaxed only after that machinery is in place.

Safety properties:
- staging failure aborts the backup; a partial tree is never handed to restic
- a post-mount pass asserts every target is a mountpoint and the root is
  non-empty, so this cannot regress into the empty backup it exists to prevent
- apply.sh patches crud.py last, so a partial failure leaves the guard intact
  rather than exposing "guard removed, traversal missing"
- every injected block no-ops when _truecloud_nested is absent
- unmountable/locked datasets are skipped and reported, never dropped silently
- scoped to cloud_backup; cloudsync has no teardown wired in, so its guard stays

The staging root is stable per task, so restic can find its parent snapshot
between runs; stock's timestamped .zfs path changes every run and forces a
full re-scan.

Add CI (shellcheck, bash -n, ruff, pytest on 3.11-3.13), including tests that
compile the *_BLOCK strings, which are Python source appended to live
middlewared modules and were previously unchecked.

Also: sync stale version strings, untrack a committed .pyc, gitignore
__pycache__.
2026-07-12 21:20:22 +00:00

77 lines
2.5 KiB
Bash
Executable File

#!/bin/bash
# recover.sh — emergency recovery if middlewared won't start after installing truecloud-patch.
#
# Run this from the TrueNAS shell (local console, SSH, or debug shell):
#
# bash /mnt/tank/truenas-truecloud-patch/recover.sh
#
# What it does:
# 1. Creates a "disabled" file in the repo root — apply.sh checks for this
# file at boot and skips all patching, so the next boot is always clean.
# 2. Unmounts any active truecloud overlays so the original /usr files are
# visible immediately (no reboot required).
# 3. Restarts middlewared against the unpatched files.
#
# To re-enable the patch after investigating:
# rm /mnt/tank/truenas-truecloud-patch/disabled
# bash /mnt/tank/truenas-truecloud-patch/patch/apply.sh
# systemctl restart middlewared
VERSION="0.3.0"
PATCH_DIR="$(cd "$(dirname "$0")" && pwd)"
echo "=== TrueNAS TrueCloud Provider Patch v${VERSION} — Recover ==="
echo ""
if [ "$(id -u)" -ne 0 ]; then
echo "ERROR: must be run as root." >&2
exit 1
fi
if [ ! -d "$PATCH_DIR" ]; then
echo "ERROR: $PATCH_DIR not found — truecloud-patch may not be installed." >&2
exit 1
fi
touch "$PATCH_DIR/disabled"
echo "Kill switch set: $PATCH_DIR/disabled created."
echo "Unmounting truecloud overlays ..."
_any=0
for _tag in mw ui; do
if mount | grep -qF "truecloud-${_tag} on "; then
_mnt=$(mount | grep "truecloud-${_tag} on " | awk '{print $3}' | head -1)
if umount "$_mnt" 2>/dev/null; then
echo " Unmounted: $_mnt"
_any=1
else
echo " WARNING: Could not unmount $_mnt — a reboot will restore original files."
fi
fi
done
[ "$_any" -eq 0 ] && echo " No overlays active."
# Cancel a deferred boot restart if one is still queued — we restart ourselves.
systemctl stop truecloud-mw-restart.service 2>/dev/null
systemctl reset-failed truecloud-mw-restart.service 2>/dev/null
echo "Restarting middlewared ..."
if systemctl restart middlewared; then
echo ""
echo "middlewared started successfully."
echo "Your system is back to normal (Storj-only TrueCloud Backup)."
else
echo ""
echo "WARNING: middlewared did not start cleanly even with the patch disabled."
echo "The problem is unrelated to truecloud-patch."
echo "Check the system log for details:"
echo " journalctl -u middlewared -n 50"
exit 1
fi
echo ""
echo "To re-enable the patch once you have investigated:"
echo " rm $PATCH_DIR/disabled"
echo " bash $PATCH_DIR/patch/apply.sh"
echo " systemctl restart middlewared"