The robot's computer · Chapter 08 · Time: 2 hours, plus 10-20 minutes per backup · Level: Intermediate · Status: Partly test-built
A golden set of the robot's whole disk on the robot and on bigbuddy, checked by checksum and proven by extracting it in a container; and the three written procedures for going back, none of which has been run on this robot.
Every risky step in this guide is reversible only if you can put the disk back the way it was. A golden set is
that safety net: one script on the robot copies the partition table, the small boot partitions and every file of
the root filesystem into a dated folder, with a checksum for each file. You copy the folder to bigbuddy, check the
checksums there, and prove the copy is complete by unpacking it inside the build container from chapter 4. The
first goal of the whole rebuild (docs/status.md, 2026-09-26) was a stable foundation plus a backup of it on
bigbuddy; after that, each milestone the owner approved got its own set.
What has been done on this robot: five golden sets exist on both machines; the oldest and one later set passed the
restore test. What has not been done: no set has ever been written back onto the robot. The restore procedures at
the end of this chapter are written down and checked piece by piece, but not run.
New idea: a golden set
The robot's NVMe has 16 partitions (chapter 5). Partitions 2-15 are small (kernel, device tree, recovery, the
EFI system partitionesp, reserved space), from 512 KB to 480 MB; they are copied raw, byte for byte, and
compressed. Partition 16,APP, is the root filesystem (357.7 GB, most of it empty); it is copied file by file
withtar, keeping owners, permissions, extended attributes and ACLs, so it can be unpacked onto any ext4
filesystem. The partition table is saved withsgdisk --backup, which keeps each partition's name and unique
GUID; the boot lineroot=PARTUUID=ef01d088-...depends on that GUID. Partition 1,APP_old, holds the
factory install and is not copied.
The script lives in ~/golden on the robot, next to the sets it makes.
On your laptop:
ssh rosorin 'mkdir -p ~/golden'
scp -q ~/CCode/rosorin-pro/scripts/make_golden.sh rosorin:golden/make_golden.sh
scripts/make_golden.sh is 42 lines. Build it up in five parts.
On the robot:
#!/bin/bash
# Golden image of the ROSOrin foundation. Run as: sudo bash ~/golden/make_golden.sh
# Reads the disk; writes only to $OUT (excluded from the rootfs archive).
set -euo pipefail
[ "$(id -u)" = 0 ] || { echo "run with sudo"; exit 1; }
DISK=/dev/nvme0n1
OUT=/home/burgerbarn/golden/$(date -u +%Y%m%dT%H%MZ)
mkdir -p "$OUT"
[ "$(findmnt -no SOURCE /)" = "${DISK}p16" ] || { echo "root is not ${DISK}p16, abort"; exit 1; }
Root is needed to read raw partitions. The output folder is named by the UTC time, for example 20260929T1829Z.
The last line stops the script unless the running system's root is partition 16, so it cannot run on the factory
install by mistake.
On the robot:
echo "== GPT"
sgdisk --backup="$OUT/gpt.sgdisk" "$DISK"
sgdisk -p "$DISK" > "$OUT/gpt.txt"
lsblk -bo NAME,SIZE,PARTLABEL,PARTUUID,FSTYPE,UUID "$DISK" > "$OUT/lsblk.txt"
gpt.sgdisk is the binary backup that a restore loads. gpt.txt and lsblk.txt are for you to read.
On the robot:
echo "== boot partitions p2..p15 (raw, zstd)"
umount /boot/efi
trap 'mountpoint -q /boot/efi || mount /boot/efi' EXIT
for n in $(seq 2 15); do
dd if="${DISK}p$n" bs=4M status=none | zstd -q -T0 -19 -o "$OUT/p$n.img.zst"
done
mount /boot/efi
Partition 10 is the EFI system partition, mounted at /boot/efi. It is unmounted first so its image is
consistent. The trap remounts it when the script exits for any reason, including an error halfway. zstd -19 is
slow and small; these partitions are small enough that it does not matter.
On the robot:
echo "== rootfs (file-level tar, zstd)"
cp /boot/extlinux/extlinux.conf /etc/nv_boot_control.conf /etc/fstab "$OUT/"
tune2fs -l "${DISK}p16" | grep -E "Filesystem features|Inode size|Block size|Filesystem volume name|Filesystem UUID" > "$OUT/rootfs_tune2fs.txt"
find / -xdev \( -type f -o -type l \) -not -path "/home/burgerbarn/golden/*" \
-not -path "/tmp/*" -not -path "/var/tmp/*" | wc -l > "$OUT/rootfs_filecount.txt" # same exclusions as the tar
Three config files are copied for reference, the filesystem's settings are noted, and the files and symlinks on /
are counted with the same exclusions the archive uses. The restore test compares against this count.
On the robot:
# GNU tar exits 1 when a live file changes while read (warning); 2 = fatal.
set +e
tar -C / --one-file-system --numeric-owner --xattrs --xattrs-include='*' --acls \
--exclude=./home/burgerbarn/golden --exclude='./tmp/*' --exclude='./var/tmp/*' \
-cpf - . 2> "$OUT/tar_warnings.txt" | zstd -q -T0 -10 -o "$OUT/rootfs.tar.zst"
rc=("${PIPESTATUS[@]}"); set -e
echo "tar rc=${rc[0]} zstd rc=${rc[1]}"; cat "$OUT/tar_warnings.txt"
[ "${rc[0]}" -le 1 ] && [ "${rc[1]}" -eq 0 ] || { echo "rootfs archive FAILED"; exit 1; }
--one-file-system stays on partition 16 (no /proc, /sys, /boot/efi).--numeric-owner stores user and group numbers, not names, so a restore on another machine keeps the robot's--xattrs --xattrs-include='*' --acls keep extended attributes and ACLs (file capabilities live in extendedset +e and PIPESTATUS let the script accept 1 and fail only on 2 or on a zstd error.On the robot:
echo "== checksums"
( cd "$OUT" && sha256sum -- * > SHA256SUMS )
chown -R burgerbarn:burgerbarn /home/burgerbarn/golden
du -sh "$OUT"; ls -la "$OUT"
echo "GOLDEN_DONE $OUT"
One SHA-256 line per file, then the set is handed to your user so you can copy it without sudo.
~/golden/make_golden.sh (identical to the robot's copy, checked 2026-10-07):
On the robot:
#!/bin/bash
# Golden image of the ROSOrin foundation. Run as: sudo bash ~/golden/make_golden.sh
# Reads the disk; writes only to $OUT (excluded from the rootfs archive).
set -euo pipefail
[ "$(id -u)" = 0 ] || { echo "run with sudo"; exit 1; }
DISK=/dev/nvme0n1
OUT=/home/burgerbarn/golden/$(date -u +%Y%m%dT%H%MZ)
mkdir -p "$OUT"
[ "$(findmnt -no SOURCE /)" = "${DISK}p16" ] || { echo "root is not ${DISK}p16, abort"; exit 1; }
echo "== GPT"
sgdisk --backup="$OUT/gpt.sgdisk" "$DISK"
sgdisk -p "$DISK" > "$OUT/gpt.txt"
lsblk -bo NAME,SIZE,PARTLABEL,PARTUUID,FSTYPE,UUID "$DISK" > "$OUT/lsblk.txt"
echo "== boot partitions p2..p15 (raw, zstd)"
umount /boot/efi
trap 'mountpoint -q /boot/efi || mount /boot/efi' EXIT
for n in $(seq 2 15); do
dd if="${DISK}p$n" bs=4M status=none | zstd -q -T0 -19 -o "$OUT/p$n.img.zst"
done
mount /boot/efi
echo "== rootfs (file-level tar, zstd)"
cp /boot/extlinux/extlinux.conf /etc/nv_boot_control.conf /etc/fstab "$OUT/"
tune2fs -l "${DISK}p16" | grep -E "Filesystem features|Inode size|Block size|Filesystem volume name|Filesystem UUID" > "$OUT/rootfs_tune2fs.txt"
find / -xdev \( -type f -o -type l \) -not -path "/home/burgerbarn/golden/*" \
-not -path "/tmp/*" -not -path "/var/tmp/*" | wc -l > "$OUT/rootfs_filecount.txt" # same exclusions as the tar
# GNU tar exits 1 when a live file changes while read (warning); 2 = fatal.
set +e
tar -C / --one-file-system --numeric-owner --xattrs --xattrs-include='*' --acls \
--exclude=./home/burgerbarn/golden --exclude='./tmp/*' --exclude='./var/tmp/*' \
-cpf - . 2> "$OUT/tar_warnings.txt" | zstd -q -T0 -10 -o "$OUT/rootfs.tar.zst"
rc=("${PIPESTATUS[@]}"); set -e
echo "tar rc=${rc[0]} zstd rc=${rc[1]}"; cat "$OUT/tar_warnings.txt"
[ "${rc[0]}" -le 1 ] && [ "${rc[1]}" -eq 0 ] || { echo "rootfs archive FAILED"; exit 1; }
echo "== checksums"
( cd "$OUT" && sha256sum -- * > SHA256SUMS )
chown -R burgerbarn:burgerbarn /home/burgerbarn/golden
du -sh "$OUT"; ls -la "$OUT"
echo "GOLDEN_DONE $OUT"
If it fails
- The first run, on 2026-09-26, had no
PIPESTATUShandling. tar exited 1 because a log file changed while it
was read,set -e -o pipefailkilled the script in the middle, and the half-made set20260926T1856Zhad
to be thrown away. Theset +eblock above is the fix (docs/lessons.md 2026-09-26).- Set
20260928T0110Zcame out withoutrootfs_filecount.txt: an edit had put the# same exclusionscomment
in front of the> "$OUT/rootfs_filecount.txt"redirect, so the count went nowhere. The file above has the
comment after the redirect. After every run, check thatrootfs_filecount.txtexists.root is not /dev/nvme0n1p16, abort: you are on another system (the factory install, or a restored disk with
different numbering).- If
/boot/efiis not mounted afterwards, runsudo mount /boot/efi. After the first run it was mounted
again as/dev/nvme0n1p10.
Make a set at every milestone, and before any step you might want to undo. The robot keeps running while the set is
made; the run took about 10 minutes for a 9.2 GB set on 2026-09-29. Run it in the background so a dropped ssh
session does not kill it, and watch the log:
On the robot:
nohup sudo -n bash ~/golden/make_golden.sh > /tmp/golden.log 2>&1 &
tail -f /tmp/golden.log
Check
The log starts with== GPTandThe operation has completed successfully.(fromsgdisk). The end of a run
on 2026-09-27 06:11 UTC:tar rc=1 zstd rc=0 3.5G /home/burgerbarn/golden/20260927T0611Z GOLDEN_DONE /home/burgerbarn/golden/20260927T0611Z
tar rc=1is the accepted warning;tar_warnings.txtlists which files changed.
Then check the set on the robot itself (use your set's name):
On the robot:
cd ~/golden/20260929T1829Z && sha256sum -c --quiet SHA256SUMS && echo robot SHA256 OK; du -sh .
ls ~/golden/20260929T1829Z
Check
On 2026-09-29 this printedrobot SHA256 OKand9.2G .. A set holds 25 files:p2.img.zstto
p15.img.zst,rootfs.tar.zst,gpt.sgdisk,gpt.txt,lsblk.txt,extlinux.conf,
nv_boot_control.conf,fstab,rootfs_filecount.txt,rootfs_tune2fs.txt,tar_warnings.txt, and
SHA256SUMS, which lists the other 24.
The sets that exist now, on the robot (~/golden) and on bigbuddy (~/rosorin-golden), newest first:
| Set | Size | Files on / |
What it adds |
|---|---|---|---|
20260929T1829Z |
9.2 GB | 280,910 | camera service at boot, auto-tuck, vision stack (CUDA, cuDNN, TensorRT, ~/vision) |
20260929T1441Z |
9.2 GB | 280,828 | vision stack |
20260928T2103Z |
4.0 GB | 245,566 | Nav2, SLAM, room map, Aurora 930 driver, URDF |
20260927T1719Z |
3.5 GB | 217,149 | arm tuck and drive poses |
20260926T2041Z |
2.2 GB | 188,312 | foundation only: OS, user, ssh |
The owner deleted seven other sets on 2026-09-27 and 2026-09-28 to free space (listed in docs/restore.md).
What a set contains, and what it does not
- Everything after 2026-09-29 is in no set. Persistent journal, the DDS host-id fix, headless boot,
Wi-Fi power save off and the Bluetooth fix (chapters 6, 7, 9) all came later. Make a new set before you rely
on any of them.- A set made today would be much larger. On 2026-10-07 the root filesystem held 73 GB outside
~/golden,
including 29 GB of container images in/var/lib/containerdand 17 GB of recordings in~/bags. The size
and run time of such a set have not been measured. The robot had 237 GB free.- The sets from
20260928T2103Zon contain a secret. The depth camera's vendor driver arrived from the
factory partition with a GitLab access token inside its git remote URL (chapter 15 removes it from the live
robot). The old sets keep it. Treat every set as private: it stays on the robot and on bigbuddy. Nothing goes
toserverwithout the owner's explicit OK, and the owner's external SSD on bigbuddy is not for robot data
(owner, 2026-10-01).- Not in any set: the QSPI bootloader (it lives on the Jetson module's own flash chip, not on the NVMe) and the
contents ofAPP_old.
Run this on the Mac. It streams the set as a tar archive from the robot to bigbuddy over two ssh connections and,
on bigbuddy, checks every file against SHA256SUMS before saying OK.
On your laptop:
ssh bigbuddy 'mkdir -p /home/burgerbarn/rosorin-golden'
SET=20260929T1829Z
ssh rosorin "tar -C /home/burgerbarn/golden -cf - $SET" | ssh bigbuddy "tar -C /home/burgerbarn/rosorin-golden -xf - && cd /home/burgerbarn/rosorin-golden/$SET && sha256sum -c --quiet SHA256SUMS && echo SHA256 OK \$(wc -l < SHA256SUMS) files && df -h /home | tail -1"
Set SET to the folder name make_golden.sh printed. bigbuddy's login shell is fish; this command line works
there as written (it was run this way on 2026-09-27). Later sets were copied with scp -3 -q "rosorin:golden/$G/*" "bigbuddy:rosorin-golden/$G/" and then checked with the same sha256sum -c.
Check
On 2026-09-27 17:24 UTC:SHA256 OK 24 files
sha256sum -c --quietprints nothing for good files and aFAILEDline for each bad one; theechoonly runs
if all 24 match.
If it fails
- bigbuddy's disk ran full: 98 % used on 2026-09-29 (21 GB free) while sets were being copied. The owner
cleared space (451 GB free on 2026-10-07). Checkdf -h /homeon bigbuddy before you copy a large set.- In commands typed on the Mac, an unquoted
~is expanded by the Mac's shell into/Users/matty/...before
ssh sends it. That is why every remote path above is written out as/home/burgerbarn/...
(docs/lessons.md 2026-09-26).- bigbuddy suspends when idle, and ssh traffic does not keep it awake (docs/lessons.md 2026-09-26). If the copy
stops, check that bigbuddy is awake.
A checksum proves the copy matches the robot's files. It does not prove the set can rebuild a disk. The restore
test does that without touching any disk: it runs inside the jetson-flash:r36.4.3 container on bigbuddy (Ubuntu
22.04, built in chapter 4), which has the same sgdisk, zstd, fsck.vfat and mke2fs 1.46.5 as the robot.
On bigbuddy the folder of sets is ~/rosorin-golden; inside the container it appears as /g.
scripts/restore_test.sh, 24 lines, run inside the container with the sets mounted at /g:
On bigbuddy:
set -uo pipefail
G=/g/${1:-20260926T2041Z}; T=/g/restore-test
fail=0
echo "== zstd integrity"
for f in $G/*.zst; do zstd -tq "$f" || { echo "BAD $f"; fail=1; }; done; echo "zst files tested: $(ls $G/*.zst | wc -l)"
The set name is the first argument. set -e is deliberately missing: every check runs, and each failure sets
fail=1. First, zstd -t decompresses every archive in memory to prove none is damaged.
On bigbuddy:
echo "== GPT restore onto sparse disk-size file"
truncate -s $((1000215216*512)) /tmp/disk.img
sgdisk --load-backup=$G/gpt.sgdisk /tmp/disk.img >/dev/null
sgdisk -p /tmp/disk.img | sed -n "/^Number/,\$p" > /tmp/gpt_restored.txt
sed -n "/^Number/,\$p" $G/gpt.txt > /tmp/gpt_orig.txt
diff /tmp/gpt_orig.txt /tmp/gpt_restored.txt && echo "GPT partition table identical" || fail=1
rm -f /tmp/disk.img
Make an empty file exactly the size of the robot's NVMe (1,000,215,216 sectors of 512 bytes; truncate makes it
sparse, so it uses no real space), load the saved partition table onto it, and compare the result with the robot's
own table.
On bigbuddy:
echo "== ESP image fsck"
zstd -dcq $G/p10.img.zst > /tmp/esp.img && fsck.vfat -n /tmp/esp.img | tail -2; rm -f /tmp/esp.img
Unpack the EFI system partition image and check its FAT filesystem read-only (-n).
On bigbuddy:
echo "== rootfs extract + count"
rm -rf $T; mkdir -p $T
zstd -dcq $G/rootfs.tar.zst | tar -C $T --numeric-owner --xattrs --xattrs-include="*" --acls -xpf - || fail=1
got=$(find $T -xdev \( -type f -o -type l \) | wc -l); want=$(cat $G/rootfs_filecount.txt)
echo "restored files+links: $got robot count: $want diff: $((want-got))"
Unpack the whole root filesystem the way a real restore would, and count what came out against the robot's count.
On bigbuddy:
for f in boot/Image boot/initrd boot/extlinux/extlinux.conf etc/nv_boot_control.conf etc/hostname home/burgerbarn/.ssh/authorized_keys etc/ssh/sshd_config.d/10-local.conf; do test -e $T/$f && echo "present $f" || { echo "MISSING $f"; fail=1; }; done
stat -c "%u:%g %a %n" $T/home/burgerbarn/.ssh/authorized_keys $T/usr/bin/sudo
grep -c . $T/etc/nv_boot_control.conf
rm -rf $T
echo "RESULT fail=$fail"
The files the robot cannot boot or be reached without must be present: kernel, initrd, boot menu, NVIDIA's boot
control file, hostname, your ssh key and the ssh config. Owners and modes must survive: your key must be 1000:1000 600, and sudo must keep its setuid bit (4755). Then the unpacked tree is deleted.
The complete file, ~/rosorin-golden/restore_test.sh on bigbuddy (identical to the repo copy, checked
2026-10-07):
On bigbuddy:
set -uo pipefail
G=/g/${1:-20260926T2041Z}; T=/g/restore-test
fail=0
echo "== zstd integrity"
for f in $G/*.zst; do zstd -tq "$f" || { echo "BAD $f"; fail=1; }; done; echo "zst files tested: $(ls $G/*.zst | wc -l)"
echo "== GPT restore onto sparse disk-size file"
truncate -s $((1000215216*512)) /tmp/disk.img
sgdisk --load-backup=$G/gpt.sgdisk /tmp/disk.img >/dev/null
sgdisk -p /tmp/disk.img | sed -n "/^Number/,\$p" > /tmp/gpt_restored.txt
sed -n "/^Number/,\$p" $G/gpt.txt > /tmp/gpt_orig.txt
diff /tmp/gpt_orig.txt /tmp/gpt_restored.txt && echo "GPT partition table identical" || fail=1
rm -f /tmp/disk.img
echo "== ESP image fsck"
zstd -dcq $G/p10.img.zst > /tmp/esp.img && fsck.vfat -n /tmp/esp.img | tail -2; rm -f /tmp/esp.img
echo "== rootfs extract + count"
rm -rf $T; mkdir -p $T
zstd -dcq $G/rootfs.tar.zst | tar -C $T --numeric-owner --xattrs --xattrs-include="*" --acls -xpf - || fail=1
got=$(find $T -xdev \( -type f -o -type l \) | wc -l); want=$(cat $G/rootfs_filecount.txt)
echo "restored files+links: $got robot count: $want diff: $((want-got))"
for f in boot/Image boot/initrd boot/extlinux/extlinux.conf etc/nv_boot_control.conf etc/hostname home/burgerbarn/.ssh/authorized_keys etc/ssh/sshd_config.d/10-local.conf; do test -e $T/$f && echo "present $f" || { echo "MISSING $f"; fail=1; }; done
stat -c "%u:%g %a %n" $T/home/burgerbarn/.ssh/authorized_keys $T/usr/bin/sudo
grep -c . $T/etc/nv_boot_control.conf
rm -rf $T
echo "RESULT fail=$fail"
Copy it to bigbuddy once, then run it for each set from the Mac:
On your laptop:
ssh bigbuddy 'cat > /home/burgerbarn/rosorin-golden/restore_test.sh' < ~/CCode/rosorin-pro/scripts/restore_test.sh
SET=20260929T1829Z
ssh bigbuddy "docker run --rm -v /home/burgerbarn/rosorin-golden:/g:z jetson-flash:r36.4.3 bash /g/restore_test.sh $SET 2>&1"
-v .../rosorin-golden:/g:z mounts the folder of sets into the container at /g (:z lets SELinux on Fedora
allow it). --rm deletes the container afterwards.
Check
The first set,20260926T2041Z, on 2026-09-26 20:54 UTC (the command log cut the last lines):== zstd integrity zst files tested: 15 == GPT restore onto sparse disk-size file GPT partition table identical == ESP image fsck fsck.fat 4.2 (2021-01-31) /tmp/esp.img: 5 files, 221/129022 clusters == rootfs extract + count restored files+links: 188312 robot count: 188312 diff: 0 present boot/Image present boot/initrd present boot/extlinux/extlinux.conf present etc/nv_boot_control.conf present etc/hostname present home/burgerbarn/.ssh/authorized_keys present etc/ssh/sshd_config.d/10-local.conf 1000:1000 600 /g/restore-test/home/burgerbarn/.ssh/authorized_keysand
docs/status.mdrecordsfail=0for it. Set20260927T1719Zon 2026-09-27:GPT partition table identical restored files+links: 217058 robot count: 217149 diff: 91 RESULT fail=0
RESULT fail=0is the line that matters. The script reports a difference in the file count but does not fail
on it; the count is taken a moment before the archive while the system runs.
A second check from the same day shows why the saved table matters: loaded onto a blank file, it gives back the
partitions' unique GUIDs, so the boot line keeps working.
On your laptop:
ssh bigbuddy 'docker run --rm -v /home/burgerbarn/rosorin-golden:/g:z jetson-flash:r36.4.3 bash -c "truncate -s \$((1000215216*512)) /tmp/d.img && sgdisk --load-backup=/g/20260926T2041Z/gpt.sgdisk /tmp/d.img >/dev/null && for n in 1 10 16; do sgdisk -i \$n /tmp/d.img | grep -E \"unique GUID|Partition name\" | tr -s \" \" | tr \"\n\" \" \"; echo; done; rm /tmp/d.img"'
Check
On 2026-09-26 21:01 UTC:Partition unique GUID: 59FCA1C4-9E3E-406D-8532-CAD56C869912 Partition name: 'APP_old' Partition unique GUID: 2929D5B9-C45C-40D3-8FB9-4E7464327232 Partition name: 'esp' Partition unique GUID: EF01D088-5975-4263-B620-2D0F5137DFE1 Partition name: 'APP'
If it fails
docker: ... Unable to find image 'jetson-flash:r36.4.3': build the container first (chapter 4).
docker images | grep jetson-flashon bigbuddy listsjetson-flash:r36.4.3(image4a08465dc9c1).restore_test.shwrites/g/restore-testinside the mounted folder, so the mount must be writable and
bigbuddy needs free space for one fully unpacked root filesystem.- Which sets were restore-tested:
20260926T2041Zand20260927T1719Z(bothfail=0), plus sets the owner
later deleted. For20260928T2103Z,20260929T1441Zand20260929T1829Zthe command log shows only the
checksum check on both machines, no restore test. Run the test on any set you plan to rely on.
docs/restore.md describes three procedures. They are copied here verbatim. None of them has been run on this
robot. The only related step ever run is the forward one in chapter 5: renaming the partitions so APP points at
the new system.
Rules that apply to all three:
jetson-flash:r36.4.3orphan_file, which the robot'smke2fs 1.46.5 (30-Dec-2021), bigbuddy says mke2fs 1.47.3 (8-Jul-2025) with orphan_file in its defaults.--numeric-owner --xattrs --xattrs-include='*' --acls -p.root=PARTUUID=ef01d088-5975-4263-b620-2d0f5137dfe1. If partition 16 is recreated by any meanssgdisk --load-backup, it gets a new GUID, and boot/extlinux/extlinux.conf must be fixed.APP. Only one partition may have that name.The factory system is still on partition 1 (APP_old, 117.8 GB, checked 2026-10-07). Swapping the names makes it
the one that boots:
On your laptop:
ssh -t rosorin 'sudo sgdisk -c 16:APP_new -c 1:APP /dev/nvme0n1 && sudo systemctl reboot'
The factory system comes up as user ubuntu on its own address (its Wi-Fi used 192.168.1.108). To return, from the
factory system:
On the robot:
sudo sgdisk -c 1:APP_old -c 16:APP /dev/nvme0n1 && sudo systemctl reboot
Partition 16 keeps its GUID because it is reformatted, not recreated, so extlinux.conf stays valid. Boot the
factory system with procedure A, then from the Mac (the archive is decompressed on bigbuddy because whether the
factory system has zstd is unverified):
On your laptop:
ssh rosorin-factory 'sudo umount /dev/nvme0n1p16 2>/dev/null; sudo mkfs.ext4 -F -L APP /dev/nvme0n1p16 && sudo mkdir -p /mnt/app && sudo mount /dev/nvme0n1p16 /mnt/app'
ssh bigbuddy 'zstd -dc /home/burgerbarn/rosorin-golden/20260926T2041Z/rootfs.tar.zst' \
| ssh rosorin-factory "sudo tar -C /mnt/app --numeric-owner --xattrs --xattrs-include='*' --acls -xpf -"
ssh rosorin-factory 'sudo find /mnt/app -xdev \( -type f -o -type l \) | wc -l' # expect 188312
ssh rosorin-factory 'sudo umount /mnt/app && sudo sgdisk -c 1:APP_old -c 16:APP /dev/nvme0n1 && sudo systemctl reboot'
Substitute your set's name and its count from rootfs_filecount.txt. The factory system's mke2fs is Ubuntu 22.04's
1.46.5. Afterwards: ssh burgerbarn@<ip>, hostname must print rosorin, and systemctl --failed must be empty.
rosorin-factory means ubuntu@<factory IP>. There is no such alias in the Mac's ~/.ssh/config today; add one
before you need it. The factory image has passwordless sudo for ubuntu.
This needs the NVMe in a USB enclosure on a Linux host such as bigbuddy (where it appears as /dev/sdX), or the
robot in recovery mode with NVIDIA's initrd flash (chapter 5). The target disk must be at least as large as the
original (1,000,215,216 sectors), because partition 16 ends at the original disk's end. On a smaller disk the
procedure is unverified and needs a hand-made partition 16 plus a PARTUUID fix. All steps run as root inside the
build container:
On bigbuddy:
D=/dev/sdX # CHECK THIS. Everything on it is overwritten.
G=/home/burgerbarn/rosorin-golden/20260926T2041Z
docker run --rm -it --privileged -v /dev:/dev -v $G:/g:ro,z jetson-flash:r36.4.3 bash -c "
set -e
cd /g && sha256sum -c --quiet SHA256SUMS
sgdisk --load-backup=/g/gpt.sgdisk $D && sgdisk -e $D && partprobe $D && sleep 2
for n in \$(seq 2 15); do zstd -dcq /g/p\$n.img.zst | dd of=${D}\$n bs=4M conv=fsync status=none; done
mkfs.ext4 -F -L APP ${D}16
mkdir -p /mnt/app && mount ${D}16 /mnt/app
zstd -dcq /g/rootfs.tar.zst | tar -C /mnt/app --numeric-owner --xattrs --xattrs-include='*' --acls -xpf -
find /mnt/app -xdev \( -type f -o -type l \) | wc -l # expect 188312
sgdisk -i 16 $D | grep 'unique GUID' # expect EF01D088-5975-4263-B620-2D0F5137DFE1
umount /mnt/app
"
bigbuddy's login shell is fish: open bash first, then type these lines. Partition device names differ by host:
/dev/sdX2 on a USB host, /dev/nvme0n1p2 on the robot; adjust ${D}\$n. After C, partition 1 (APP_old) is
listed in the table but empty; its name does not match APP, so it is harmless. The first boot after C does not
grow the filesystem: NVIDIA's one-time resize trigger was used up on the original first boot.
Safety
Procedure C overwrites every byte of the disk named inD. bigbuddy has its own disks and, at times, the
owner's external SSD with his files on it; a wrong letter destroys them. CheckDwithlsblkimmediately before you run
it, with the robot's disk plugged in and nothing else.
Not test-built
- A, B and C have never been run on this robot. A is the reverse of the rename that was run on 2026-09-26.
B's and C's pieces were proven separately (GPT restore with GUIDs, rootfs extract with owners and count,
ESP image check), but not end to end, and never followed by a boot.- Whether the factory system has
zstdis unverified (B decompresses on bigbuddy for that reason).- A smaller target disk than 1,000,215,216 sectors is unverified.
- The QSPI bootloader is not covered by any set. NVIDIA's initrd flash (README_initrd_flash.txt, Workflow 4) is
the way to rewrite it; nothing in this project has done so.- The factory state before the rebuild is covered separately: the owner made a full clone of the original disk,
and source and notes from it were copied tobigbuddy:~/rosorin-backup-2026-09-26(docs/status.md).
ls ~/golden on the robot lists your new set next to make_golden.sh, and rootfs_filecount.txt exists in it.cd ~/golden/<set> && sha256sum -c --quiet SHA256SUMS && echo OK prints OK on the robot.SHA256 OK 24 files on bigbuddy in ~/rosorin-golden/<set>.restore_test.sh <set> in the container ends with RESULT fail=0 and diff: close to 0.Where this comes from
scripts/make_golden.sh(first commit de7599a 2026-09-26, last fa7c278 2026-09-27; identical to
~/golden/make_golden.shon the robot, 2026-10-07),scripts/restore_test.sh(de7599a, d96d324; identical to
bigbuddy's copy),docs/restore.md(sets, rules, procedures A/B/C),docs/lessons.md2026-09-26 (tar exit 1,
mke2fs orphan_file, quoting~, bigbuddy suspend),docs/status.md(first golden set and restore test
2026-09-26, factory clone),docs/hardware.md(NVMe size, bigbuddy storage).
Command log (UTC): first set and failed run 2026-09-26 18:56-20:41, copy and restore test 20:51-20:54, GUID
check 21:01; set20260927T0611Z2026-09-27 06:11; restore test of20260927T1719Z17:24; file-count bug
2026-09-28 01:10; set20260929T1829Z2026-09-29 18:29. survey_live_robot.md (GitLab token in sets,~/golden
28 GB). Live read-back 2026-10-07: set sizes on both machines, partition table, root filesystem usage (73 GB
outside~/golden), bigbuddy free space, container image.
← Add the missing drivers · Contents · Install ROS 2 Humble →