The robot's computer · Chapter 09 · Time: 2 hours · Level: Beginner · Status: Partly test-built
ROS 2 Humble's base packages from the official ROS apt repository with NVIDIA's L4T packages frozen, a first publisher and listener talking, the extra ROS packages later chapters need, and a start-up wait that keeps every ROS process on fast shared memory.
ROS 2 is the software framework every robot program in this guide is built on: the base driver, the LiDAR reader,
navigation, the mind. This chapter installs ROS 2 Humble, the release made for Ubuntu 22.04, which is what NVIDIA's
L4T R36.4.3 runs. Only the base packages go on the robot (ros-base, no desktop tools), because the robot has no
screen; graphical tools run on another computer. Before anything is installed, every NVIDIA L4T package is frozen
at its current version, so that an Ubuntu upgrade cannot replace the kernel or the bootloader. At the end you fix a
start-up race in the middleware that once loaded the robot to a load average of 30.
Done on this robot: the install on 2026-09-26 (third attempt; the first two failed on the freeze line, fixed in the
script you see here), the reboot and the first publish/listen test the same evening, the extra packages on
2026-09-27 to 09-29, and the start-up wait on 2026-10-05.
New idea: ROS 2 in one paragraph
A ROS 2 program is a node: one process (or part of one) with a name, such asrosorin_board_driveror
rosorin_lidar_reader. Nodes talk by publishing messages on named topics: the LiDAR reader publishes a
sensor_msgs/msg/LaserScanmessage on/scanten times a second, and any node that subscribes to/scan
receives it. Neither side knows the other; they only agree on the topic name and the message type. Chapter 2
explains services, actions, parameters and launch files.
New idea: DDS, discovery and the domain ID
ROS 2 does not move messages itself. It hands them to a middleware called DDS; on Humble the default is eProsima
Fast DDS (version 2.6 on this robot). DDS finds the other side by itself: every process announces its
publishers and subscribers on the network, and processes with matching topics connect. This is called
discovery. The domain ID (environment variableROS_DOMAIN_ID, a number, 0 if unset) splits one network
into separate ROS worlds: processes only discover others with the same number. On one machine DDS sends data
through shared memory when it can and through UDP network packets otherwise.
New idea: sourcing a ROS setup file
ROS 2 installs into/opt/ros/humble, not into/usr. A shell does not find theros2command or the ROS
Python packages until you runsource /opt/ros/humble/setup.bash, which adds them toPATH,PYTHONPATHand
several ROS variables (for exampleROS_DISTRO=humble). Your own packages, built later in a workspace such as
~/ros2_ws, get their owninstall/setup.bash; sourcing it after the system one lays your packages over the
system ones (an overlay). The order matters: whatever is sourced last wins.
The robot's apt sources include NVIDIA's L4T repository for the r36.4 channel. On 2026-09-26 that channel already
offered 36.4.4 and 36.4.7 next to the installed 36.4.3:
Anywhere:
curl -sfL https://repo.download.nvidia.com/jetson/t234/dists/r36.4/main/binary-arm64/Packages.gz | gunzip | awk '/^Package: /{p=$2} /^Version: /{if (p ~ /^nvidia-l4t-(bootloader|kernel|core|init)$/) print p, $2}' | sort -u
Check
On 2026-09-26 the list includednvidia-l4t-bootloader 36.4.3-20250107174145(installed) and
nvidia-l4t-bootloader 36.4.7-20250918154033, and the same forcore,initandkernel.
A normal apt-get upgrade would install 36.4.7: a new kernel (every module from chapter 7 would stop loading) and a
bootloader package that triggers an update of the module's QSPI flash. The decision of 2026-09-26: hold every
installed nvidia-l4t-* package at 36.4.3; an L4T update is a separate, later step with its own golden set.
Never install nvidia-jetpack
Thenvidia-jetpackmeta-package looks like the obvious way to get CUDA and TensorRT. On this robot it would
addnvidia-l4t-*36.4.7 packages and replace the system OpenCV 4.5.4 that ROS 2 Humble uses, removing 19
packages (scripts/install_vision.shheader, 2026-09-29). Chapter 20 installs CUDA, cuDNN and TensorRT one
package at a time instead.dpkg-query -W nvidia-jetpackmust keep saying
no packages found matching nvidia-jetpack.
scripts/install_ros2.sh follows the official docs.ros.org procedure for Humble ("Ubuntu (deb packages)") and adds
the freeze. The locale (LANG=en_US.UTF-8) and the Ubuntu universe repository that the procedure asks for were
already set on the stock rootfs. 36 lines, in five parts.
On the robot:
#!/bin/bash
# ROS 2 Humble (ros-base + dev tools) per docs.ros.org Humble "Ubuntu (deb packages)".
# L4T packages held at the installed version (golden image baseline). Run: sudo bash ~/setup/install_ros2.sh
set -euo pipefail
[ "$(id -u)" = 0 ] || { echo "run with sudo"; exit 1; }
export DEBIAN_FRONTEND=noninteractive
L4T_BEFORE=$(head -1 /etc/nv_tegra_release)
DEBIAN_FRONTEND=noninteractive stops package scripts from asking questions. /etc/nv_tegra_release holds the L4T
release line; the script compares it at the end.
On the robot:
echo "== hold L4T packages"
# only installed packages (dpkg also lists known-but-not-installed ones, e.g. nvidia-l4t-ccp-generic)
apt-mark hold $(dpkg-query -W -f='${db:Status-Abbrev} ${Package}\n' 'nvidia-l4t-*' | awk '$1 ~ /^[ih]i$/ {print $2}') >/dev/null
echo "held: $(apt-mark showhold | grep -c '^nvidia-l4t-')"
dpkg-query lists every nvidia-l4t-* package dpkg has ever heard of with a status code in front: ii installed,
hi installed and held, un known but not installed. awk keeps only ii and hi, and apt-mark hold freezes
them. A held package is never upgraded, removed or replaced by apt unless you unhold it.
On the robot:
echo "== apt update + upgrade (Ubuntu packages; L4T held)"
apt-get update -q
apt-get -y -q upgrade
With the holds in place this upgrades only Ubuntu's own packages.
On the robot:
echo "== ros2-apt-source"
apt-get install -y -q curl
V=$(curl -s https://api.github.com/repos/ros-infrastructure/ros-apt-source/releases/latest | grep -F '"tag_name"' | awk -F'"' '{print $4}')
[ -n "$V" ] || { echo "could not read ros-apt-source version"; exit 1; }
CODENAME=$(. /etc/os-release && echo "${UBUNTU_CODENAME:-${VERSION_CODENAME}}")
curl -fL -o /tmp/ros2-apt-source.deb "https://github.com/ros-infrastructure/ros-apt-source/releases/download/${V}/ros2-apt-source_${V}.${CODENAME}_all.deb"
dpkg -i /tmp/ros2-apt-source.deb
The ROS project ships its repository address and signing key as a small package, ros2-apt-source. The script asks
GitHub for the newest release number, builds the file name for Ubuntu jammy, downloads it and installs it. That
creates /etc/apt/sources.list.d/ros2.sources (pointing at packages.ros.org).
On the robot:
echo "== install ros-humble-ros-base + ros-dev-tools"
apt-get update -q
apt-get install -y -q ros-humble-ros-base ros-dev-tools
echo "== verify"
dpkg-query -W -f='${Package} ${Version}\n' ros-humble-ros-base ros-dev-tools ros2-apt-source
L4T_AFTER=$(head -1 /etc/nv_tegra_release)
[ "$L4T_BEFORE" = "$L4T_AFTER" ] && echo "L4T unchanged: $L4T_AFTER" || echo "WARNING L4T CHANGED: $L4T_AFTER"
bash -c 'source /opt/ros/humble/setup.bash && ros2 --help >/dev/null && echo "ros2 CLI OK" && printenv ROS_DISTRO'
systemctl --failed --no-legend
echo "ROS2_DONE"
ros-humble-ros-base is the core of ROS 2 without graphical tools. ros-dev-tools adds the build tools (colcon,
rosdep) used from chapter 11 on. The verify block prints the versions, proves the L4T release line did not
change, runs the ros2 command once in a fresh shell, and lists failed system services (it should list none).
~/setup/install_ros2.sh:
On the robot:
#!/bin/bash
# ROS 2 Humble (ros-base + dev tools) per docs.ros.org Humble "Ubuntu (deb packages)".
# L4T packages held at the installed version (golden image baseline). Run: sudo bash ~/setup/install_ros2.sh
set -euo pipefail
[ "$(id -u)" = 0 ] || { echo "run with sudo"; exit 1; }
export DEBIAN_FRONTEND=noninteractive
L4T_BEFORE=$(head -1 /etc/nv_tegra_release)
echo "== hold L4T packages"
# only installed packages (dpkg also lists known-but-not-installed ones, e.g. nvidia-l4t-ccp-generic)
apt-mark hold $(dpkg-query -W -f='${db:Status-Abbrev} ${Package}\n' 'nvidia-l4t-*' | awk '$1 ~ /^[ih]i$/ {print $2}') >/dev/null
echo "held: $(apt-mark showhold | grep -c '^nvidia-l4t-')"
echo "== apt update + upgrade (Ubuntu packages; L4T held)"
apt-get update -q
apt-get -y -q upgrade
echo "== ros2-apt-source"
apt-get install -y -q curl
V=$(curl -s https://api.github.com/repos/ros-infrastructure/ros-apt-source/releases/latest | grep -F '"tag_name"' | awk -F'"' '{print $4}')
[ -n "$V" ] || { echo "could not read ros-apt-source version"; exit 1; }
CODENAME=$(. /etc/os-release && echo "${UBUNTU_CODENAME:-${VERSION_CODENAME}}")
curl -fL -o /tmp/ros2-apt-source.deb "https://github.com/ros-infrastructure/ros-apt-source/releases/download/${V}/ros2-apt-source_${V}.${CODENAME}_all.deb"
dpkg -i /tmp/ros2-apt-source.deb
echo "== install ros-humble-ros-base + ros-dev-tools"
apt-get update -q
apt-get install -y -q ros-humble-ros-base ros-dev-tools
echo "== verify"
dpkg-query -W -f='${Package} ${Version}\n' ros-humble-ros-base ros-dev-tools ros2-apt-source
L4T_AFTER=$(head -1 /etc/nv_tegra_release)
[ "$L4T_BEFORE" = "$L4T_AFTER" ] && echo "L4T unchanged: $L4T_AFTER" || echo "WARNING L4T CHANGED: $L4T_AFTER"
bash -c 'source /opt/ros/humble/setup.bash && ros2 --help >/dev/null && echo "ros2 CLI OK" && printenv ROS_DISTRO'
systemctl --failed --no-legend
echo "ROS2_DONE"
On your laptop:
scp -q ~/CCode/rosorin-pro/scripts/install_ros2.sh rosorin:setup/
On the robot:
sudo bash ~/setup/install_ros2.sh
On 2026-09-26 the successful run started after 21:25 UTC and was finished when the robot was checked at 21:34.
Check
The script's own output was not captured on this robot (the owner ran it). Below are its final lines as the
script writes them, filled in with the values read on the robot right after the run (holds, L4T line) and on
2026-10-07 (exact package versions):held: 45 ros-humble-ros-base 0.10.0-1jammy.20260910.003535 ros-dev-tools 1.0.3 ros2-apt-source 1.3.0~jammy L4T unchanged: # R36 (release), REVISION: 4.3, GCID: 38968081, BOARD: generic, EABI: aarch64, DATE: Wed Jan 8 01:49:37 UTC 2025 ros2 CLI OK humble ROS2_DONEwith nothing listed between
humbleandROS2_DONE(no failed units). The robot holds 46 L4T packages today:
chapter 20 adds and holdsnvidia-l4t-dla-compiler.
If it fails
- The first two runs on 2026-09-26 (21:23:05 and 21:23:58 UTC) stopped at the hold line and installed no ROS
package. The script's first version held every namedpkg-query -W 'nvidia-l4t-*'printed, including
nvidia-l4t-ccp-generic, which is known to dpkg but not installed (un);apt-mark holdfails on such a
name, andset -eended the script. Theawk '$1 ~ /^[ih]i$/'filter above is the fix
(docs/lessons.md 2026-09-26).could not read ros-apt-source version: the GitHub API answered with nothing (no network, or GitHub's rate
limit for anonymous requests). Run it again later.WARNING L4T CHANGED: something upgraded L4T. Stop and restore from your golden set (chapter 8).
Version drift
The script installs whateverros2-apt-sourcerelease is newest on the day you run it (1.3.0 on 2026-09-26),
and apt installs the newest Humble packages. Your versions can be newer than this robot's; that was not
tested.
The upgrade replaced libraries that running programs still use.
On the robot:
test -f /var/run/reboot-required && { echo "REBOOT REQUIRED by:"; cat /var/run/reboot-required.pkgs 2>/dev/null; } || echo "no reboot required"
sudo systemctl reboot
Check
Before the reboot on 2026-09-26:REBOOT REQUIRED by: libc6 network-manager evolution-data-server gnome-shellAfter it,
test -f /var/run/reboot-required && echo "reboot still required" || echo "no reboot pending"printed
no reboot pending, the root was still/dev/nvme0n1p16, the kernel5.15.148-tegra, and
head -1 /etc/nv_tegra_releasestillREVISION: 4.3.
ros-base has no demo programs (ros-humble-demo-nodes-cpp is not installed on the robot), so use the ros2
command line itself as a publisher and as a listener. Open two ssh sessions to the robot.
In the first, publish the text hello five times a second on a topic called /probe:
On the robot:
source /opt/ros/humble/setup.bash
ros2 topic pub -r 5 /probe std_msgs/msg/String "{data: hello}"
std_msgs/msg/String is the message type: one field, data, holding text. "{data: hello}" fills it in YAML.
In the second, wait for one message and print it:
On the robot:
source /opt/ros/humble/setup.bash
ros2 topic echo --once /probe
Check
The listener prints, and exits:data: hello ---This is exactly what the robot printed on 2026-09-26, before and after the reboot. Stop the publisher with
Ctrl-C. Theros2command starts a small background helper (the ROS 2 daemon) that remembers what it has
discovered; stop it when you are done withros2 daemon stop.
The same test as one command, the way it was run on the robot:
On the robot:
source /opt/ros/humble/setup.bash
ros2 topic pub -r 5 /probe std_msgs/msg/String "{data: hello}" >/dev/null 2>&1 & P=$!
sleep 4; timeout 8 ros2 topic echo --once /probe; kill $P 2>/dev/null
ros2 daemon stop >/dev/null 2>&1
If it fails
- Nothing arrives, or only sometimes, and it gets worse after you log out of ssh: systemd-logind is deleting
your user's shared-memory files when your last session closes. Chapter 6 setsRemoveIPC=no
(docs/lessons.md 2026-09-28).- A publisher that sends once right after it starts can lose that message: the listener has not been discovered
yet. That is why the test publishes repeatedly (-r 5) and the listener waits up to 8 s
(docs/lessons.md, factory stack).- Back-to-back
ros2 topic pubandros2 service callcommands in one shell sometimes fail for no visible
reason. One empty result proves nothing; run it again (docs/lessons.md).- A background job started inside
ssh robot '... &'ignores Ctrl-C andkill -INT. Stop it with plain
kill(SIGTERM), as above (docs/lessons.md 2026-09-26).pkill -f probeinside an ssh command can match and kill your own ssh session, because its command line
contains the word. Kill by process number (docs/lessons.md 2026-09-26, 2026-09-30).
The robot's ~/.bashrc is the stock Ubuntu one: it does not source ROS (checked 2026-10-07). Every command in
this guide sources the setup files explicitly, and so does every service. The base service, for example
(ros2/rosorin_base/systemd/rosorin-base.service, chapter 11):
On the robot:
ExecStart=/bin/bash -c 'source /opt/ros/humble/setup.bash && source /home/burgerbarn/ext_ws/install/setup.bash && source /home/burgerbarn/ros2_ws/install/setup.bash && exec ros2 launch rosorin_base base.launch.py'
Three layers in a fixed order: the system install, ~/ext_ws (packages built from other people's source, chapter
16), then ~/ros2_ws (this project's packages, chapter 11). A package in a later layer hides the same package in an
earlier one; chapter 18 uses that to replace one Nav2 package with a fixed build.
Every robot service sets the domain explicitly, with the same line in each unit file
(rosorin-base, rosorin-nav, rosorin-camera, rosorin-mind, rosorin-api, rosorin-vslam,
rosorin-contact, rosorin-selfcheck, rosorin-explore, rosorin-practice, rosorin-task@):
On the robot:
Environment=ROS_DOMAIN_ID=0
In your ssh shell ROS_DOMAIN_ID is unset, which also means domain 0, so your commands see the services. The
project's isolated test scripts (scripts/tests/*.sh) use the opposite: they run in domain 77 with
ROS_LOCALHOST_ONLY=1, so nothing they start can reach the robot's driver. The robot talks to bigbuddy over HTTP,
never over DDS (scripts/install_dds_localhost_only.sh header).
Check
In a fresh ssh shell on the robot:$ source /opt/ros/humble/setup.bash; printenv ROS_DISTRO; echo "ROS_DOMAIN_ID=${ROS_DOMAIN_ID:-unset}" humble ROS_DOMAIN_ID=unsetAfter chapter 11,
systemctl show -p Environment rosorin-baseshowsROS_DOMAIN_ID=0.
Domain 0 is shared with your whole network
The robot's DDS is not limited to the robot:ROS_LOCALHOST_ONLYis not set on it
(scripts/install_dds_localhost_only.shis marked NOT APPLIED). Any other computer on your network running
ROS 2 in domain 0 discovers the robot's topics and can publish to them, including the topics that drive the
wheels. Chapter 12 uses this on purpose to view the robot from a desktop. Keep other ROS machines on another
domain unless you mean to talk to the robot. The factory image also had DDS open on the LAN
(docs/hardware.md).
These were installed from the same ROS repository (and NVIDIA's, for ffmpeg) between 2026-09-27 and 2026-09-29,
each when a chapter first needed it. They do not depend on each other's order; install them now so later chapters
only build and configure. The ROS ones were installed with --no-install-recommends, which tells apt to install
only what a package requires, not the extra packages it merely recommends. The record does not say why; it keeps
the robot to what is used.
On the robot:
sudo DEBIAN_FRONTEND=noninteractive apt-get install -y -q ffmpeg v4l-utils
sudo apt-get install -y --no-install-recommends ros-humble-robot-localization
sudo apt-get install -y --no-install-recommends ros-humble-slam-toolbox
sudo apt-get install -y --no-install-recommends ros-humble-navigation2 ros-humble-nav2-bringup
sudo apt-get install -y ros-humble-web-video-server
apt-mark showhold | grep -c nvidia-l4t
| Date (UTC) | Package | New packages | Used for |
|---|---|---|---|
| 2026-09-27 15:58 | ffmpeg 7:4.4.2-nvidia (NVIDIA's build), v4l-utils 1.22.1 |
a USB test webcam beside the robot (rosorin-cam, installed but disabled); nothing else in the repo uses them |
|
| 2026-09-28 00:09 | ros-humble-robot-localization 3.5.4 |
6 | the EKF that fuses wheel odometry and the IMU (chapter 16) |
| 2026-09-28 06:09 | ros-humble-slam-toolbox 2.6.10 |
203 (includes Qt libraries) | mapping and localization (chapter 17) |
| 2026-09-28 07:00 | ros-humble-navigation2, ros-humble-nav2-bringup 1.1.20 |
74 | navigation (chapter 18) |
| 2026-09-29 23:06 | ros-humble-web-video-server 3.1.0 |
viewing camera topics in a browser on port 8080; started by hand only, nothing starts it at boot |
Check
After each install on this robot,apt-mark showhold | grep -c nvidia-l4tstill printed45, and
head -1 /etc/nv_tegra_releasestill saidREVISION: 4.3. slam_toolbox's install printed
0 upgraded, 203 newly installed, 0 to remove and 50 not upgraded., Nav2's
0 upgraded, 74 newly installed, 0 to remove and 50 not upgraded.Read on the robot on 2026-10-07:$ dpkg-query -W -f='${Package} ${Version}\n' ros-humble-navigation2 ros-humble-nav2-bringup ros-humble-slam-toolbox ros-humble-robot-localization ros-humble-web-video-server ffmpeg v4l-utils ffmpeg 7:4.4.2-nvidia ros-humble-nav2-bringup 1.1.20-1jammy.20260910.014042 ros-humble-navigation2 1.1.20-1jammy.20260910.013941 ros-humble-robot-localization 3.5.4-1jammy.20260908.051527 ros-humble-slam-toolbox 2.6.10-1jammy.20260909.214512 ros-humble-web-video-server 3.1.0-1jammy.20260908.050220 v4l-utils 1.22.1-2build1The web server answers
<html><head><title>ROS Streamable Topic List</title>on
http://localhost:8080/once you start it with
ros2 run web_video_server web_video_server --ros-args -p port:=8080 -p address:=0.0.0.0.
Installed elsewhere in this guide, by their own scripts: MoveIt (scripts/install_moveit.sh, chapter 13), Isaac
ROS nvblox and cuVSLAM (scripts/install_isaac_ros.sh, chapters 16 and 18), CUDA, cuDNN and TensorRT
(scripts/install_vision.sh, chapter 20).
Not test-built as one block
On this robot the five installs ran one at a time over three days, in the order of the table, each with output
filters. Running them back to back as above has not been done. The Nav2 overlay in chapter 18
(install_nav2_behaviors_fix.sh) refuses to run unless the installednav2-behaviorsis exactly 1.1.20; a
newer apt release would need that fix checked again.
If it fails
- Any of these pulls an
nvidia-l4t-*upgrade or wants to remove packages: stop and read apt's list before
you answer. With the holds in place none did on this robot.- Never run
apt autoremoveon the robot. After a removal apt offers to removeinitramfs-tools, which L4T
needs to build the boot initrd (docs/lessons.md 2026-09-26).
This step fixes the worst performance fault the robot has had. It needs the base service from chapter 11; read it
now, put the script in place, and run the installer once rosorin-base.service exists.
New idea: shared memory, and Fast DDS's host id
Two DDS processes on the same computer can exchange messages through shared memory: one writes into a block of
RAM (files under/dev/shm), the other reads it. That is fast and costs almost no CPU. Processes on different
computers must use UDP packets. To decide which case applies, Fast DDS 2.6 gives each process a host id: an
MD5 hash of the computer's IPv4 addresses at the moment the process starts (Fast DDS source
src/cpp/utils/Host.hpp). Two processes with the same host id use shared memory; with different ids they
believe they are on different computers and send UDP packets to each other over the loopback interface. For
topics with guaranteed delivery ("reliable"), each side also sends acknowledgements and re-sends, and on a
busy robot that becomes a storm. The host id is visible in every endpoint's GID: the 3rd and 4th byte
(01.0f.f4.a0...has host bytesf4.a0).
What happened on 2026-10-05: with navigation always on, the robot ran at a load average of 30, all eight cores at
85-95 %, 30 % of CPU time inside the kernel. The base service had started 4 s before the Wi-Fi had its address. Its
nodes (EKF, robot_state_publisher, rf2o, LiDAR reader) and two services started right after it got host id 7f.01;
everything started later (navigation, the mind, visual odometry, the camera after a restart) got f4.a0. Between
the two groups every transform, scan and some camera images went over UDP on loopback: 65,000 packets per second
(56,000 of them acknowledgements), 5,900 kernel receive-buffer overflows per second, and robot_state_publisher
using 37 % of a core for a 10 Hz job.
The unit file already said After=network-online.target, but that target waits for nothing on this robot:
NetworkManager-wait-online.service is disabled (systemctl is-enabled prints disabled). Restarting the early
services so that all processes shared one id brought loopback traffic from 65,000 to 330 packets per second,
overflows to 0, and kernel CPU time from 220 % to 85 % of a core. The lasting fix is to make the base service, the
first ROS process at boot, wait until the address list has stopped changing.
scripts/wait_addresses.sh, 14 lines:
On the robot:
#!/bin/bash
# Runs before the first ROS process (rosorin-base ExecStartPre). Fast DDS 2.6 gives each process a host id = MD5 of the
# machine's IPv4 addresses AT PROCESS START (src/cpp/utils/Host.hpp); processes with different ids do not use shared
# memory with each other and fall back to UDP on loopback. 2026-10-05: base started 4 s before Wi-Fi had its address
# -> two groups -> 65,000 packets/s, 5,900 kernel receive-buffer overflows/s, load average 30.
# Waits until the address list has been the same for 6 s (at most 45 s; the robot must also start without Wi-Fi).
The header is the whole story in five lines.
On the robot:
prev=""; same=0
for i in $(seq 1 45); do
cur=$(ip -4 -o addr show | awk '$2!="lo"{print $2"="$4}' | sort | paste -sd,)
Once a second, for at most 45 seconds, build one line from every IPv4 address except loopback, as
interface=address pairs sorted and joined by commas, for example docker0=172.17.0.1/16,wlP1p1s0=192.168.1.108/24.
On the robot:
if [ -n "$cur" ] && [ "$cur" = "$prev" ]; then same=$((same+1)); else same=0; fi
[ $same -ge 6 ] && { echo "addresses stable after ${i}s: $cur"; exit 0; }
prev=$cur; sleep 1
done
echo "addresses not stable after 45 s (now: ${cur:-none}) - starting anyway"; exit 0
Count how many seconds in a row the line has not changed. Six unchanged seconds means the network has settled:
print the list and let the service start. After 45 s it gives up and starts anyway, because the robot must also work
with no Wi-Fi at all. It always exits 0, so it can delay the service but never stop it.
The complete file, ~/setup/wait_addresses.sh:
On the robot:
#!/bin/bash
# Runs before the first ROS process (rosorin-base ExecStartPre). Fast DDS 2.6 gives each process a host id = MD5 of the
# machine's IPv4 addresses AT PROCESS START (src/cpp/utils/Host.hpp); processes with different ids do not use shared
# memory with each other and fall back to UDP on loopback. 2026-10-05: base started 4 s before Wi-Fi had its address
# -> two groups -> 65,000 packets/s, 5,900 kernel receive-buffer overflows/s, load average 30.
# Waits until the address list has been the same for 6 s (at most 45 s; the robot must also start without Wi-Fi).
prev=""; same=0
for i in $(seq 1 45); do
cur=$(ip -4 -o addr show | awk '$2!="lo"{print $2"="$4}' | sort | paste -sd,)
if [ -n "$cur" ] && [ "$cur" = "$prev" ]; then same=$((same+1)); else same=0; fi
[ $same -ge 6 ] && { echo "addresses stable after ${i}s: $cur"; exit 0; }
prev=$cur; sleep 1
done
echo "addresses not stable after 45 s (now: ${cur:-none}) - starting anyway"; exit 0
scripts/install_dds_hostid_fix.sh adds a systemd drop-in: a small file in
/etc/systemd/system/rosorin-base.service.d/ whose lines are added to the unit without editing the unit file.
On the robot:
#!/bin/bash
# rosorin-base waits for a stable IPv4 address list before any ROS process starts (scripts/wait_addresses.sh).
# Run on the robot: sudo bash ~/setup/install_dds_hostid_fix.sh
# Rollback: sudo rm -r /etc/systemd/system/rosorin-base.service.d/10-wait-addresses.conf; sudo systemctl daemon-reload
set -euo pipefail
[ "$(id -u)" = 0 ] || { echo "run with sudo"; exit 1; }
mkdir -p /etc/systemd/system/rosorin-base.service.d
cat > /etc/systemd/system/rosorin-base.service.d/10-wait-addresses.conf <<'UNIT'
[Service]
ExecStartPre=/bin/bash /home/burgerbarn/setup/wait_addresses.sh
TimeoutStartSec=120
UNIT
systemctl daemon-reload
systemctl cat rosorin-base | grep -A3 "10-wait-addresses"
ExecStartPre runs the wait before the service's own command. TimeoutStartSec=120 gives the start enough time to
include up to 45 s of waiting. daemon-reload makes systemd read the new file, and the last line shows it merged into
the unit.
Copy both files now; run the installer after chapter 11 has installed rosorin-base.service:
On your laptop:
scp -q ~/CCode/rosorin-pro/scripts/wait_addresses.sh ~/CCode/rosorin-pro/scripts/install_dds_hostid_fix.sh rosorin:setup/
On the robot:
bash ~/setup/wait_addresses.sh
sudo bash ~/setup/install_dds_hostid_fix.sh
Check
On 2026-10-05 20:14 UTC the two commands printed (installer output first, as run then):[Service] ExecStartPre=/bin/bash /home/burgerbarn/setup/wait_addresses.sh TimeoutStartSec=120 addresses stable after 7s: docker0=172.17.0.1/16,wlP1p1s0=192.168.1.108/24After the next power-on (2026-10-05 22:59 UTC):
$ journalctl -u rosorin-base -b --no-pager -o cat | grep "addresses stable" addresses stable after 11s: docker0=172.17.0.1/16,wlP1p1s0=192.168.1.108/24and on the boot of 2026-10-07,
addresses stable after 12s. Your list shows the interfaces your robot has at
that moment (docker0exists only once Docker is installed;enP8p1s0appears when the Ethernet cable is in).
Once the base and the other services run, check that every ROS process has the same host id. This lists, for a few
busy topics, the first four GID bytes of each endpoint and the nodes that have them:
On the robot:
source /opt/ros/humble/setup.bash; for t in /tf /joint_states /aurora/rgb/image_raw /scan /imu/data_raw; do ros2 topic info -v $t 2>/dev/null | awk "/^Node name:/{n=\$3} /^GID/{print substr(\$2,1,11), n}"; done | sort -u | awk "{h[\$1]=h[\$1]\" \"\$2} END{for(k in h) print k, \":\", h[k]}"
Check
One line means one host id: correct. On 2026-10-05 20:06 UTC, after the restart that fixed it:01.0f.f4.a0 : aurora contact_monitor ekf_filter_node look_around_scan rf2o_laser_odometry robot_api robot_state_publisher rosorin_board_driver rosorin_lidar_reader transform_listener_impl_...Two lines are the fault. Before the restart the same check gave one line starting
01.0f.f4.a0 :(navigation,
the mind, visual odometry, the camera) and a second starting01.0f.7f.01 :(the EKF and the other early
nodes).
If it fails
- Two host ids after a boot: some process started before the addresses settled. Restart the processes with the
old id (on 2026-10-05:rosorin-base, thenrosorin-contactandrosorin-api) and check again. The robot's
self-care (chapter 23) checks for this every 15 minutes and restarts processes with a stale id.- The same split can happen later without a reboot: if the Wi-Fi drops and comes back with a different address
list, processes started after that get a new id. The self-care rule is the cure; it has not been tested on a
real split (docs/status.md 2026-10-05).After=network-online.targetalone does not help on this robot, becauseNetworkManager-wait-onlineis
disabled (docs/lessons.md 2026-10-05).- Fresh nodes see no topics at all, or only after a delay, and it changes when ssh sessions close: that is the
logindRemoveIPCfault from chapter 6, a different cause. SettingROS_LOCALHOST_ONLY=1did not help
there and was reverted.- Not applied, by design:
scripts/install_dds_localhost_only.sh(marked NOT APPLIED, 2026-10-06) and
ros2/rosorin_base/config/fastdds_shm.xml(tested on an isolated domain, not deployed).
Rollback, from the installer's header:
sudo rm -r /etc/systemd/system/rosorin-base.service.d/10-wait-addresses.conf; sudo systemctl daemon-reload.
Order on a fresh build
On this robot the drop-in was added on 2026-10-05, to a base service that had run for nine days. Installing it
right after chapter 11, as this guide does, is the same file on the same unit, but that order has not been
run.
apt-mark showhold | grep -c nvidia-l4t prints 45 (46 after chapter 20), andhead -1 /etc/nv_tegra_release still says REVISION: 4.3.dpkg-query -W nvidia-jetpack says no packages found matching nvidia-jetpack.data: hello and ---.dpkg-query -W ros-humble-navigation2 ros-humble-slam-toolbox ros-humble-robot-localization lists versions~/setup/wait_addresses.sh exists on the robot, and after chapter 11 the base service's journal showsaddresses stable after ... at every boot.Where this comes from
scripts/install_ros2.sh(commit 6a0f55b, 2026-09-26),scripts/wait_addresses.shand
scripts/install_dds_hostid_fix.sh(9ed6e30, 2026-10-05),scripts/install_dds_localhost_only.shheader,
scripts/install_vision.shheader (nvidia-jetpack),ros2/rosorin_base/systemd/rosorin-base.service,
systemd/*.service(ROS_DOMAIN_ID=0),scripts/tests/*.sh(domain 77). Docs:docs/decisions.md
2026-09-26 (L4T hold, ros-base only),docs/lessons.md2026-09-26 (dpkgunpackages, autoremove, SIGINT,
pkill), 2026-09-28 (RemoveIPC), 2026-10-05 (host id),docs/status.md2026-09-26 (ROS 2 install) and
2026-10-05 (the overload and its fix),docs/odometry.md,docs/navigation.md,docs/hardware.md(ffmpeg).
Command log (UTC): NVIDIA repo versions 2026-09-26 21:22, failed runs 21:23, hold fix 21:24:56, pub/sub test
21:34 and after reboot 21:36; ffmpeg 2026-09-27 15:58; robot-localization 2026-09-28 00:09, slam-toolbox 06:09,
Nav2 07:00; web-video-server 2026-09-29 23:06; host-id diagnosis and fix 2026-10-05 20:05-20:14, first boot
with the wait 22:59. sources/gaps_from_cmdlog.md sections 4 and 5. Live read-back 2026-10-07: package versions,
46 holds, no nvidia-jetpack, no demo nodes,~/.bashrcwithout ROS,NetworkManager-wait-onlinedisabled,
boot lineaddresses stable after 12s.
← Make a backup you can restore · Contents · Talk to the controller board →