The robot's own mind · Chapter 20 · Time: 3 hours, about 30 minutes of it waiting for downloads and engine builds · Level: Intermediate · Status: Partly test-built
NVIDIA's CUDA, cuDNN and TensorRT installed from the robot's own L4T apt repo without disturbing ROS 2, a Python environment in ~/vision, two NanoOWL TensorRT engines, and your first open-vocabulary detection on one camera frame, with the measured speed of each detector.
The robot finds Matt with its camera. To follow a person who moves, it has to run a neural network on every camera
frame, about 14 times a second, and only the Jetson's GPU is fast enough for that. This chapter installs NVIDIA's GPU
libraries in a way that leaves the ROS 2 install from chapter 9 untouched, builds the detector the robot uses, and
runs it by hand on one frame so you can see what it reports. Chapter 21 puts the same detector inside the robot's
mind.
Every step below was run on this robot by hand on 2026-09-29 (face recognition on 2026-09-30), in this order, with
the failures shown in the "If it fails" boxes. scripts/install_vision.sh was written from those steps afterwards
and has never been run as one piece. That is why the badge says partly test-built.
New idea: GPU, CUDA and cuDNN
A neural network is mostly multiplying large tables of numbers. A CPU does that a few numbers at a time; a GPU
does thousands at once. The Jetson Orin NX has its GPU on the same chip as the CPU, and the two share the same
16 GB of RAM: memory the GPU takes is memory the robot's other processes lose (chapter 1 tells what happened when
a big model took too much). CUDA is NVIDIA's way for programs to run code on the GPU; it comes as libraries
plus a compiler (nvcc). cuDNN is a CUDA library of ready-made neural-network building blocks
(convolutions, attention). PyTorch uses both.
New idea: inference, fp16 and TensorRT engines
Running a trained network on new input is called inference. TensorRT is NVIDIA's inference optimiser:
you give it a trained network (here as an ONNX file, a standard network format) and it measures which way of
computing each layer is fastest on this GPU, then saves the result as an engine file. The engines here use
fp16, 16-bit floating point numbers instead of 32-bit: half the memory traffic, nearly the same answers. An
engine is tied to the GPU and the TensorRT version it was built with; if either changes, build it again.
New idea: open-vocabulary detection
An ordinary object detector knows a fixed list of classes it was trained on. An open-vocabulary detector takes
the class names as text at run time: you write"a person","a guitar","a capo", and it scores every
candidate box in the image against every phrase. The phrases are called prompts. The robot uses OWL-ViT (a
Google model) through NVIDIA's NanoOWL, which runs OWL-ViT's image half as a TensorRT engine and keeps the text
half in PyTorch. The image is resized to 768x768 and cut into square patches: patch32 cuts 32x32-pixel
squares (24x24 = 576 patches, faster), patch16 cuts 16x16 squares (48x48 = 2304 patches, finer, slower).
OWLv2 is the newer, heavier model of the same family; it runs in plain PyTorch here and is better on small
things. Scores are between 0 and 1 but real hits are often only 0.15-0.6, so each prompt gets its own threshold.
New idea: non-maximum suppression (NMS)
A detector reports many overlapping boxes for one object. NMS keeps the best-scoring box and drops every other
box of the same label that overlaps it by more than a set fraction. The overlap measure is IoU, intersection
over union: the area two boxes share divided by the area they cover together (1.0 = identical, 0 = apart).
nvidia-l4t-* packages held. Check: apt-mark showholdrosorin-camera publishes /aurora/rgb/image_raw (needed only for the experiment at the end).20260929T1441Z.Never install nvidia-jetpack, never autoremove
The two apt mistakes that can break this robot both sit close to this chapter. Do not install the
nvidia-jetpackmeta-package (next section), and never runapt autoremove: after a removal apt offers to
removeinitramfs-tools, which the Jetson needs to boot (docs/lessons.md, 2026-09-26). apt prints
"Use 'apt autoremove' to remove them." during this chapter. Ignore it.
The install script and the vision code live in the repo on your laptop. Copy them over:
On your laptop:
cd ~/CCode/rosorin-pro
ssh rosorin-wifi 'mkdir -p ~/setup ~/vision'
scp scripts/install_vision.sh scripts/rollback_vision.sh rosorin-wifi:setup/
scp vision/*.py vision/README.md rosorin-wifi:vision/
Not test-built
On this robot the vision files arrived one or a few at a time withscpas they were written. Copying them all
at once, as above, gives the same result (the live~/vision/*.pyfiles are identical to the repo's), but this
exact command was not run.
NVIDIA's usual instruction is sudo apt install nvidia-jetpack, which pulls in CUDA, cuDNN, TensorRT and much more.
On this robot that would undo work from earlier chapters. Look before you install anything: apt-get -s simulates an
install and changes nothing.
On the robot:
apt-get -s install nvidia-jetpack 2>/dev/null | grep -E "^Inst" > /tmp/jp_sim.txt
wc -l < /tmp/jp_sim.txt
grep "nvidia-l4t" /tmp/jp_sim.txt | head
grep -E "\[[^]]+\]" /tmp/jp_sim.txt | head -5
What the simulation showed on 2026-09-29
122 packages to install, four of themnvidia-l4t-*at 36.4.7 while the robot runs 36.4.3, and an upgrade of
the OpenCV development package from 4.5.4:On the robot:
122 Inst nvidia-l4t-cudadebuggingsupport (12.6-34622040.0 L4T Jetson r36.4:stable [arm64]) Inst nvidia-l4t-jetson-multimedia-api (36.4.7-20250918154033 L4T Jetson r36.4:stable [arm64]) Inst nvidia-l4t-dla-compiler (36.4.7-20250918154033 L4T Jetson r36.4:stable [arm64]) Inst nvidia-l4t-gstreamer (36.4.7-20250918154033 L4T Jetson r36.4:stable [arm64]) Inst libopencv-dev [4.5.4+dfsg-9ubuntu4] (4.8.
Why only the pieces
ROS 2 Humble is built against the system OpenCV 4.5.4; the meta-package replaces it and, per the header of
scripts/install_vision.sh, would remove 19 packages. It also adds L4T 36.4.7 packages on top of a 36.4.3 system
whose L4T packages are held on purpose (chapter 9). So the robot installs only the packages vision needs, from the
same NVIDIA apt repo the stock L4T system already lists (repo.download.nvidia.com/jetson/{common,t234,ffmpeg} r36.4).
Simulate first. What you want to see is 0 upgraded and 0 to remove, and no nvidia-l4t line:
On the robot:
P="cuda-libraries-12-6 cuda-cudart-dev-12-6 cuda-nvcc-12-6 cuda-cupti-12-6 libcudnn9-cuda-12 tensorrt-libs libnvinfer-bin python3-libnvinfer libnvinfer-dev libnvonnxparsers-dev libnvinfer-plugin-dev python3-venv python3-pip"
apt-get -s install $P > /tmp/min_sim.txt 2>&1
grep -iE "upgraded|remove" /tmp/min_sim.txt | head -3
grep "^Remv" /tmp/min_sim.txt | head
grep "^Inst" /tmp/min_sim.txt | grep -E "l4t|opencv|\[" | head
Check
On 2026-09-29 the simulation of the CUDA/cuDNN/TensorRT part (the list withoutcuda-cupti-12-6,
python3-venvandpython3-pip) printed:On the robot:
Use 'apt autoremove' to remove them. 0 upgraded, 37 newly installed, 0 to remove and 56 not upgraded.With the three extra packages the number of new packages is larger.
0 upgradedand0 to removemust stay.
NoRemvlines, nonvidia-l4toropencvlines.
Then install. These are lines 11-13 of install_vision.sh:
On the robot:
sudo apt-get install -y cuda-libraries-12-6 cuda-cudart-dev-12-6 cuda-nvcc-12-6 cuda-cupti-12-6 libcudnn9-cuda-12 \
tensorrt-libs libnvinfer-bin python3-libnvinfer libnvinfer-dev libnvonnxparsers-dev libnvinfer-plugin-dev \
python3-venv python3-pip
| Package | What it gives you |
|---|---|
cuda-libraries-12-6 |
The CUDA 12.6 runtime libraries, in /usr/local/cuda-12.6/lib64 (56 .so entries on this robot). Live version 12.6.11-1. |
cuda-cudart-dev-12-6 |
CUDA runtime headers and link libraries, for building code against CUDA. |
cuda-nvcc-12-6 |
nvcc, the CUDA compiler. |
cuda-cupti-12-6 |
CUPTI, NVIDIA's profiling library. PyTorch refuses to import without libcupti.so.12 (see the fail box in step 5). |
libcudnn9-cuda-12 |
cuDNN 9. Live version 9.3.0.75-1. |
tensorrt-libs |
The TensorRT 10.3 runtime. Live version 10.3.0.30-1+cuda12.5. |
libnvinfer-bin |
trtexec at /usr/src/tensorrt/bin/trtexec. NanoOWL's engine builder calls this program. |
python3-libnvinfer |
import tensorrt for Python 3.10. |
libnvinfer-dev, libnvonnxparsers-dev, libnvinfer-plugin-dev |
TensorRT headers, its ONNX reader and its plugin library. |
python3-venv, python3-pip |
Python virtual environments and pip. |
On the robot:
ls /usr/local/cuda-12.6/lib64 | grep -c so; ls /usr/src/tensorrt/bin/trtexec
Check
On the robot:56 /usr/src/tensorrt/bin/trtexec
If it fails
python3 -c "import tensorrt"right after this step fails with
ImportError: libnvdla_compiler.so: cannot open shared object file: No such file or directory. That is expected
at this point: step 3 fixes it. Any other error from apt: stop and read it before retrying; never answer it with
apt autoremoveorapt --fix-broken installwithout simulating first.
TensorRT can also run networks on the Orin's DLA (deep-learning accelerator) blocks. The engines in this guide do not
use them, but TensorRT's Python module loads the DLA compiler library when it is imported, and that library comes in a
separate package that apt did not pull in. Look at the versions on offer:
On the robot:
apt-cache policy nvidia-l4t-dla-compiler | head -12; dpkg -l nvidia-l4t-core | tail -1 | awk '{print $2,$3}'
Check
On the robot:nvidia-l4t-dla-compiler: Installed: (none) Candidate: 36.4.7-20250918154033 Version table: 36.4.7-20250918154033 600 600 https://repo.download.nvidia.com/jetson/common r36.4/main arm64 Packages 36.4.4-20250616085344 600 600 https://repo.download.nvidia.com/jetson/common r36.4/main arm64 Packages 36.4.3-20250107174145 600 600 https://repo.download.nvidia.com/jetson/common r36.4/main arm64 Packages 36.4.0-20240912212859 600 600 https://repo.download.nvidia.com/jetson/common r36.4/main arm64 Packages nvidia-l4t-core 36.4.3-20250107174145
apt would pick 36.4.7. The robot's L4T core is 36.4.3, so install the version that matches it exactly and hold it so
a later apt upgrade cannot move it. Then refresh the dynamic linker's cache. These are lines 14-16 of the script:
On the robot:
sudo apt-get install -y nvidia-l4t-dla-compiler=36.4.3-20250107174145 && sudo apt-mark hold nvidia-l4t-dla-compiler
sudo ldconfig
python3 -c "import tensorrt; print('tensorrt', tensorrt.__version__)"
Check
apt saysnvidia-l4t-dla-compiler set on hold.and the import prints the TensorRT version
(the recorded run printed it astrt 10.3.0):On the robot:
tensorrt 10.3.0
dpkg -l nvidia-l4t-dla-compilerstarts withhi(installed, held). Live on the robot:
hi nvidia-l4t-dla-compiler 36.4.3-20250107174145.
If it fails
On 2026-09-29 the import still failed after the package was installed, with the samelibnvdla_compiler.so
error. The file was there (/usr/lib/aarch64-linux-gnu/nvidia/libnvdla_compiler.so, 8159168 bytes) but the
linker cache did not list it yet;ldconfig -p | grep nvdlashowed onlylibnvdla_runtime.so.sudo ldconfig
fixed it. That is why line 15 of the script exists.
The vision code gets its own Python environment (a venv) so that nothing pip installs can replace a system
package that ROS 2 depends on. It is created with --system-site-packages, so it can still see the system's
tensorrt module (from python3-libnvinfer) and the system OpenCV cv2 that ROS 2 uses. Lines 17-21:
On the robot:
IDX=https://pypi.jetson-ai-lab.io/jp6/cu126/+simple
mkdir -p ~/vision/src ~/vision/engines && cd ~/vision
[ -d venv ] || python3 -m venv --system-site-packages venv
. venv/bin/activate
pip install --upgrade pip setuptools wheel
On the robot:
head -3 ~/vision/venv/pyvenv.cfg
Check
On the robot:home = /usr/bin include-system-site-packages = true version = 3.10.12
The PyTorch wheels on the normal Python package index (PyPI) for this kind of ARM board run on the CPU only. NVIDIA's
Jetson AI Lab publishes wheels built for JetPack 6 and CUDA 12.6 at pypi.jetson-ai-lab.io/jp6/cu126. Line 22 of the
script installs everything from that index and nowhere else:
On the robot:
pip install --index-url $IDX "torch==2.8.0" "torchvision==0.23.0" "numpy<2" "transformers==4.46.3" "onnx<1.18" pillow "matplotlib<3.9"
| Pin | Why |
|---|---|
torch==2.8.0, torchvision==0.23.0 |
A matched pair from the Jetson index (torchvision's compiled ops must match torch). Newer pairs existed on 2026-09-29 (torch 2.9.1-2.11.0); 2.8.0 is the one installed and tested here. |
numpy<2 |
Lets the venv keep the system's numpy 1.21.5 (live pip list), the one the system OpenCV and ROS 2 packages came with. |
transformers==4.46.3 |
Hugging Face's model library: OWL-ViT's text half for NanoOWL, and all of OWLv2. |
onnx<1.18 |
NanoOWL's engine builder exports the image encoder to ONNX first; without onnx the build fails (step 7). Live 1.17.0. |
pillow, matplotlib<3.9 |
Satisfied by the system copies (Pillow 9.0.1, matplotlib 3.5.1 live). |
Check the GPU from Python. This is the check that was run on 2026-09-29:
On the robot:
cd ~/vision && . venv/bin/activate
python - <<'PY'
import torch, torchvision
print('torch', torch.__version__, 'cuda', torch.cuda.is_available(), torch.version.cuda, 'tv', torchvision.__version__, torch.cuda.get_device_name(0))
from torchvision.ops import roi_align
x = torch.rand(1, 3, 64, 64, device='cuda'); print('roi_align', roi_align(x, [torch.tensor([[0, 0, 32, 32.]], device='cuda')], output_size=(8, 8)).shape)
PY
Check
On the robot:torch 2.8.0 cuda True 12.6 tv 0.23.0 Orin roi_align torch.Size([1, 3, 8, 8])
cuda Truemeans torch sees the GPU.roi_alignis one of torchvision's compiled GPU operations; if it runs,
torchvision matches torch.
If it fails
torch 2.8.0+cpu cuda FalseandAssertionError: Torch not compiled with CUDA enabled: pip took the CPU wheel
from PyPI. On 2026-09-29 this happened with--extra-index-url https://pypi.org/simpleadded. Use
--index-urlalone, thenpip uninstall -y torch torchvisionand install again with--no-cache-dir.ImportError: libcupti.so.12: cannot open shared object file:cuda-cupti-12-6is missing. Install it
(sudo apt-get install -y cuda-cupti-12-6 && sudo ldconfig). It is in the step 2 list for this reason.- The index answers on
pypi.jetson-ai-lab.io(HTTP 200 on 2026-09-29). The olderpypi.jetson-ai-lab.dev
address answered HTTP 410 (gone).
Both are installed from their GitHub source, checked out at the commits that were tested on this robot.
Lines 23-27:
On the robot:
cd ~/vision/src
[ -d torch2trt ] || git clone https://github.com/NVIDIA-AI-IOT/torch2trt.git
[ -d nanoowl ] || git clone https://github.com/NVIDIA-AI-IOT/nanoowl.git
(cd torch2trt && git checkout -q 4e820ae && pip install --no-build-isolation --no-deps .)
(cd nanoowl && git checkout -q fb553de && pip install --no-build-isolation --no-deps .)
4e820ae, 2024-05-03) loads a TensorRT engine as if it were a PyTorchTRTModule to run the image encoder engine (nanoowl/owl_predictor.py line 383).fb553de, 2025-02-05) is the OWL-ViT predictor and the engine builder.--no-build-isolation builds with the packages already in the venv instead of a temporary environment that--no-deps stops pip from "fixing" dependencies by pulling other versions of torch or numpy from PyPI.Check the imports from a directory other than ~/vision/src, so Python imports the installed package and not the
source folder next to it:
On the robot:
cd /tmp && ~/vision/venv/bin/python -c "import transformers, torch2trt, nanoowl; from nanoowl.owl_predictor import OwlPredictor; print('imports ok', transformers.__version__)"
Check
On the robot:imports ok 4.46.3Live
pip listin the venv showstorch2trt 0.5.0andnanoowl 0.0.0(NanoOWL does not set a version).
If it fails
On 2026-09-29 the first try usedpip install -e .(editable) for nanoowl and failed:
ERROR: Project file:///home/burgerbarn/vision/src/nanoowl uses a build backend that is missing the 'build_editable' hook, thenModuleNotFoundError: No module named 'nanoowl.owl_predictor'. Install it with
pip install --no-build-isolation --no-deps .(no-e), as above. Run
pip install setuptools wheelfirst if pip complains about either.
First confirm torch sees the GPU (script line 29), then build one engine per patch size (lines 30-33). Each build
downloads the OWL-ViT weights from Hugging Face, exports the image encoder to an ONNX file in /tmp, and runs
trtexec on it with fp16 at a fixed input of 1x3x768x768.
On the robot:
cd ~/vision && . venv/bin/activate
python -c "import torch; assert torch.cuda.is_available(); print('torch', torch.__version__, 'cuda ok')"
for P in 16 32; do
[ -s engines/owl_image_encoder_patch$P.engine ] || \
python -m nanoowl.build_image_encoder_engine engines/owl_image_encoder_patch$P.engine --model_name google/owlvit-base-patch$P
done
echo VISION_INSTALLED
Each build took 3-4 minutes on this robot (patch32 started 06:01:22, file written 06:04; patch16 started 06:05:58,
written 06:09). The [ -s ... ] || part skips an engine that already exists, so the loop is safe to run again.
Check
The end of the patch16 build log shows thetrtexeccall NanoOWL made:On the robot:
TensorRT.trtexec [TensorRT v100300] # /usr/src/tensorrt/bin/trtexec --onnx=/tmp/tmp99b7iqo5/image_encoder.onnx --saveEngine=engines/owl_image_encoder_patch16.engine --fp16 --shapes=image:1x3x768x768Then
ls -la ~/vision/engines/*.engine:On the robot:
-rw-rw-r-- 1 burgerbarn burgerbarn 181804636 Sep 29 06:09 /home/burgerbarn/vision/engines/owl_image_encoder_patch16.engine -rw-rw-r-- 1 burgerbarn burgerbarn 183555548 Sep 29 06:04 /home/burgerbarn/vision/engines/owl_image_encoder_patch32.engine
If it fails
torch.onnx.OnnxExporterError: Module onnx is not installed!in the build log, and no engine file: the first
build on 2026-09-29 failed this way. Installonnx<1.18from the Jetson index (it is in the step 5 line) and
build again.- A long build with no engine yet is normal for 3-4 minutes. Run it with
nohup ... > engines/build.log 2>&1 &
if your ssh session may drop, and watchtail -c 400 ~/vision/engines/build.log.- Warnings in the build log about normalization layers and ONNX opset 17 appeared on this robot too and did not
stop the build.
from_pretrained() in the transformers library downloads a model's weights the first time it is asked for, into
~/.cache/huggingface/hub/, and reads them from there afterwards. Nothing on the robot downloads them explicitly:
google/owlvit-base-patch16 and google/owlvit-base-patch32 (NanoOWL alsogoogle/owlv2-base-patch16-ensemble (2026-09-29).On the robot:
du -sh ~/.cache/huggingface/hub/models--*
Check (live, 2026-10-07)
On the robot:593M /home/burgerbarn/.cache/huggingface/hub/models--google--owlv2-base-patch16-ensemble 585M /home/burgerbarn/.cache/huggingface/hub/models--google--owlvit-base-patch16 587M /home/burgerbarn/.cache/huggingface/hub/models--google--owlvit-base-patch32OWLv2 appears after its first use (the speed test in "Measured speeds" below does that).
If it fails
The first load needs internet. If the robot is offline,from_pretrainedfails; the cache is part of the golden
sets made after 2026-09-29, so a restored robot has it.
The nightly owner-face job in chapter 21 uses facenet-pytorch (a face detector, MTCNN, plus a face-embedding
network, InceptionResnetV1). It is installed without its dependencies so that pip leaves the Jetson CUDA torch alone
(docs/decisions.md 2026-09-30: "pip --no-deps: keeps the Jetson CUDA torch"). Check torch before and after:
On the robot:
cd ~/vision && . venv/bin/activate
python -c "import torch; print('before torch', torch.__version__, torch.cuda.is_available())"
pip install --no-deps facenet-pytorch==2.6.0
python -c "import torch; print('after torch', torch.__version__, torch.cuda.is_available()); from facenet_pytorch import MTCNN, InceptionResnetV1; print('facenet ok')"
Check
On the robot:before torch 2.8.0 True Successfully installed facenet-pytorch-2.6.0 after torch 2.8.0 True facenet okThe first run of the face job downloads the VGGFace2 weights (107 MB) to
~/.cache/torch/checkpoints/20180402-114759-vggface2.pt.
Not test-built
The command run on 2026-09-30 had no version:pip install --no-deps facenet-pytorch. It installed 2.6.0, which
is what is live. The pin above makes that explicit; it was not run with the pin.
For reference, here is scripts/install_vision.sh as it is in the repo. Steps 2-7 are its lines in order; step 9
(facenet) is not in it.
On the robot:
#!/bin/bash
# Open-vocabulary detection on the robot (2026-09-29, owner OK: vision runs on the robot).
# NVIDIA only: CUDA 12.6 libs, cuDNN 9, TensorRT 10.3 from the robot's NVIDIA L4T apt repo (r36.4) - NOT the
# nvidia-jetpack metapackage (it would add nvidia-l4t-* 36.4.7 packages and replace the system OpenCV 4.5.4 that
# ROS 2 Humble uses, removing 19 packages). nvidia-l4t-dla-compiler pinned to 36.4.3 (= held L4T) and held.
# Python: venv ~/vision/venv (--system-site-packages for tensorrt/cv2); torch 2.8.0 + torchvision 0.23.0 CUDA
# wheels from NVIDIA Jetson AI Lab (pypi.jetson-ai-lab.io/jp6/cu126); transformers 4.46.3; onnx 1.17;
# torch2trt 4e820ae and nanoowl fb553de (NVIDIA-AI-IOT, MIT / Apache-2.0).
# Engines: OWL-ViT base patch16 (default) + patch32, fp16, 768x768. Rollback: bash rollback_vision.sh
set -euo pipefail
sudo apt-get install -y cuda-libraries-12-6 cuda-cudart-dev-12-6 cuda-nvcc-12-6 cuda-cupti-12-6 libcudnn9-cuda-12 \
tensorrt-libs libnvinfer-bin python3-libnvinfer libnvinfer-dev libnvonnxparsers-dev libnvinfer-plugin-dev \
python3-venv python3-pip
sudo apt-get install -y nvidia-l4t-dla-compiler=36.4.3-20250107174145 && sudo apt-mark hold nvidia-l4t-dla-compiler
sudo ldconfig
python3 -c "import tensorrt; print('tensorrt', tensorrt.__version__)"
IDX=https://pypi.jetson-ai-lab.io/jp6/cu126/+simple
mkdir -p ~/vision/src ~/vision/engines && cd ~/vision
[ -d venv ] || python3 -m venv --system-site-packages venv
. venv/bin/activate
pip install --upgrade pip setuptools wheel
pip install --index-url $IDX "torch==2.8.0" "torchvision==0.23.0" "numpy<2" "transformers==4.46.3" "onnx<1.18" pillow "matplotlib<3.9"
cd src
[ -d torch2trt ] || git clone https://github.com/NVIDIA-AI-IOT/torch2trt.git
[ -d nanoowl ] || git clone https://github.com/NVIDIA-AI-IOT/nanoowl.git
(cd torch2trt && git checkout -q 4e820ae && pip install --no-build-isolation --no-deps .)
(cd nanoowl && git checkout -q fb553de && pip install --no-build-isolation --no-deps .)
cd ~/vision
python -c "import torch; assert torch.cuda.is_available(); print('torch', torch.__version__, 'cuda ok')"
for P in 16 32; do
[ -s engines/owl_image_encoder_patch$P.engine ] || \
python -m nanoowl.build_image_encoder_engine engines/owl_image_encoder_patch$P.engine --model_name google/owlvit-base-patch$P
done
echo VISION_INSTALLED
Not test-built
The script was written at 06:10 on 2026-09-29 from the hand steps above, which include the two fixes found on the
way (CUPTI in the apt line, onnx in the pip line). It has never been run on a clean system as one piece. If you run
it (bash ~/setup/install_vision.sh), it stops at the first failing line (set -euo pipefail); fix that line
with the fail boxes above and run it again: every step either skips work already done or repeats it harmlessly.
Now use the detector the way the robot first did: grab one camera frame, then run NanoOWL on it. Two small programs
do this. They are already in ~/vision from step 1; read them before you run them.
vision/grab.py subscribes to the camera's colour and depth topics, keeps the first message of each, and saves them.
It runs with the system Python and ROS 2 (no venv needed).
The top creates a dictionary got and a node whose two subscriptions each store their first message
(setdefault keeps the first and ignores the rest):
On the robot:
"""Save one frame of /aurora/rgb/image_raw (and depth) to ~/vision/frame_rgb.png / frame_depth.npy."""
import numpy as np, rclpy, cv2
from rclpy.node import Node
from sensor_msgs.msg import Image
got = {}
class G(Node):
def __init__(self):
super().__init__('grab')
for k, t in (('rgb', '/aurora/rgb/image_raw'), ('depth', '/aurora/depth/image_raw')):
self.create_subscription(Image, t, lambda m, k=k: got.setdefault(k, m), 5)
Then it spins until both have arrived. There is no time limit in the loop, so run it under timeout:
On the robot:
rclpy.init(); n = G()
while len(got) < 2: rclpy.spin_once(n, timeout_sec=1)
A ROS Image message is a flat byte buffer plus width, height and encoding. np.frombuffer(...).reshape(...) turns
it into an image array. OpenCV writes images in blue-green-red order; the Aurora publishes bgr8, so the array is
written as it is (an rgb8 image would be converted first). Depth is 16-bit, one value per pixel, in millimetres:
On the robot:
m = got['rgb']; print('rgb', m.width, m.height, m.encoding, m.header.frame_id)
a = np.frombuffer(m.data, np.uint8).reshape(m.height, m.width, -1)
cv2.imwrite('frame_rgb.png', cv2.cvtColor(a, cv2.COLOR_RGB2BGR) if m.encoding == 'rgb8' else a)
d = got['depth']; print('depth', d.width, d.height, d.encoding, d.header.frame_id)
np.save('frame_depth.npy', np.frombuffer(d.data, np.uint16).reshape(d.height, d.width))
The complete file, ~/vision/grab.py:
On the robot:
"""Save one frame of /aurora/rgb/image_raw (and depth) to ~/vision/frame_rgb.png / frame_depth.npy."""
import numpy as np, rclpy, cv2
from rclpy.node import Node
from sensor_msgs.msg import Image
got = {}
class G(Node):
def __init__(self):
super().__init__('grab')
for k, t in (('rgb', '/aurora/rgb/image_raw'), ('depth', '/aurora/depth/image_raw')):
self.create_subscription(Image, t, lambda m, k=k: got.setdefault(k, m), 5)
rclpy.init(); n = G()
while len(got) < 2: rclpy.spin_once(n, timeout_sec=1)
m = got['rgb']; print('rgb', m.width, m.height, m.encoding, m.header.frame_id)
a = np.frombuffer(m.data, np.uint8).reshape(m.height, m.width, -1)
cv2.imwrite('frame_rgb.png', cv2.cvtColor(a, cv2.COLOR_RGB2BGR) if m.encoding == 'rgb8' else a)
d = got['depth']; print('depth', d.width, d.height, d.encoding, d.header.frame_id)
np.save('frame_depth.npy', np.frombuffer(d.data, np.uint16).reshape(d.height, d.width))
Run it:
On the robot:
source /opt/ros/humble/setup.bash
cd ~/vision && timeout 20 python3 grab.py
Check
On the robot:rgb 640 400 bgr8 rgb_camera_link depth 640 400 mono16 depth_camera_link
~/vision/frame_rgb.pngand~/vision/frame_depth.npynow exist. Copy the picture to your laptop to look at it:
scp rosorin-wifi:vision/frame_rgb.png .
If it fails
- It prints nothing and
timeoutends it after 20 s: no frames. Checktimeout 8 ros2 topic hz /aurora/rgb/image_raw
(it printedaverage rate: 13.848on 2026-10-07). On 2026-10-07 the camera service was "active" after a boot but
published nothing for 23 minutes;sudo systemctl restart rosorin-camerabrought frames back.vision/README.mdstill says to start the camera withros2 launch deptrum-ros-driver-aurora930 aurora930_launch.pyby hand. That was beforerosorin-camera.serviceexisted (chapter 15); the docs differ,
the service is what runs.
vision/detect.py loads NanoOWL with one engine, scores your prompts on the image, prints each hit with its score,
box and the median depth inside the box, and saves a copy of the picture with the boxes drawn. It does not need ROS.
It starts with its usage text and imports, and the default prompts it was written for (the owner wanted the robot to
find his microphone, TV and speakers):
On the robot:
"""Open-vocabulary detection on the robot (NanoOWL = OWL-ViT with a TensorRT image encoder, NVIDIA-AI-IOT, Apache-2.0).
Run in ~/vision/venv. Usage:
python detect.py frame_rgb.png [--depth frame_depth.npy] [--prompts "a microphone,a television,a speaker"] [--th 0.1]
Writes <image>_det.png (boxes) and prints label, score, box and median depth inside the box (metres; depth is
the Aurora depth image at the same 640x400 size, assumed roughly aligned with RGB - first cut)."""
import argparse, os, time
import cv2
import numpy as np
import PIL.Image
from nanoowl.owl_predictor import OwlPredictor
HERE = os.path.dirname(os.path.abspath(__file__))
PROMPTS = 'a microphone,a television,a speaker'
The options: the image, an optional depth file, the prompts, the score threshold (0.1), and which engine to use:
On the robot:
ap = argparse.ArgumentParser()
ap.add_argument('image')
ap.add_argument('--depth')
ap.add_argument('--prompts', default=PROMPTS)
ap.add_argument('--th', type=float, default=0.1)
ap.add_argument('--model', default='16', choices=['16', '32'], help='OWL-ViT base patch size (16 = finer, slower)')
a = ap.parse_args()
Loading: OwlPredictor takes the Hugging Face model name (for the text half and the box head) and the path of the
image-encoder engine. encode_text turns the prompts into vectors once; they are reused for every image:
On the robot:
t = time.time()
pred = OwlPredictor(f'google/owlvit-base-patch{a.model}',
image_encoder_engine=os.path.join(HERE, 'engines', f'owl_image_encoder_patch{a.model}.engine'))
text = [p.strip() for p in a.prompts.split(',') if p.strip()]
enc = pred.encode_text(text)
print(f'load {time.time() - t:.1f}s')
Prediction runs three times and reports the time of the last one: the first call of a GPU program is slow while it
warms up. pad_square=False lets NanoOWL stretch the 640x400 frame to its square input instead of padding it:
On the robot:
img = PIL.Image.open(a.image).convert('RGB')
for i in range(3): # first call warms up
t = time.time()
out = pred.predict(image=img, text=text, text_encodings=enc, threshold=a.th, pad_square=False)
dt = time.time() - t
print(f'predict {dt * 1000:.0f} ms')
For depth it takes the middle half of each box (a quarter trimmed from every side, so the edges of the object and the
background behind it are left out), keeps the pixels with a depth reading (> 0) and reports their median in metres
and the share of pixels that had a reading:
On the robot:
depth = np.load(a.depth) if a.depth else None
vis = cv2.cvtColor(np.array(img), cv2.COLOR_RGB2BGR)
for lab, sc, box in zip(out.labels.tolist(), out.scores.tolist(), out.boxes.tolist()):
x0, y0, x1, y1 = [int(round(v)) for v in box]
d = ''
if depth is not None:
cx0, cx1 = x0 + (x1 - x0) // 4, x1 - (x1 - x0) // 4
cy0, cy1 = y0 + (y1 - y0) // 4, y1 - (y1 - y0) // 4
roi = depth[max(0, cy0):cy1, max(0, cx0):cx1]
v = roi[roi > 0]
d = f' depth {np.median(v) / 1000:.2f} m ({v.size / max(1, roi.size):.0%} valid)' if v.size else ' depth n/a'
print(f'{text[lab]:14} {sc:.2f} box ({x0},{y0})-({x1},{y1}){d}')
cv2.rectangle(vis, (x0, y0), (x1, y1), (0, 200, 255), 2)
cv2.putText(vis, f'{text[lab]} {sc:.2f}', (x0, max(12, y0 - 4)), cv2.FONT_HERSHEY_SIMPLEX, 0.45, (0, 200, 255), 1)
cv2.imwrite(os.path.splitext(a.image)[0] + '_det.png', vis)
The complete file, ~/vision/detect.py:
On the robot:
"""Open-vocabulary detection on the robot (NanoOWL = OWL-ViT with a TensorRT image encoder, NVIDIA-AI-IOT, Apache-2.0).
Run in ~/vision/venv. Usage:
python detect.py frame_rgb.png [--depth frame_depth.npy] [--prompts "a microphone,a television,a speaker"] [--th 0.1]
Writes <image>_det.png (boxes) and prints label, score, box and median depth inside the box (metres; depth is
the Aurora depth image at the same 640x400 size, assumed roughly aligned with RGB - first cut)."""
import argparse, os, time
import cv2
import numpy as np
import PIL.Image
from nanoowl.owl_predictor import OwlPredictor
HERE = os.path.dirname(os.path.abspath(__file__))
PROMPTS = 'a microphone,a television,a speaker'
ap = argparse.ArgumentParser()
ap.add_argument('image')
ap.add_argument('--depth')
ap.add_argument('--prompts', default=PROMPTS)
ap.add_argument('--th', type=float, default=0.1)
ap.add_argument('--model', default='16', choices=['16', '32'], help='OWL-ViT base patch size (16 = finer, slower)')
a = ap.parse_args()
t = time.time()
pred = OwlPredictor(f'google/owlvit-base-patch{a.model}',
image_encoder_engine=os.path.join(HERE, 'engines', f'owl_image_encoder_patch{a.model}.engine'))
text = [p.strip() for p in a.prompts.split(',') if p.strip()]
enc = pred.encode_text(text)
print(f'load {time.time() - t:.1f}s')
img = PIL.Image.open(a.image).convert('RGB')
for i in range(3): # first call warms up
t = time.time()
out = pred.predict(image=img, text=text, text_encodings=enc, threshold=a.th, pad_square=False)
dt = time.time() - t
print(f'predict {dt * 1000:.0f} ms')
depth = np.load(a.depth) if a.depth else None
vis = cv2.cvtColor(np.array(img), cv2.COLOR_RGB2BGR)
for lab, sc, box in zip(out.labels.tolist(), out.scores.tolist(), out.boxes.tolist()):
x0, y0, x1, y1 = [int(round(v)) for v in box]
d = ''
if depth is not None:
cx0, cx1 = x0 + (x1 - x0) // 4, x1 - (x1 - x0) // 4
cy0, cy1 = y0 + (y1 - y0) // 4, y1 - (y1 - y0) // 4
roi = depth[max(0, cy0):cy1, max(0, cx0):cx1]
v = roi[roi > 0]
d = f' depth {np.median(v) / 1000:.2f} m ({v.size / max(1, roi.size):.0%} valid)' if v.size else ' depth n/a'
print(f'{text[lab]:14} {sc:.2f} box ({x0},{y0})-({x1},{y1}){d}')
cv2.rectangle(vis, (x0, y0), (x1, y1), (0, 200, 255), 2)
cv2.putText(vis, f'{text[lab]} {sc:.2f}', (x0, max(12, y0 - 4)), cv2.FONT_HERSHEY_SIMPLEX, 0.45, (0, 200, 255), 1)
cv2.imwrite(os.path.splitext(a.image)[0] + '_det.png', vis)
Run it on your frame with both engines:
On the robot:
cd ~/vision && . venv/bin/activate
python detect.py frame_rgb.png --depth frame_depth.npy --model 32
python detect.py frame_rgb.png --depth frame_depth.npy --model 16
On your laptop:
scp rosorin-wifi:vision/frame_rgb_det.png .
What it printed on 2026-09-29
One 640x400 frame from the tucked arm, dim room, facing the TV and the fireplace speakers. Patch32:On the robot:
load 3.1s predict 28 ms a speaker 0.14 box (289,47)-(328,91) depth 3.01 m (100% valid) a speaker 0.15 box (279,110)-(330,159) depth n/a a television 0.54 box (378,138)-(527,240) depth n/aPatch16, same frame:
On the robot:
load 3.5s predict 115 ms a speaker 0.21 box (278,113)-(326,159) depth n/a a television 0.53 box (384,142)-(519,235) depth n/aYour room gives other boxes and scores. What counts:
loada few seconds,predicttens of milliseconds for
patch32 and around 100 for patch16, and boxes on things you can see inframe_rgb_det.png.
Things to notice in that output, all from the robot's record (vision/README.md):
a speaker at 0.14 with depth 3.01 m was the lampshade above the speaker, not a speaker. Adding"a lamp" to the prompts gave the lampshade its own label (a lamp 0.10) on the patch16 run. That is the ideadepth n/a: the Aurora's depth has no reading on dark or shiny surfaces. Placing"a tv" scored 0.43 where "a television" scored"a speaker cabinet", "a mic" and "a lampshade" found nothing:On the robot:
python detect.py frame_rgb.png --depth frame_depth.npy --model 32 --prompts "a tv,a speaker cabinet,a mic,a lampshade,a guitar"
If it fails
FileNotFoundErrorfor an.enginefile: step 7 did not finish for that patch size.- Lines like
[TRT] [W] Using default stream in enqueueV3() may lead to performance issues ...and, on later
starts,[TRT] [W] Using an engine plan file across different models of devices is not recommendedare printed
by TensorRT. Both appeared in the recorded runs whose results are shown above, and the robot's mind has run with
these engines since 2026-09-29.- Running detection while the mind (chapter 21) is running is fine: on 2026-10-05 a detector test ran on the GPU
beside the mind for about a minute. Both share the GPU, so your timings will be slower than the ones above.
Three detectors are installed. These are all the speed measurements in the record, each on the robot's own camera
frames.
| Detector | One frame by hand | Live, 20 s on the camera stream (bench_live.py) |
Used by |
|---|---|---|---|
| NanoOWL patch32 | 28 ms predict, 3 prompts (detect.py) |
14.2 fps (the camera's own rate, so camera-limited), latency p50 49 ms / p95 52 ms, GPU busy 53 %, board 8.5 W, 5 prompts | the mind, all the time |
| NanoOWL patch16 | 115 ms predict, 3 prompts (detect.py) |
11.6 fps, latency p50 69 ms / p95 77 ms, GPU busy 73 %, board 21.9 W, 5 prompts | detect.py default |
| OWLv2 base patch16 ensemble | 640 ms for the network alone; 946 ms through make_detector with box decoding and NMS; 13.5 s to load |
not run live | scan.py, teach.py, floor_check.py |
The live numbers come from vision/bench_live.py (2026-09-29 23:00 UTC). A second run on 2026-09-30, with the
since-disabled on-board language model loaded, gave patch32 14.1 fps and GPU 56 % alone, 14.1 fps and GPU 72 % while
the language model was answering: the detector kept its rate. The camera itself publishes about 14 frames a second
(ros2 topic hz measured 13.848 on 2026-10-07), so patch32 keeps up with every frame.
To measure your own robot (no motion, camera running):
On the robot:
source /opt/ros/humble/setup.bash
cd ~/vision && . venv/bin/activate
for m in 32 16; do python bench_live.py --model $m --seconds 20 2>&1 | grep -v Warn | tail -3; done
Check
On the robot:NanoOWL patch32, 5 prompts: 283 frames in 20 s = 14.2 fps (camera-limited if ~= camera rate) | latency ms p50 49 p95 52 GPU busy 53% (max 99%) | CPU avg 19% | RAM 3697 MB | board power 8.5 W seen: {'a person': 0.28, 'a face': 0.27, 'a cat': 0.14} NanoOWL patch16, 5 prompts: 232 frames in 20 s = 11.6 fps (camera-limited if ~= camera rate) | latency ms p50 69 p95 77 GPU busy 73% (max 99%) | CPU avg 16% | RAM 3591 MB | board power 21.9 W seen: {'a person': 0.18}(Recorded with the mind not yet running. With
rosorin-mindactive your numbers include its load.)
The OWLv2 timing was taken with make_detector from vision/scan.py on a saved frame (tuck_now.png then; use
your frame_rgb.png). The first call loads the model (and downloads it on a new robot); the second is timed:
On the robot:
source /opt/ros/humble/setup.bash
cd ~/vision && . venv/bin/activate && python -c "
import time, PIL.Image, sys; sys.argv=['x']
from scan import make_detector
t = ['a microphone', 'a television', 'a speaker', 'a lamp']
d = make_detector('owlv2', '16', t)
img = PIL.Image.open('frame_rgb.png').convert('RGB')
d(img, 0.15); s = time.time(); r = d(img, 0.15); print('ms', round((time.time()-s)*1000))
for x in r: print(x[0], round(x[1], 2), [int(v) for v in x[2]])
" 2>&1 | grep -vE "arn|^ "
Check
On 2026-09-29, on the same scene as the NanoOWL run above:On the robot:
ms 946 a television 0.62 [411, 113, 529, 194] a speaker 0.6 [299, 81, 339, 129] a speaker 0.57 [87, 55, 136, 106] a speaker 0.36 [562, 168, 580, 209] a lamp 0.34 [305, 12, 340, 83] a microphone 0.17 [190, 50, 205, 68]OWLv2 found both fireplace speakers (0.60 and 0.57; NanoOWL patch16: 0.21 and a miss) and the microphone (0.17;
NanoOWL: missed), at about 30 times the cost per frame. The owner confirmed the microphone's position.
The docstring of make_detector says "~0.64 s/frame" for OWLv2 and "~0.12 s/frame" for NanoOWL; those are the
network-only OWLv2 time and the patch16 single-frame time. The docs differ from the measurements in the table; the
table is what was measured.
Everything on the robot that detects objects gets its detector from one function, make_detector() in
vision/scan.py (lines 45-110). scan.py is a larger file (the older scripted look-around scan, 326 lines); this
function is the part the mind uses. It returns a function detect(image, threshold) that gives a list of
(label, score, [x0, y0, x1, y1]), best first.
The signature and docstring:
On the robot:
def make_detector(kind, model, text, iou=0.3, per_label=3, decoys=()):
"""-> detect(PIL image, threshold) = [(label, score, [x0, y0, x1, y1]), ...], per-label NMS, best first.
owlv2 (default): google/owlv2-base-patch16-ensemble, HF transformers fp16, ~0.64 s/frame on the Orin, much
better on small objects (2026-09-29, same frame: speakers 0.60/0.57, mic 0.17 vs NanoOWL 0.21/missed/missed).
nanoowl: OWL-ViT + TensorRT, ~0.12 s/frame."""
import torch
from torchvision.ops import box_iou, nms
kind is 'owlv2' or 'nanoowl', model the patch size ('16' or '32', NanoOWL only), text the prompt list.
iou=0.3 is the NMS overlap limit, per_label=3 the most boxes kept per label, decoys the labels that exist only
to catch look-alikes.
The OWLv2 branch loads the model in fp16 on the GPU and defines raw(), which returns labels, scores, boxes and
one embedding per box. It decodes the boxes itself instead of calling the library's post-processing, so each box keeps
its class embedding (a "fingerprint" of what is in the box, which vision/teach.py uses to learn from examples).
OWLv2 pads the image to a square from the top-left, so boxes are scaled by the longer side:
On the robot:
if kind == 'owlv2':
from transformers import Owlv2ForObjectDetection, Owlv2Processor
name = 'google/owlv2-base-patch16-ensemble'
proc = Owlv2Processor.from_pretrained(name)
net = Owlv2ForObjectDetection.from_pretrained(name, torch_dtype=torch.float16).cuda().eval()
def raw(img, th):
inp = proc(text=[text], images=img, return_tensors='pt').to('cuda')
inp['pixel_values'] = inp['pixel_values'].half()
with torch.no_grad():
out = net(**inp)
s = max(img.size) # OWLv2 pads to a square (top-left)
# same maths as proc.post_process_object_detection (sigmoid of the best query logit, cxcywh -> xyxy
# scaled to the padded square), done here so each box keeps its class embedding ("fingerprint")
scores, labels = torch.sigmoid(out.logits[0].float()).max(-1)
keep = scores > th
b = out.pred_boxes[0][keep].float()
boxes = torch.stack([b[:, 0] - b[:, 2] / 2, b[:, 1] - b[:, 3] / 2,
b[:, 0] + b[:, 2] / 2, b[:, 1] + b[:, 3] / 2], 1) * s
emb = torch.nn.functional.normalize(out.class_embeds[0][keep].float(), dim=-1)
return labels[keep], scores[keep], boxes, emb
The NanoOWL branch is the same OwlPredictor call as in detect.py. The encoded prompts are kept in a small dict
st so that add_labels (below) can replace them later:
On the robot:
else:
from nanoowl.owl_predictor import OwlPredictor
pred = OwlPredictor(f'google/owlvit-base-patch{model}',
image_encoder_engine=os.path.join(HERE, 'engines', f'owl_image_encoder_patch{model}.engine'))
st = {'enc': pred.encode_text(text)}
def raw(img, th):
o = pred.predict(image=img, text=text, text_encodings=st['enc'], threshold=th, pad_square=False)
return o.labels, o.scores.float(), o.boxes.float(), None
detect() is shared by both. For each label it runs NMS (IoU 0.3) and keeps the best three. Then the decoy rule:
a target box is dropped when a decoy box overlaps it by IoU 0.5 or more and scores at least as high. A lampshade that
scores as "a speaker" is removed when "a lamp" scores the same box higher:
On the robot:
def detect(img, th, with_embeds=False):
"""with_embeds: 4-tuples (label, score, box, embedding) - owlv2 only (embedding None for nanoowl)."""
labels, scores, boxes, emb = raw(img, th)
out = []
for li in set(labels.tolist()):
idx = torch.nonzero(labels == li).squeeze(1)
keep = nms(boxes[idx].cpu(), scores[idx].cpu(), iou)[:per_label]
out += [(text[li], float(scores[idx[k]]), boxes[idx[k]].tolist(),
emb[idx[k]].cpu() if emb is not None else None) for k in keep]
# a target box that mostly overlaps a stronger decoy box (lamp, tower fan) is that decoy, not the target
dec = [d for d in out if d[0] in decoys]
if dec:
db = torch.tensor([d[2] for d in dec])
out = [d for d in out if d[0] in decoys or not any(
s >= d[1] and v >= 0.5 for s, v in zip((x[1] for x in dec), box_iou(torch.tensor([d[2]]), db)[0].tolist()))]
out = sorted(out, key=lambda d: -d[1])
return out if with_embeds else [d[:3] for d in out]
add_labels teaches new words while the detector is running. It appends to the same text list and, for NanoOWL,
re-encodes the prompts once (OWLv2 encodes the text on every call anyway). The mind uses this when the owner names an
object (chapter 21):
On the robot:
def add_labels(new):
"""Teach new object names live (owner 2026-10-01: learn the moment the owner names something)."""
new = [t for t in new if t and t not in text]
if new:
text.extend(new)
if kind != 'owlv2':
st['enc'] = pred.encode_text(text)
return new
detect.add_labels = add_labels
return detect
Why two detectors
NanoOWL patch32 keeps up with the camera at half the GPU, which is what following a person needs. OWLv2 is about
30 times slower but finds small things NanoOWL misses (the microphone, the speaker in shadow), which is what a
careful look at a still scene needs. Both come from the same factory, so a caller switches with one argument.
The older scripted look-around inscan.pyuses OWLv2 and adds a per-label score floor and a minimum number of
views (vision/scan_filter.py:a microphone0.15 / 1 view,a televisionanda speaker0.5 / 2 views,
anything else 0.3 / 2). Scripted arm scans were ruled out on 2026-10-02; the mind does not use that filter.
The mind's prompts and thresholds are in vision/attention.py lines 24-30 (the mind imports them from there):
On the robot:
PROMPTS = ['a face', 'a person', 'a hand']
TH = {'a face': 0.2, 'a person': 0.25, 'a hand': 0.22}
# everyday objects: give the robot's thinking something to be curious about (2026-09-30: it only ever saw 'nothing specific')
OBJECTS = ['a cup', 'a phone', 'a bag', 'a box', 'a cable', 'a shoe', 'a guitar', 'a laptop']
TH.update({o: 0.2 for o in OBJECTS})
GAIN, DEAD_DEG, VMAX = 1.8, 1.5, 1.25 # 1/s, deg, rad/s (owner 2026-10-01: head faster; was 1.6, 0.7; 2.5 overshot +-18 deg)
LOST_S, LOOK_S, SWEEP_V = 1.5, 2.5, 0.25 # s, s, rad/s
How the mind (vision/mind.py) uses them:
PROMPTS + OBJECTS (line 120), adds every word the owner has taught it~/vision/learned_objects.json (line 121, 17 words on 2026-10-05), and adds 'an object' so that unknownTH for the words above, LEARNED_TH = 0.2 for taught words and 'an object' (lines 41 and 332):On the robot:
dets = [d for d in self.detect(PIL.Image.fromarray(rgb), 0.12) if d[1] >= TH.get(d[0], LEARNED_TH)]
What the record says about these numbers:
docs/research/whole_project_review.md, finding 6).These are open problems, listed again in chapter 21 and chapter 31.
scripts/rollback_vision.sh removes what install_vision.sh added:
On the robot:
#!/bin/bash
# Undo install_vision.sh: remove ~/vision/venv + engines + src, and the NVIDIA packages it added.
set -euo pipefail
rm -rf ~/vision/venv ~/vision/engines ~/vision/src ~/.cache/huggingface/hub/models--google--owlvit-base-patch*
sudo apt-mark unhold nvidia-l4t-dla-compiler
sudo apt-get remove -y nvidia-l4t-dla-compiler cuda-libraries-12-6 cuda-cudart-dev-12-6 cuda-nvcc-12-6 cuda-cupti-12-6 \
libcudnn9-cuda-12 tensorrt-libs libnvinfer-bin python3-libnvinfer libnvinfer-dev libnvonnxparsers-dev libnvinfer-plugin-dev
sudo apt-get autoremove -y
echo VISION_ROLLED_BACK
Do not run this script as it is
Its line 8 issudo apt-get autoremove -y. The project's own rule is never to autoremove on the robot, because
apt then removesinitramfs-tools, which the Jetson needs to boot (docs/lessons.md, 2026-09-26). Run lines 4-7
by hand instead, and simulate the remove first (apt-get -s remove ...) to see what else it would take with it.
The script has never been run on this robot.
What the rollback leaves behind: the OWLv2 weights in ~/.cache/huggingface/hub/models--google--owlv2-base-patch16-ensemble,
the face weights in ~/.cache/torch/checkpoints/, and the vision code in ~/vision. Delete those by hand if you want
them gone. Everything that uses the venv (the mind, the robot API, the nightly timers; chapters 21-22) stops working
after a rollback, so stop those services first.
The cleaner way back is the golden set made right after this chapter on this robot: 20260929T1441Z ("+ vision
stack") in docs/restore.md, restored as in chapter 8.
python3 -c "import tensorrt; print(tensorrt.__version__)" prints 10.3.0.apt-mark showhold | grep dla prints nvidia-l4t-dla-compiler, and apt-mark showhold | grep -c nvidia-l4ttorch.cuda.is_available() is True and torch.__version__ is 2.8.0.~/vision/engines/ holds two engines of about 182 MB each.detect.py on your own frame draws boxes on things you can see, with patch32 in tens of milliseconds.Where this comes from
scripts/install_vision.sh,scripts/rollback_vision.sh(commit abe3f88, 2026-09-29);vision/grab.py,
vision/detect.py,vision/bench_live.py,vision/scan.pylines 45-110,vision/scan_filter.py,
vision/attention.pylines 24-30,vision/mind.pylines 40-41, 120-122, 332;vision/README.md(first results,
OWLv2);docs/decisions.md2026-09-30 (facenet --no-deps);docs/lessons.md2026-09-26 (never autoremove);
docs/restore.md(golden sets 20260929T1441Z, 20260929T1829Z);docs/capabilities.md;
docs/research/whole_project_review.mdfindings 6, 27, 32, 35. Command logsources/cmdlog_robot.md
2026-09-29 05:40-06:47 UTC (jetpack simulation, apt install, DLA compiler, torch, nanoowl, engine builds, first
detections, OWLv2 test), 2026-09-29 23:00 UTC (bench_live), 2026-09-30 01:53 and 03:05 UTC (facenet, bench with the
on-board model), 2026-10-07 01:36 UTC (camera rate). Live read-only check of the robot on 2026-10-07 (package
versions, venvpyvenv.cfg,pip list, Hugging Face cache sizes, engine files, nanoowlowl_predictor.py
lines 383 and 443).