The robot's own mind · Chapter 23 · Time: 3 hours · Level: Intermediate · Status: Done on this robot
Every 15 minutes the robot measures itself, compares the numbers with its own normal, fixes the few things it is allowed to fix, measures again to prove the fix, and asks you out loud for the rest. You build the loop, install its timer, and prove it with injected faults.
A robot that has to run for a day without help needs to notice its own faults. On 2026-10-05 the robot's load
average sat at 30 for hours because of two faults nobody saw until a person measured them: a maths library spinning
idle threads, and processes that could not share memory. The self-care loop (selfcare/, run by
rosorin-selfcheck.timer) is the robot doing that measuring itself. Its rule: code decides what is true and what
is allowed; the language model on bigbuddy may only choose which extra check to run next, and it never acts.
New idea: evidence first, then rules, then the model
The pattern comes from small open ROS 2 diagnosis projects (RobotOps, RoboDiag): fixed read-only checks
produce numbered evidence; a short playbook in code maps known patterns to a small allow-list of
repairs; every repair is verified by measuring again ("a zero exit code counts for nothing") and rolled back
if it did not help. A language model is used only where no rule applies, and only to pick the next check and
name a cause that cites evidence numbers. On 2026-10-05 the robot's own brain model (Gemma 4 12B on bigbuddy)
was tested with that day's real symptoms: it chose the right next check, but it proposed no remedy, repeated a
check it had already run, and for a low battery it did not ask to be plugged in. So the remedies and the
sentences for the owner come from code.
New idea: a busy-waiting thread
A thread that has nothing to do should sleep. A busy-waiting thread instead loops, asking the kernel "anything
for me?" (thesched_yieldcall) thousands of times a second. It shows up as high CPU use with most of it in
the kernel. On 2026-10-05 the camera odometry node (vslam_odom, chapter 16) used 2 of 8 cores while standing
still because three worker threads of OpenBLAS, the maths library numpy uses, were spinning like that.
OPENBLAS_NUM_THREADS=1stops it: 0.33 core afterwards. Self-care recognises that pattern: one thread at= 40 % of a core with >= 60 % of its time in the kernel, in a process that has OpenBLAS loaded.
| File | Lines | Job |
|---|---|---|
selfcare/common.py |
66 | Paths, the owner's off switch, small helpers, the evidence Ledger. |
selfcare/checks.py |
281 | The read-only checks. Each adds one piece of evidence. |
selfcare/baseline.py |
22 | What is normal for this robot: the median of its own healthy rounds. |
selfcare/expected.json |
65 | Seed values (until there is history), limits, topic owners, the robot's units. |
selfcare/playbook.py |
184 | Known patterns and the allow-listed actions, each verified. |
selfcare/investigate.py |
61 | The brain-guided extra checks, from a closed menu. |
selfcare/say.py |
41 | The queue of sentences for Buddy to say, with quiet hours. |
selfcare/self_check.py |
56 | One round: measure, decide, investigate, learn, report. |
systemd/rosorin-selfcheck.service, .timer |
13, 9 | Run a round 5 min after boot and then every 15 min. |
scripts/install_selfcare.sh |
12 | Install the unit and timer, enable the timer. |
All of it runs with the system python3 and the ROS 2 Python packages, not the vision venv. Write the files in your
repo on your laptop (~/CCode/rosorin-pro/selfcare/), then copy them to ~/selfcare on the robot.
Start with where things live and the owner's switch. If the file ~/selfcare/DISABLED exists, the robot keeps
measuring but changes nothing. SKILL_UNITS lists the wheel units of chapter 24: while one runs, the robot is driving.
On your laptop:
"""Self-care: shared pieces - where things are kept, the evidence ledger, small helpers.
The robot notices, diagnoses, fixes what it is allowed to fix, verifies, and asks its owner for the rest
(docs/research/doing_it_itself_reference.md). Code decides what is true and what is allowed; the brain model only
chooses which further check to run and names a cause."""
import json, os, subprocess, time
HOME = os.path.expanduser('~')
DATA = os.path.join(HOME, 'selfcare', 'data')
HERE = os.path.dirname(os.path.abspath(__file__))
DISABLED = os.path.join(HOME, 'selfcare', 'DISABLED') # owner's switch: touch this file and nothing is changed
SKILL_UNITS = ('rosorin-explore', 'rosorin-practice', 'rosorin-task@come', 'rosorin-task@home', 'rosorin-task@sethome')
os.makedirs(DATA, exist_ok=True)
Then small helpers. save() writes to a temporary file and renames it, so a reader never sees half a file (the
robot API reads latest.json and say_queue.json while self-care writes them). run() never raises: a check that
cannot run a command gets a return code of 1 and the error text.
On your laptop:
def load(path, default):
try:
return json.load(open(path))
except (OSError, ValueError):
return default
def save(path, obj):
tmp = path + '.tmp'
json.dump(obj, open(tmp, 'w'), indent=1)
os.replace(tmp, path)
def append(path, obj):
with open(path, 'a') as f:
f.write(json.dumps(obj) + '\n')
def run(cmd, timeout=30):
"""(return code, stdout) of a command; never raises."""
try:
r = subprocess.run(cmd, capture_output=True, text=True, timeout=timeout)
return r.returncode, r.stdout.strip()
except (OSError, subprocess.TimeoutExpired) as e:
return 1, str(e)
def unit_active(unit):
return run(['systemctl', 'is-active', '--quiet', unit])[0] == 0
def skill_running():
return next((u for u in SKILL_UNITS if unit_active(u)), None)
Last, the ledger. Every measurement becomes one numbered item (E1, E2, ...). Findings, actions and the brain's
causes refer to these numbers, so every claim can be traced to a measurement. brief() is the one-line summary the
brain gets.
On your laptop:
class Ledger:
"""Numbered evidence. Every finding, action and diagnosis refers to these ids."""
def __init__(self):
self.items = []
def add(self, check, status, text, facts=None, args=None):
e = dict(id=f'E{len(self.items) + 1}', check=check, args=args or {}, status=status, text=text,
facts=facts or {}, t=round(time.time(), 1))
self.items.append(e)
return e
def get(self, check):
return next((e for e in reversed(self.items) if e['check'] == check), None)
def brief(self):
return ' '.join(f"{e['id']} {e['check']}: {e['text']}" for e in self.items)
The complete file, selfcare/common.py:
On your laptop:
"""Self-care: shared pieces - where things are kept, the evidence ledger, small helpers.
The robot notices, diagnoses, fixes what it is allowed to fix, verifies, and asks its owner for the rest
(docs/research/doing_it_itself_reference.md). Code decides what is true and what is allowed; the brain model only
chooses which further check to run and names a cause."""
import json, os, subprocess, time
HOME = os.path.expanduser('~')
DATA = os.path.join(HOME, 'selfcare', 'data')
HERE = os.path.dirname(os.path.abspath(__file__))
DISABLED = os.path.join(HOME, 'selfcare', 'DISABLED') # owner's switch: touch this file and nothing is changed
SKILL_UNITS = ('rosorin-explore', 'rosorin-practice', 'rosorin-task@come', 'rosorin-task@home', 'rosorin-task@sethome')
os.makedirs(DATA, exist_ok=True)
def load(path, default):
try:
return json.load(open(path))
except (OSError, ValueError):
return default
def save(path, obj):
tmp = path + '.tmp'
json.dump(obj, open(tmp, 'w'), indent=1)
os.replace(tmp, path)
def append(path, obj):
with open(path, 'a') as f:
f.write(json.dumps(obj) + '\n')
def run(cmd, timeout=30):
"""(return code, stdout) of a command; never raises."""
try:
r = subprocess.run(cmd, capture_output=True, text=True, timeout=timeout)
return r.returncode, r.stdout.strip()
except (OSError, subprocess.TimeoutExpired) as e:
return 1, str(e)
def unit_active(unit):
return run(['systemctl', 'is-active', '--quiet', unit])[0] == 0
def skill_running():
return next((u for u in SKILL_UNITS if unit_active(u)), None)
class Ledger:
"""Numbered evidence. Every finding, action and diagnosis refers to these ids."""
def __init__(self):
self.items = []
def add(self, check, status, text, facts=None, args=None):
e = dict(id=f'E{len(self.items) + 1}', check=check, args=args or {}, status=status, text=text,
facts=facts or {}, t=round(time.time(), 1))
self.items.append(e)
return e
def get(self, check):
return next((e for e in reversed(self.items) if e['check'] == check), None)
def brief(self):
return ' '.join(f"{e['id']} {e['check']}: {e['text']}" for e in self.items)
selfcare/checks.py is 281 lines; read it in the repo. Every check is read-only and cheap enough to run next to a
driving robot. Each one adds one item to the ledger with a status of PASS, WARN or FAIL, one line of text, and the
raw numbers (facts).
| Check | What it measures | FAIL / WARN when |
|---|---|---|
services |
systemctl is-active for each unit in expected.json; failed rosorin* units; rosorin-nav NRestarts. |
FAIL: a unit is not running. |
host_ids (in ros_checks) |
Fast DDS host id of every ROS endpoint (bytes 2-3 of its GID), matched to robot processes by the low 16 bits of their pid. | FAIL: some processes carry another host id than the checker (chapter 9 explains why that breaks shared memory). |
topic_rates (in ros_checks) |
Messages counted for 4 s on /scan, /imu/data_raw, /joint_states, /wheel/odom, /odometry/filtered, /odom_vslam, /tf, /aurora/rgb/camera_info. |
FAIL: a topic below half its usual rate. |
nav_status (in ros_checks) |
The navigation supervisor's latched /nav/status (chapter 18). |
FAIL: no status; WARN: not ready. |
driver (in ros_checks) |
The board driver's state: wheels, e-stop, plugged. | FAIL: no driver state. |
cpu_by_process |
CPU of every process in a rosorin-* cgroup over 4 s, from /proc/<pid>/stat. |
FAIL: above max(usual x 1.6, usual + 30) %; WARN: load average above 18. |
cpu_by_thread |
Per thread of one process over 3 s: busy threads, kernel share, OpenBLAS loaded, its thread setting. Run by the playbook or the brain, not every round. | FAIL: a thread >= 40 % with >= 60 % in the kernel. |
network |
UDP packets received and receive-buffer overflows per second, from /proc/net/snmp. |
FAIL: > 8000 packets/s or > 20 overflows/s. |
resources |
Free disk in ~, MemAvailable, the hottest thermal zone. |
FAIL: < 15 GB, < 1200 MB, > 88 C. |
battery |
Median of the last minute of ~/battery/<day>.csv (written by the driver), trend over ~10 min, ~/battery/plugged.json. |
FAIL: unplugged at or below 10.9 V; WARN: no reading in 2 min. |
run_outcomes |
The summary line of the last 5 ~/explore/<run>/log.jsonl (chapter 24). |
FAIL: each of the last three failed more goals than it reached, or had an e-stop. |
journal_errors |
Error lines of one service since boot. Run by the brain only. | WARN: any. |
Two pieces are worth reading closely. First, which process belongs to which service: Linux puts every process
started by a systemd unit into that unit's control group, and /proc/<pid>/cgroup names it.
On your laptop:
def unit_of(pid):
try:
m = re.search(r'/([^/]+\.service)', open(f'/proc/{pid}/cgroup').read())
return m.group(1)[:-8] if m else None
except OSError:
return None
Second, the rule for "too much CPU": compared with this robot's own normal (base.usual), not with a fixed number.
On your laptop:
high = {}
for n, v in cpu.items():
usual = base.usual(f'cpu.{n}', exp['cpu'].get(n, exp['cpu_default']))
if v > max(usual * 1.6, usual + 30):
high[n] = (round(v), round(usual))
And the busy-wait signature from the idea box above, calibrated on 2026-10-05 against a reproduced OpenBLAS fault:
On your laptop:
spin = [h for h in hot if h['cpu'] >= 40 and h['kernel_share'] >= 0.6]
try:
blas = 'libopenblas' in open(f'/proc/{pid}/maps').read()
env = dict(x.split('=', 1) for x in open(f'/proc/{pid}/environ').read().split('\0') if '=' in x)
except OSError:
blas, env = False, {}
/proc/<pid>/maps lists every library a process has loaded; /proc/<pid>/environ its environment. Reading both
needs the same user, which is why self-care runs as burgerbarn like the services it watches.
If it fails
- A Python node is missing from the CPU check. Nodes started through ROS's installed wrapper run as
/usr/bin/python3 .../lib/<package>/<node>; the first version named them allpython3and skipped them,
so the injected fault invslam_odomwas invisible (2026-10-05).name_of()now takes the second word when
the first is a Python interpreter.- The camera's rate reads low. A best-effort reader of the 768 kB camera image loses frames with Fast DDS's
default 512 kB shared-memory segment, so the check counts the small/aurora/rgb/camera_infoinstead.
What is "usual" for this robot is learned from its own healthy rounds. Until a value has 30 samples, the seed from
expected.json is used; after that, the median of the last 300.
On your laptop:
"""What is normal for THIS robot: the median of its own healthy measurements (last 300 per value). Until a value has
30 samples the seed from expected.json is used."""
import os
from common import DATA, load, save
PATH = os.path.join(DATA, 'baseline.json')
class Baseline:
def __init__(self):
self.d = load(PATH, {})
On your laptop:
def usual(self, key, seed):
v = self.d.get(key, [])
return sorted(v)[len(v) // 2] if len(v) >= 30 else seed
def learn(self, values):
for k, x in values.items():
if isinstance(x, (int, float)):
self.d[k] = (self.d.get(k, []) + [round(float(x), 2)])[-300:]
save(PATH, self.d)
The median, not the mean: one bad round does not move it. The complete file, selfcare/baseline.py:
On your laptop:
"""What is normal for THIS robot: the median of its own healthy measurements (last 300 per value). Until a value has
30 samples the seed from expected.json is used."""
import os
from common import DATA, load, save
PATH = os.path.join(DATA, 'baseline.json')
class Baseline:
def __init__(self):
self.d = load(PATH, {})
def usual(self, key, seed):
v = self.d.get(key, [])
return sorted(v)[len(v) // 2] if len(v) >= 30 else seed
def learn(self, values):
for k, x in values.items():
if isinstance(x, (int, float)):
self.d[k] = (self.d.get(k, []) + [round(float(x), 2)])[-300:]
save(PATH, self.d)
The seeds were measured on 2026-10-05 after that day's fixes, robot standing, navigation ready, mind running. CPU
is in percent of one core. owner says which unit to restart when a topic goes quiet. units is the list the
services check watches.
On your laptop:
{
"_comment": "What is normal until the robot has its own history (>= 30 healthy snapshots per value). Measured 2026-10-05 after the fixes of that day, robot standing, navigation ready, mind running. cpu = % of one core. The camera is watched through its small camera_info topic: a best-effort reader of the 768 kB image loses frames with the default 512 kB shared-memory segment (scripts/tests/dds_large_data.py) and would read a false low rate.",
"cpu": {
"mind": 65,
"contact_monitor": 45,
"rosorin_board_driver": 45,
"aurora": 45,
"vslam_odom": 40,
"nav_supervisor": 15,
"ekf_filter_node": 20,
"robot_state_publisher": 8,
"amcl": 20,
"controller_server": 25,
"planner_server": 20,
"bt_navigator": 20,
"behavior_server": 20,
"smoother_server": 15,
"collision_monitor": 15,
"nvblox_node": 30,
"rf2o_laser_odometry": 12,
"robot_api": 15,
"rosorin_lidar_reader": 12,
"depth_gate": 12
},
"cpu_default": 25,
"rates": {
"/scan": 10,
"/imu/data_raw": 106,
"/joint_states": 10,
"/wheel/odom": 50,
"/odometry/filtered": 45,
"/odom_vslam": 9,
"/tf": 60,
"/aurora/rgb/camera_info": 15
},
"owner": {
"/scan": "rosorin-base",
"/imu/data_raw": "rosorin-base",
"/joint_states": "rosorin-base",
"/wheel/odom": "rosorin-base",
"/odometry/filtered": "rosorin-base",
"/tf": "rosorin-base",
"/odom_vslam": "rosorin-vslam",
"/aurora/rgb/camera_info": "rosorin-camera"
},
"limits": {
"udp_packets_per_s": 8000,
"rcvbuf_errors_per_s": 20,
"load1": 18,
"temp_c": 88,
"disk_free_gb": 15,
"ram_avail_mb": 1200,
"battery_ask_v": 10.9,
"battery_urgent_v": 10.4,
"lost_min": 15
},
"units": [
"rosorin-base",
"rosorin-camera",
"rosorin-vslam",
"rosorin-contact",
"rosorin-api",
"rosorin-mind",
"rosorin-nav"
]
}
selfcare/playbook.py is 184 lines; read it in the repo. Its header states the limits, and the code enforces them:
On your laptop:
"""Known patterns -> what the robot does about them. Every action is from a short allow-list, is logged with the
evidence it rests on, is verified by measuring again, and is rolled back when it did not help.
Never: wheels, goals, deleting data, anything outside its own services and its own learned settings.
Changes to services happen only while the robot is idle (wheels disabled, no wheel skill running)."""
import os, time
import checks, say
from common import DATA, DISABLED, Ledger, append, run, skill_running, unit_active
ALLOWED_UNITS = ('rosorin-base', 'rosorin-camera', 'rosorin-vslam', 'rosorin-contact', 'rosorin-api', 'rosorin-mind', 'rosorin-nav')
ACTIONS = os.path.join(DATA, 'actions.jsonl')
DRY = bool(os.environ.get('SELFCARE_DRY'))
def idle(L):
d = (L.get('driver') or {}).get('facts', {})
return d.get('enabled') is False and skill_running() is None
def allowed():
return not os.path.exists(DISABLED) and not DRY
idle() needs the driver to say the wheels are disabled (not merely "unknown") and no wheel unit to be running.
allowed() is false when the owner's switch is set or the round is a dry run. Every action goes through sudo -n
(no password prompt, fail instead) and is appended to ~/selfcare/data/actions.jsonl with the evidence ids it rests
on.
Restarting services has one trap: navigation depends on the base, and a navigation stack started within seconds of
the old one being killed fails to come up (chapter 18). So restarts follow one order, and navigation comes back no
sooner than 20 s after it was stopped:
On your laptop:
def restart_units(units):
"""Restart some of the robot's own services in a safe order. Navigation depends on the base and must not come
back within seconds of being stopped (docs/research/always_on_navigation_reference.md section 6)."""
units = [u for u in units if u in ALLOWED_UNITS]
nav = 'rosorin-nav' in units or ('rosorin-base' in units and unit_active('rosorin-nav'))
if nav:
sudo('systemctl', 'stop', 'rosorin-nav'); t_stop = time.time()
if 'rosorin-base' in units:
sudo('systemctl', 'restart', 'rosorin-base'); time.sleep(8)
rest = [u for u in units if u not in ('rosorin-base', 'rosorin-nav')]
if rest:
sudo('systemctl', 'restart', *rest); time.sleep(8)
if nav:
time.sleep(max(0.0, 20.0 - (time.time() - t_stop)))
sudo('systemctl', 'start', 'rosorin-nav')
return units
decide() goes through the evidence in this order:
| # | Pattern | What the robot does | Guard |
|---|---|---|---|
| 1 | host_ids FAIL: processes that do not share memory with the rest |
Restart those units, wait 25 s, check host ids and network again; if still split, queue "Some of my programs are not talking to each other properly..." | allowed and idle |
| 2 | A process far above its usual CPU, confirmed by a second measurement | cpu_by_thread on it; if spinning + OpenBLAS + not one thread: write drop-in 90-selfcare-openblas.conf with OPENBLAS_NUM_THREADS=1, restart, measure; roll back if not better |
allowed, idle, unit in the allow-list |
| 3 | One of its units is not running | reset-failed, start, check after 15 s; if it will not start, queue "One of my programs, ..., will not start." |
allowed; the base only while idle |
| 4 | A topic below half its usual rate, confirmed by a second measurement | Restart the unit that owns it (from expected.json), at most once per hour per unit, measure again |
allowed and idle |
| 5 | Battery unplugged at or below 10.9 V | Queue "My battery is getting low. Please plug me in." (urgent); at or below 10.4 V "My battery is very low. Please plug me in now." | cancelled when plugged in |
| 6 | Navigation status lost for more than 15 min |
Queue "I do not know where I am on my map. If you moved me, please put me back at my home spot." | |
| 7 | Heavy kernel network traffic without a host-id split; low resources; the last three drives failed; navigation restarted 3+ times since the last round | Report as a finding; resources and fresh failed drives are also said out loud; failed drives older than 2 h are only noted |
Rule 2 is the one that changes the system, so here it is whole:
On your laptop:
for name, (now, usual) in high.items():
t = checks.cpu_by_thread(L, name)
f = t['facts']; unit = f.get('unit')
if f.get('spinning') and f.get('openblas') and f.get('openblas_threads') != '1' and unit in ALLOWED_UNITS:
if not allowed() or not idle(L):
finding(f'{name} wastes CPU in idle maths-library threads', [c, t], 'deferred' if allowed() else 'not allowed', unit=unit); continue
d = f'/etc/systemd/system/{unit}.service.d'; conf = f'{d}/90-selfcare-openblas.conf'
log('openblas_one_thread', [c['id'], t['id']], unit=unit, file=conf)
sudo('mkdir', '-p', d)
run(['sudo', '-n', 'bash', '-c', f'printf "[Service]\\nEnvironment=OPENBLAS_NUM_THREADS=1\\n" > {conf}'])
sudo('systemctl', 'daemon-reload'); restart_units([unit]); time.sleep(20)
V = Ledger(); v = checks.cpu_by_process(V, exp, base)
after = v['facts']['cpu'].get(name, 0)
if after <= max(usual * 1.4, usual + 20):
finding(f'{name} wasted CPU in idle maths-library threads', [c, t], f'set its maths library to one thread and restarted {unit}',
'fixed', before=now, after=round(after))
else:
sudo('rm', '-f', conf); sudo('systemctl', 'daemon-reload'); restart_units([unit])
finding(f'{name} uses more CPU than usual', [c, t], 'tried one maths thread: no change, rolled back', 'not fixed', before=now, after=round(after))
else:
finding(f'{name} uses {now} % of a core (usual {usual} %)', [c, t])
The fix is a systemd drop-in (chapter 22 explains drop-ins), so it survives reboots and is removed by deleting one
file. The 90- prefix sorts it after any other drop-in, so its Environment= line wins.
Rule 4 limits itself to one restart per unit per hour with a trick: it uses the speech queue's repeat guard.
say.enqueue(f'restarted_{u}', '', repeat_s=3600, expires_s=0) returns True at most once an hour per unit, and the
entry has no text and expires at once, so nothing is ever spoken.
Findings that no rule explains (action none, result open) go to the brain: the same Gemma model on bigbuddy the
mind uses (chapter 21), reached at BRAIN_URL. It gets the numbered evidence and a closed menu of checks.
On your laptop:
"""For findings no rule explains: the brain (the model on bigbuddy the mind already uses) chooses which further
read-only check to run, one at a time, and names a cause citing evidence ids. Code runs the checks, refuses repeats
and unknown names, and throws away any cause that cites evidence that does not exist. The brain never acts.
Tested 2026-10-05 with the day's real symptoms: valid choices in ~1 s each (docs/research/doing_it_itself_reference.md)."""
import json, os
import checks
URL = os.environ.get('BRAIN_URL', 'http://192.168.1.111:1234/v1/chat/completions')
MODEL = os.environ.get('BRAIN_MODEL', 'google/gemma-4-12b')
MENU = {
'cpu_by_thread': 'for one process: which threads use the CPU and whether they are busy-waiting. args {"process": name}',
'journal_errors': 'error lines of one of my services since boot. args {"service": "rosorin-nav|rosorin-base|rosorin-camera|rosorin-vslam|rosorin-mind|rosorin-contact|rosorin-api"}',
'network': 'traffic between my own programs that goes through the kernel, and dropped packets',
'resources': 'disk, memory, temperature',
'run_outcomes': 'results of my last drives',
'battery': 'battery voltage, trend, plugged in or not',
'done': 'enough evidence. args {"cause": "<one sentence>", "evidence": ["E1", ...]}',
'unknown': 'the evidence does not show a cause',
}
SYS = ('You are the robot ROSOrin examining a problem with yourself. You get numbered evidence. Choose ONE next step from '
'this list only:\n' + '\n'.join(f'- {k}: {v}' for k, v in MENU.items()) +
'\nDo not repeat a check that is already in the evidence. Answer with JSON only: '
'{"check": name, "args": {...}, "why": "<short, citing evidence ids>"}')
The loop runs at most four steps. reasoning_effort: 'none' matters: without it the model spent its whole token
budget on hidden thinking and returned an empty answer (tested 2026-10-05). Code then enforces what the prompt only
asks: a done must cite evidence ids that exist, an unknown check or a repeat stops the investigation.
On your laptop:
def investigate(L, finding, exp, steps=4):
try:
import requests
except ImportError:
return None
done = set()
for _ in range(steps):
user = f"Problem: {finding['what']}. Evidence: {L.brief()}"[:6000]
try:
r = requests.post(URL, timeout=30, json={'model': MODEL, 'temperature': 0.2, 'max_tokens': 200, 'reasoning_effort': 'none',
'messages': [{'role': 'system', 'content': SYS}, {'role': 'user', 'content': user}]}).json()
t = r['choices'][0]['message']['content']; d = json.loads(t[t.index('{'):t.rindex('}') + 1])
except Exception as e:
return dict(cause=None, note=f'brain not usable: {type(e).__name__}')
c, a = d.get('check'), d.get('args') or {}
key = (c, json.dumps(a, sort_keys=True))
if c == 'done':
ids = {e['id'] for e in L.items}; cited = [x for x in a.get('evidence', []) if isinstance(x, str)]
if cited and all(x in ids for x in cited):
return dict(cause=str(a.get('cause', ''))[:240], evidence=cited)
return dict(cause=None, note='the brain named a cause without valid evidence')
if c == 'unknown' or c not in MENU or key in done:
return dict(cause=None, note='no cause found' if c == 'unknown' else f'stopped: {c} not usable or repeated')
done.add(key)
The chosen check is then run by code, with its arguments checked (a service name must be one of the robot's
units):
On your laptop:
if c == 'cpu_by_thread':
checks.cpu_by_thread(L, str(a.get('process', '')))
elif c == 'journal_errors':
s = str(a.get('service', ''))
if s in exp['units']: checks.journal_errors(L, s)
else: return dict(cause=None, note='stopped: unknown service')
elif c == 'network': checks.network(L, exp)
elif c == 'resources': checks.resources(L, exp)
elif c == 'run_outcomes': checks.run_outcomes(L)
elif c == 'battery': checks.battery(L, exp)
return dict(cause=None, note='no cause within the step limit')
The complete file, selfcare/investigate.py:
On your laptop:
"""For findings no rule explains: the brain (the model on bigbuddy the mind already uses) chooses which further
read-only check to run, one at a time, and names a cause citing evidence ids. Code runs the checks, refuses repeats
and unknown names, and throws away any cause that cites evidence that does not exist. The brain never acts.
Tested 2026-10-05 with the day's real symptoms: valid choices in ~1 s each (docs/research/doing_it_itself_reference.md)."""
import json, os
import checks
URL = os.environ.get('BRAIN_URL', 'http://192.168.1.111:1234/v1/chat/completions')
MODEL = os.environ.get('BRAIN_MODEL', 'google/gemma-4-12b')
MENU = {
'cpu_by_thread': 'for one process: which threads use the CPU and whether they are busy-waiting. args {"process": name}',
'journal_errors': 'error lines of one of my services since boot. args {"service": "rosorin-nav|rosorin-base|rosorin-camera|rosorin-vslam|rosorin-mind|rosorin-contact|rosorin-api"}',
'network': 'traffic between my own programs that goes through the kernel, and dropped packets',
'resources': 'disk, memory, temperature',
'run_outcomes': 'results of my last drives',
'battery': 'battery voltage, trend, plugged in or not',
'done': 'enough evidence. args {"cause": "<one sentence>", "evidence": ["E1", ...]}',
'unknown': 'the evidence does not show a cause',
}
SYS = ('You are the robot ROSOrin examining a problem with yourself. You get numbered evidence. Choose ONE next step from '
'this list only:\n' + '\n'.join(f'- {k}: {v}' for k, v in MENU.items()) +
'\nDo not repeat a check that is already in the evidence. Answer with JSON only: '
'{"check": name, "args": {...}, "why": "<short, citing evidence ids>"}')
def investigate(L, finding, exp, steps=4):
try:
import requests
except ImportError:
return None
done = set()
for _ in range(steps):
user = f"Problem: {finding['what']}. Evidence: {L.brief()}"[:6000]
try:
r = requests.post(URL, timeout=30, json={'model': MODEL, 'temperature': 0.2, 'max_tokens': 200, 'reasoning_effort': 'none',
'messages': [{'role': 'system', 'content': SYS}, {'role': 'user', 'content': user}]}).json()
t = r['choices'][0]['message']['content']; d = json.loads(t[t.index('{'):t.rindex('}') + 1])
except Exception as e:
return dict(cause=None, note=f'brain not usable: {type(e).__name__}')
c, a = d.get('check'), d.get('args') or {}
key = (c, json.dumps(a, sort_keys=True))
if c == 'done':
ids = {e['id'] for e in L.items}; cited = [x for x in a.get('evidence', []) if isinstance(x, str)]
if cited and all(x in ids for x in cited):
return dict(cause=str(a.get('cause', ''))[:240], evidence=cited)
return dict(cause=None, note='the brain named a cause without valid evidence')
if c == 'unknown' or c not in MENU or key in done:
return dict(cause=None, note='no cause found' if c == 'unknown' else f'stopped: {c} not usable or repeated')
done.add(key)
if c == 'cpu_by_thread':
checks.cpu_by_thread(L, str(a.get('process', '')))
elif c == 'journal_errors':
s = str(a.get('service', ''))
if s in exp['units']: checks.journal_errors(L, s)
else: return dict(cause=None, note='stopped: unknown service')
elif c == 'network': checks.network(L, exp)
elif c == 'resources': checks.resources(L, exp)
elif c == 'run_outcomes': checks.run_outcomes(L)
elif c == 'battery': checks.battery(L, exp)
return dict(cause=None, note='no cause within the step limit')
The robot never connects out, so it cannot call Buddy. Instead it keeps a queue in ~/selfcare/data/say_queue.json;
Buddy on bigbuddy asks for the next sentence through the robot API (GET /say) every 15 s, speaks it, and confirms
(POST /say/ack, chapter 22).
enqueue() adds a sentence under a key, unless the same key was queued or spoken within repeat_s (2 h by
default). A dry run queues nothing. Entries older than a day are dropped.
On your laptop:
"""What the robot wants to say out loud to its owner. A queue on the robot; Buddy (bigbuddy) polls it through the
robot API (GET /say, POST /say/ack) and speaks. The robot never connects out.
priority: urgent (always spoken) | normal (not during quiet hours 22:30-07:30)."""
import os, time
from common import DATA, load, save
PATH = os.path.join(DATA, 'say_queue.json')
def enqueue(key, text, priority='normal', repeat_s=7200, expires_s=1800):
"""Add a message unless the same key was queued or spoken within repeat_s. Returns True when added."""
if os.environ.get('SELFCARE_DRY'): # measure-only round: nothing is queued
return False
q = load(PATH, []); now = time.time()
q = [m for m in q if now - m['t'] < 86400]
last = max((m.get('said_t') or m['t'] for m in q if m['key'] == key), default=0)
if now - last < repeat_s:
save(PATH, q); return False
q.append(dict(id=f'{int(now * 1000)}', key=key, text=text, priority=priority, t=now, expires=now + expires_s, said_t=None))
save(PATH, q); return True
pending() picks what to say next: not yet said, not expired, and, between 22:30 and 07:30 local time, urgent only.
Urgent first, then oldest first. ack() marks a sentence as said.
On your laptop:
def cancel(key):
q = load(PATH, []); save(PATH, [m for m in q if not (m['key'] == key and m.get('said_t') is None)])
def pending():
"""The next message to speak, or None."""
now = time.time(); lt = time.localtime(now); quiet = (lt.tm_hour, lt.tm_min) >= (22, 30) or (lt.tm_hour, lt.tm_min) < (7, 30)
c = [m for m in load(PATH, []) if m.get('said_t') is None and m['expires'] > now and (m['priority'] == 'urgent' or not quiet)]
c.sort(key=lambda m: (m['priority'] != 'urgent', m['t']))
return c[0] if c else None
def ack(mid):
q = load(PATH, [])
for m in q:
if m['id'] == mid:
m['said_t'] = time.time()
save(PATH, q)
The complete file, selfcare/say.py:
On your laptop:
"""What the robot wants to say out loud to its owner. A queue on the robot; Buddy (bigbuddy) polls it through the
robot API (GET /say, POST /say/ack) and speaks. The robot never connects out.
priority: urgent (always spoken) | normal (not during quiet hours 22:30-07:30)."""
import os, time
from common import DATA, load, save
PATH = os.path.join(DATA, 'say_queue.json')
def enqueue(key, text, priority='normal', repeat_s=7200, expires_s=1800):
"""Add a message unless the same key was queued or spoken within repeat_s. Returns True when added."""
if os.environ.get('SELFCARE_DRY'): # measure-only round: nothing is queued
return False
q = load(PATH, []); now = time.time()
q = [m for m in q if now - m['t'] < 86400]
last = max((m.get('said_t') or m['t'] for m in q if m['key'] == key), default=0)
if now - last < repeat_s:
save(PATH, q); return False
q.append(dict(id=f'{int(now * 1000)}', key=key, text=text, priority=priority, t=now, expires=now + expires_s, said_t=None))
save(PATH, q); return True
def cancel(key):
q = load(PATH, []); save(PATH, [m for m in q if not (m['key'] == key and m.get('said_t') is None)])
def pending():
"""The next message to speak, or None."""
now = time.time(); lt = time.localtime(now); quiet = (lt.tm_hour, lt.tm_min) >= (22, 30) or (lt.tm_hour, lt.tm_min) < (7, 30)
c = [m for m in load(PATH, []) if m.get('said_t') is None and m['expires'] > now and (m['priority'] == 'urgent' or not quiet)]
c.sort(key=lambda m: (m['priority'] != 'urgent', m['t']))
return c[0] if c else None
def ack(mid):
q = load(PATH, [])
for m in q:
if m['id'] == mid:
m['said_t'] = time.time()
save(PATH, q)
Other parts of the robot use the same queue: the mind's camera watchdog (chapter 21) and the explorer's "I could
not get back home" (chapter 24).
Quiet hours run on the robot's clock, which is UTC
pending()usestime.localtime(). The robot's time zone isEtc/UTCand bigbuddy's is
America/Los_Angeles(both read withtimedatectlon 2026-10-07). So the quiet hours are 22:30-07:30 UTC,
which is 15:30-00:30 Pacific daylight time: normal sentences are held back in the afternoon and evening and
spoken from 00:30 at night. Urgent ones are always spoken. The project record does not mention this. Setting
the robot's time zone (sudo timedatectl set-timezone America/Los_Angeles) would move the quiet hours to the
owner's night, but it also changes every local timestamp the robot writes (self-care's daily report files,
room-scan folder names, the mind'slocal_time); that change has not been tried on this robot.
The order of a round: note whether the robot is driving, run every check, let the playbook decide, send what is left
to the brain.
On your laptop:
"""One self-care round (rosorin-selfcheck.service, every 15 min and after every wheel skill): measure, compare with
what is normal for this robot, act where a rule allows it, verify, ask the brain about what is left, tell the owner
what is his. Writes ~/selfcare/data/reports/<day>.jsonl and latest.json.
Usage: self_check.py [--dry] (measure and report only) Env SELFCARE_TEST_BATTERY="10.8,0" fakes a battery for tests."""
import json, os, sys, time
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
if '--dry' in sys.argv:
os.environ['SELFCARE_DRY'] = '1'
os.environ.setdefault('OPENBLAS_NUM_THREADS', '1')
import checks, playbook
from baseline import Baseline
from common import DATA, DISABLED, HERE, Ledger, append, load, save, skill_running
from investigate import investigate
--dry is turned into an environment variable before the other modules are imported, because playbook.py and
say.py read it when they load. The checker itself sets OPENBLAS_NUM_THREADS=1, so it cannot catch the fault it
looks for.
On your laptop:
def main():
t0 = time.time(); exp = load(os.path.join(HERE, 'expected.json'), {}); base = Baseline(); L = Ledger()
driving = skill_running()
checks.services(L, exp)
try:
checks.ros_checks(L, exp, base)
except Exception as e:
L.add('ros', 'FAIL', f'could not look at the ROS graph: {type(e).__name__}: {str(e)[:100]}')
checks.cpu_by_process(L, exp, base)
checks.network(L, exp)
checks.resources(L, exp)
checks.battery(L, exp)
checks.run_outcomes(L)
findings = playbook.decide(L, exp, base)
for f in findings: # what no rule explained goes to the brain (it only looks)
if f['action'] == 'none' and f['result'] == 'open':
f['brain'] = investigate(L, f, exp)
A round is healthy when every finding was only noted or not confirmed, and every check except the drive history
passed. Only a healthy round taken while standing teaches the baseline: a faulty or driving robot must not become
"normal".
On your laptop:
healthy = all(f['result'] in ('noted', 'not confirmed') for f in findings) and \
all(e['status'] == 'PASS' for e in L.items if e['check'] != 'run_outcomes')
if healthy and not driving: # normal is learned only from healthy, standing moments
vals = {f'cpu.{n}': v for n, v in L.get('cpu_by_process')['facts']['cpu'].items()}
vals.update({f'rate.{t}': r for t, r in (L.get('topic_rates') or {'facts': {'rates': {}}})['facts']['rates'].items()})
base.learn(vals)
else:
save(os.path.join(DATA, 'baseline.json'), base.d)
The report goes to a daily file and to latest.json (which the API's /selfcare serves), and a readable summary to
the journal:
On your laptop:
rep = dict(t=round(t0, 1), time=time.strftime('%Y-%m-%d %H:%M:%S'), took_s=round(time.time() - t0, 1), healthy=healthy,
driving=driving, disabled=os.path.exists(DISABLED), evidence=L.items, findings=findings)
os.makedirs(os.path.join(DATA, 'reports'), exist_ok=True)
append(os.path.join(DATA, 'reports', time.strftime('%Y%m%d') + '.jsonl'), rep)
save(os.path.join(DATA, 'latest.json'), rep)
print(f"self-check {rep['time']} ({rep['took_s']} s): " + ('healthy' if healthy else f'{len(findings)} finding(s)'))
for e in L.items:
print(f" {e['id']:>3s} {e['status']:4s} {e['check']}: {e['text']}"[:230])
for f in findings:
print(f" -> {f['what']} | {f['action']} | {f['result']}" + (f" | brain: {f['brain']}" if f.get('brain') else ''))
if __name__ == '__main__':
main()
The complete file, selfcare/self_check.py:
On your laptop:
"""One self-care round (rosorin-selfcheck.service, every 15 min and after every wheel skill): measure, compare with
what is normal for this robot, act where a rule allows it, verify, ask the brain about what is left, tell the owner
what is his. Writes ~/selfcare/data/reports/<day>.jsonl and latest.json.
Usage: self_check.py [--dry] (measure and report only) Env SELFCARE_TEST_BATTERY="10.8,0" fakes a battery for tests."""
import json, os, sys, time
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
if '--dry' in sys.argv:
os.environ['SELFCARE_DRY'] = '1'
os.environ.setdefault('OPENBLAS_NUM_THREADS', '1')
import checks, playbook
from baseline import Baseline
from common import DATA, DISABLED, HERE, Ledger, append, load, save, skill_running
from investigate import investigate
def main():
t0 = time.time(); exp = load(os.path.join(HERE, 'expected.json'), {}); base = Baseline(); L = Ledger()
driving = skill_running()
checks.services(L, exp)
try:
checks.ros_checks(L, exp, base)
except Exception as e:
L.add('ros', 'FAIL', f'could not look at the ROS graph: {type(e).__name__}: {str(e)[:100]}')
checks.cpu_by_process(L, exp, base)
checks.network(L, exp)
checks.resources(L, exp)
checks.battery(L, exp)
checks.run_outcomes(L)
findings = playbook.decide(L, exp, base)
for f in findings: # what no rule explained goes to the brain (it only looks)
if f['action'] == 'none' and f['result'] == 'open':
f['brain'] = investigate(L, f, exp)
healthy = all(f['result'] in ('noted', 'not confirmed') for f in findings) and \
all(e['status'] == 'PASS' for e in L.items if e['check'] != 'run_outcomes')
if healthy and not driving: # normal is learned only from healthy, standing moments
vals = {f'cpu.{n}': v for n, v in L.get('cpu_by_process')['facts']['cpu'].items()}
vals.update({f'rate.{t}': r for t, r in (L.get('topic_rates') or {'facts': {'rates': {}}})['facts']['rates'].items()})
base.learn(vals)
else:
save(os.path.join(DATA, 'baseline.json'), base.d)
rep = dict(t=round(t0, 1), time=time.strftime('%Y-%m-%d %H:%M:%S'), took_s=round(time.time() - t0, 1), healthy=healthy,
driving=driving, disabled=os.path.exists(DISABLED), evidence=L.items, findings=findings)
os.makedirs(os.path.join(DATA, 'reports'), exist_ok=True)
append(os.path.join(DATA, 'reports', time.strftime('%Y%m%d') + '.jsonl'), rep)
save(os.path.join(DATA, 'latest.json'), rep)
print(f"self-check {rep['time']} ({rep['took_s']} s): " + ('healthy' if healthy else f'{len(findings)} finding(s)'))
for e in L.items:
print(f" {e['id']:>3s} {e['status']:4s} {e['check']}: {e['text']}"[:230])
for f in findings:
print(f" -> {f['what']} | {f['action']} | {f['result']}" + (f" | brain: {f['brain']}" if f.get('brain') else ''))
if __name__ == '__main__':
main()
On your laptop:
cd ~/CCode/rosorin-pro
ssh rosorin-wifi 'mkdir -p ~/selfcare ~/setup'
rsync -a --exclude __pycache__ selfcare/ rosorin-wifi:selfcare/
scripts/ops/deploy.sh selfcare (chapter 29) copies later changes, but only selfcare/*.py: after a change to
expected.json, copy that file yourself.
A dry round measures, decides and reports, but changes nothing and queues nothing. Source ROS and the workspace the
same way the unit will:
On the robot:
source /opt/ros/humble/setup.bash && source ~/ros2_ws/install/setup.bash
cd ~/selfcare && python3 self_check.py --dry
Check
A round takes about 14 s and lists numbered evidence. The first round on the robot (2026-10-05 20:23 UTC, the
lines that fit the log):self-check 2026-10-05 20:23:49 (13.7 s): 1 finding(s) E1 PASS services: all the robot's services are running E2 PASS host_ids: all processes carry the same middleware host id (f4a0) E3 PASS topic_rates: all sensor and odometry topics at their usual rates E4 PASS nav_status: navigation ready, scan fits the map 0.68 E5 PASS driver: wheels disabled, e-stop None, plugged in E6 WARN cpu_by_process: every process near its usual CPU; load average 14.9, all robot processes 382 % E7 PASS network: 1013 UDP packets/s received, 0 kernel receive-buffer overflows/s (normal: a few hundred, 0)
If it fails
E? FAIL ros: could not look at the ROS graph: ROS was not sourced in this shell, orROS_DOMAIN_ID
differs from the robot's (0). The unit sets both.E? FAIL nav_status: navigation publishes no status:rosorin-nav(chapter 18) is not running yet.
Self-care works without it; that check fails until it is there.- Every rate reads low on a loaded robot: see the next section; that is why the unit is not at idle
priority.
systemd/rosorin-selfcheck.service:
On your laptop:
[Unit]
Description=ROSOrin self-care round: measure itself, fix what it may, ask for the rest (selfcare/self_check.py)
After=rosorin-base.service
[Service]
Type=oneshot
User=burgerbarn
Environment=ROS_DOMAIN_ID=0
# NOT idle scheduling: the first round at idle priority was starved, counted too few messages and restarted the
# camera for nothing (2026-10-05). A round costs ~3 s of CPU; a measuring tool must get to run.
Nice=5
TimeoutStartSec=300
ExecStart=/bin/bash -c 'source /opt/ros/humble/setup.bash && source /home/burgerbarn/ros2_ws/install/setup.bash && cd /home/burgerbarn/selfcare && exec python3 self_check.py'
systemd/rosorin-selfcheck.timer:
On your laptop:
[Unit]
Description=ROSOrin self-care round every 15 minutes (first one 5 minutes after boot)
[Timer]
OnBootSec=5min
OnUnitInactiveSec=15min
[Install]
WantedBy=timers.target
New idea: oneshot services and timers
AType=oneshotservice runs one program to the end and is then inactive again;systemctl startwaits until
it has finished (up toTimeoutStartSec). A timer unit starts the service of the same name on a schedule:
OnBootSec=5minonce, 5 minutes after boot;OnUnitInactiveSec=15min15 minutes after each round ended, so
rounds never overlap. The timer is what you enable; the service has no[Install]section.
Why Nice=5 and not idle priority
The design said self-work should run at idle CPU priority so it never slows driving. The first real round on
2026-10-05 ran that way on a loaded robot: it got so little CPU that it counted too few camera messages, read
"/aurora/rgb/image_raw 6.2 Hz (usual 14)", and restarted the camera for nothing. A measuring tool that cannot
run measures wrong.Nice=5is slightly below normal priority, and since the same day every topic restart
needs a second, confirming measurement.
The explorer script of chapter 24 also controls the timer's service: it stops rosorin-selfcheck at the start of
every drive (background self-work yields to driving) and starts one round after every drive.
The install script, scripts/install_selfcare.sh:
On your laptop:
#!/bin/bash
# The robot's self-care loop: rosorin-selfcheck.service + timer (docs/research/doing_it_itself_reference.md).
# Run on the robot: sudo bash ~/setup/install_selfcare.sh
# Off switch for the owner (nothing is changed any more, it still measures): touch ~/selfcare/DISABLED
# Rollback: sudo systemctl disable --now rosorin-selfcheck.timer; sudo rm /etc/systemd/system/rosorin-selfcheck.{service,timer};
# sudo rm -f /etc/systemd/system/rosorin-*.service.d/90-selfcare-*.conf; sudo systemctl daemon-reload
set -euo pipefail
[ "$(id -u)" = 0 ] || { echo "run with sudo"; exit 1; }
install -m 0644 /home/burgerbarn/setup/rosorin-selfcheck.service /home/burgerbarn/setup/rosorin-selfcheck.timer /etc/systemd/system/
systemctl daemon-reload
systemctl enable --now rosorin-selfcheck.timer
systemctl list-timers --no-pager | grep selfcheck
The rollback lines also remove every drop-in self-care ever wrote (90-selfcare-*.conf).
On your laptop:
cd ~/CCode/rosorin-pro
scp -q systemd/rosorin-selfcheck.service systemd/rosorin-selfcheck.timer scripts/install_selfcare.sh rosorin-wifi:setup/
On the robot:
sudo bash ~/setup/install_selfcare.sh
Check
The timer is enabled and listed. From the install on 2026-10-05 20:24 UTC:Created symlink /etc/systemd/system/timers.target.wants/rosorin-selfcheck.timer → /etc/systemd/system/rosorin-selfcheck.timer. n/a n/a Mon 2026-10-05 20:24:40 UTC 21ms ago rosorin-selfcheck.timer rLater,
systemctl list-timers rosorin-selfcheck.timer --no-pagershows the next round about 15 minutes
after the last one (read 2026-10-07):NEXT LEFT LAST PASSED UNIT ACTIVATES Wed 2026-10-07 20:45:27 UTC 14min left Wed 2026-10-07 20:30:04 UTC 1min 4s ago rosorin-selfcheck.timer rosorin-selfcheck.service
Read a round's output at any time:
On the robot:
journalctl -u rosorin-selfcheck -b --no-pager -o cat | grep -v -E "pam_unix|COMMAND=" | tail -20
tail -5 ~/selfcare/data/actions.jsonl
On the robot:
touch ~/selfcare/DISABLED
With the file present the robot still runs every round, measures, reports and investigates, but it restarts, starts
and changes nothing: those findings say not allowed (or deferred or not allowed), and latest.json carries
"disabled": true. Read from the code: the switch does not silence the speech queue, so it still asks to be plugged
in; only a dry run queues nothing. Remove or rename the file to let it act again.
The owner set the switch on 2026-10-05 at 20:42:38 UTC, 20 minutes after the loop went live. It stayed on until
2026-10-07 01:39 UTC, when the owner said "i need autonomy" and the file was moved aside:
On the robot:
mv ~/selfcare/DISABLED ~/selfcare/DISABLED.off-20261007
Check
ls ~/selfcareshowsDISABLED.off-20261007and noDISABLED(read from the robot 2026-10-07), and the newest
round saysdisabled False.
A self-care rule is only proven by making its fault happen. These are the tests run on 2026-10-05 between 20:27 and
20:33 UTC, in the order they were run.
Safety
The wheels must be disabled for all of this: the playbook acts only while idle, and these tests stop and restart
robot services. Put the robot on the stand or on the charger. Afterwards read back that every service is active
and the arm, mind mode and navigation are as before. On 2026-10-05 stopping one service for a test left the robot
with its head down for hours, because nobody read it back.
Read the starting state first:
On the robot:
source /opt/ros/humble/setup.bash
timeout 5 ros2 topic echo --once /rosorin_board_driver/state --field data | grep -o '"enabled": [a-z]*\|"estop": [a-z]*\|"plugged": [a-z]*' | tr "\n" " "; echo
systemctl is-active rosorin-base rosorin-camera rosorin-vslam rosorin-contact rosorin-api rosorin-mind rosorin-nav | paste -sd" "
Stop the contact monitor (chapter 19) and give the camera odometry node the OpenBLAS fault. vslam_odom.py sets
OPENBLAS_NUM_THREADS=1 itself with setdefault, so a value in the unit's environment overrides it; a drop-in
with 8 threads brings the fault back.
On the robot:
sudo systemctl stop rosorin-contact
sudo mkdir -p /etc/systemd/system/rosorin-vslam.service.d
printf "[Service]\nEnvironment=OPENBLAS_NUM_THREADS=8\n" | sudo tee /etc/systemd/system/rosorin-vslam.service.d/00-test-fault.conf >/dev/null
sudo systemctl daemon-reload; sudo systemctl restart rosorin-vslam
sleep 30
top -b -n 2 -d 4 -w 120 | awk "/^top/{n++} n==2" | grep vslam_odom | awk "{print \"vslam_odom with the fault: \" \$9 \" %\"}"
Now let the robot look at itself, and read what it did:
On the robot:
sudo systemctl start rosorin-selfcheck
journalctl -u rosorin-selfcheck --since "-4min" --no-pager -o cat | grep -E "ACTION|->" | cut -c1-250
systemctl is-active rosorin-contact
ls /etc/systemd/system/rosorin-vslam.service.d/
top -b -n 2 -d 4 -w 120 | awk "/^top/{n++} n==2" | grep vslam_odom | awk "{print \"vslam_odom now: \" \$9 \" %\"}"
Check
The fault costs about 1.7 cores; the round starts the contact monitor and puts the odometry node back on one
maths thread. On 2026-10-05 it took two rounds, because the first one exposed the naming bug described under
checks.py; the lines from both rounds:vslam_odom with the fault: 168.4 % ACTION {'t': 1791232116.6, 'action': 'start_unit', 'evidence': ['E1'], 'unit': 'rosorin-contact'} -> rosorin-contact was not running | started rosorin-contact | fixed ACTION {'t': 1791232176.1, 'action': 'openblas_one_thread', 'evidence': ['E6', 'E11'], 'unit': 'rosorin-vslam', 'file': '/etc/systemd/system/rosorin-vslam.service.d/90-selfcare-openblas.conf'}The status record for the second round: found at 147 % against a usual 40 %, 3 threads busy-waiting with 80 %
of their time in the kernel, remedy applied, measured 29 % afterwards. The robot's fix is the file
90-selfcare-openblas.confnext to the test fault;lsshould list both (the cleanup that followed removed
both). With the current code one round should do both repairs; that has not been run.
Remove the test fault and the robot's fix, so the node is back on its own setting:
On the robot:
sudo rm -f /etc/systemd/system/rosorin-vslam.service.d/00-test-fault.conf /etc/systemd/system/rosorin-vslam.service.d/90-selfcare-openblas.conf
sudo rmdir /etc/systemd/system/rosorin-vslam.service.d 2>/dev/null
sudo systemctl daemon-reload; sudo systemctl restart rosorin-vslam
SELFCARE_TEST_BATTERY="volts,plugged" replaces the battery reading with a fake one that fell 0.2 V in the last
10 minutes. This is a real round (not dry), so the playbook may also act on anything real it finds. Buddy polls the
queue every 15 s and may speak the sentence before you delete it.
On the robot:
rm -f ~/selfcare/data/say_queue.json
source /opt/ros/humble/setup.bash; cd ~/selfcare
SELFCARE_TEST_BATTERY=10.8,0 python3 self_check.py 2>&1 | grep -E "battery|->" | cut -c1-200
python3 -c 'import say; m = say.pending(); print("queued:", m and (m["priority"], m["text"]))'
rm -f ~/selfcare/data/say_queue.json
Check
The battery check fails and an urgent request is queued. From 2026-10-05 20:30 UTC:E9 FAIL battery: 10.80 V, not plugged in, falling 1.33 V per hour, about 23 min to the 10.3 V limit -> topics looked slow once; a second measurement was normal | none needed | not confirmed -> battery low and not plugged in | asked the owner | waiting for the owner -> the last three drives failed | told the owner | open queued: ('urgent', 'My battery is getting low. Please plug me in.')The other
->lines depend on your robot's state at that moment. The last one changed the same night: failed
drives older than 2 hours are nownoted, not told.
This needs Buddy running on bigbuddy (chapter 26). Queue one true sentence and wait for Buddy to acknowledge it.
On the robot:
cd ~/selfcare && python3 -c 'import say; print(say.enqueue("voice_test", "This is the robot. I can now speak up when I need something, for example when my battery is low.", "urgent", repeat_s=10, expires_s=120))'
for i in $(seq 1 12); do sleep 4; s=$(python3 -c 'import json; q=json.load(open("data/say_queue.json")); print(q[-1].get("said_t"))'); [ "$s" != "None" ] && { echo "spoken and acknowledged after ~$((i*4)) s"; break; }; done
Check
Buddy says the sentence out of the robot's speaker, and the queue marks it as said. From 2026-10-05 20:31 UTC,
with Buddy's journal line from bigbuddy:True spoken and acknowledged after ~16 s 13:31:45 robot: This is the robot. I can now speak up when I need something, for example when my battery is low.
Give the investigator a finding no rule covers, in a measure-only run. It needs the brain on bigbuddy (chapter 25).
On the robot:
source /opt/ros/humble/setup.bash; cd ~/selfcare && SELFCARE_DRY=1 python3 -c "
import os, json
from common import Ledger, load, HERE
import checks, investigate
from baseline import Baseline
exp = load(os.path.join(HERE, 'expected.json'), {}); L = Ledger()
L.add('cpu_by_process', 'FAIL', 'mind 140 % of a core (usual 65 %); load average 12', {})
print(investigate.investigate(L, {'what': 'mind uses 140 % of a core (usual 65 %)'}, exp))
for e in L.items[1:]: print(' ', e['id'], e['check'], e['text'][:150])"
Check
The brain asks for one sensible check, code runs it, and when the brain asks for the same check again the code
stops it. No cause is invented. From 2026-10-05 20:31 UTC:{'cause': None, 'note': 'stopped: cpu_by_thread not usable or repeated'} E2 cpu_by_thread mind: 1 busy threads 37 % (3 % in the kernel)
On the robot:
systemctl is-active rosorin-base rosorin-camera rosorin-vslam rosorin-contact rosorin-api rosorin-mind rosorin-nav | paste -sd" "
sudo systemctl start rosorin-selfcheck
journalctl -u rosorin-selfcheck --since "-60s" --no-pager -o cat | grep -E "self-check|->"
Check
Every service is active, and a normal round reports healthy. The final round on 2026-10-05 20:33 UTC:self-check 2026-10-05 20:33:02 (13.9 s): healthy -> the last three recorded drives failed (none in the last two hours) | none: old news | noted
If it fails
- The camera was restarted although it was fine: the round ran starved and counted too few frames
(2026-10-05). Check that the unit hasNice=5and thatexpected.jsonwatches/aurora/rgb/camera_info,
not the image.- The CPU fault is not found: the process is not named in the CPU table (the naming bug), or the robot was
not idle (wheels enabled or a wheel unit running: the finding then saysdeferred).brain not usable: ConnectionError: LM Studio on bigbuddy is not answering (chapter 25). The rest of
the round works without it.- Nothing happens at all, findings say
not allowed:~/selfcare/DISABLEDexists.
This round ran by itself on 2026-10-07 at 20:30 UTC (read from ~/selfcare/data/latest.json while writing this
chapter). It shows the loop at work, and one of its weak spots:
2026-10-07 20:30:26 21.2 not healthy disabled False
E1 PASS services: all the robot's services are running
E2 PASS host_ids: all processes carry the same middleware host id (f4a0)
E3 PASS topic_rates: all sensor and odometry topics at their usual rates
E4 PASS nav_status: navigation ready, scan fits the map 0.84
E5 PASS driver: wheels disabled, e-stop None, plugged in
E6 PASS cpu_by_process: every process near its usual CPU; load average 13.6, all robot processes 372 %
E7 FAIL network: 44614 UDP packets/s received, 11345 kernel receive-buffer overflows/s (normal: a few hundred, 0)
E8 PASS resources: disk 253 GB free, memory 10348 MB available, hottest sensor 63 C
E9 PASS battery: 12.67 V, plugged in
E10 PASS run_outcomes: last 2 drives: None m, 0 goals reached, 0 failed (None); 0.18 m, 0 goals reached, 12 failed (plugged in)
E11 WARN cpu_by_thread: no process named rosorin-nav
E12 PASS network: 891 UDP packets/s received, 0 kernel receive-buffer overflows/s (normal: a few hundred, 0)
E13 WARN cpu_by_thread: no process named rosorin-vslam
-> heavy traffic between my own programs through the kernel | none | open
A burst of kernel traffic (E7) with no host-id split matched no rule, so it went to the brain. The brain asked for
network again (E12: back to normal) and twice for cpu_by_thread on rosorin-nav and rosorin-vslam, which are
unit names, not process names; the code answered "no process named ..." and the investigation ended without a
cause. Nothing was changed. ~/selfcare/data/actions.jsonl held 4 lines, all from the 2026-10-05 tests: the
playbook had not needed to act since it was switched back on.
Not done
- The arm. On 2026-10-05 the claw servo (id 10) stopped answering at 19:44 UTC. The driver then refused to
wake the arm 2,907 times and the head stayed down for hours, while self-care reported healthy every round:
checks.pyhas no arm check. The driver's state carriesservo_silent(the read-back script prints it), but
no check reads it, and nothing tells the owner. The partial fix is in the driver (servo 10 is optional).- A muted or silent microphone. No check measures sound. If the array's microphone is muted or the stream
to bigbuddy carries silence, every check passes and Buddy is deaf without anyone being told. Chapter 27 has
the manual check.- Not tested: the host-id rule on a real split (it needs a network change while the robot runs), behaviour
while driving, and a whole night.- Not built: night experiments on its own recorded runs, the mind's own questions through the voice queue
(448 were waiting in a file on 2026-10-05), a stable outcome record per run.- The investigator confuses units and processes (the round above). Code keeps that harmless; it also keeps
it useless for that finding.
systemctl is-enabled rosorin-selfcheck.timer prints enabled and systemctl list-timers shows the next round.python3 self_check.py --dry) finishes in about 15 s with numbered evidence.~/selfcare/DISABLED does not exist (unless you want the robot measuring only).actions.jsonl logs start_unit.say.enqueue(...) is spoken by Buddy and acknowledged./etc/systemd/system/rosorin-vslam.service.d/.Where this comes from
selfcare/common.py,checks.py,baseline.py,expected.json,playbook.py,investigate.py,say.py,
self_check.py,systemd/rosorin-selfcheck.service,systemd/rosorin-selfcheck.timer,
scripts/install_selfcare.sh,scripts/explore_run.sh(stop/start around drives) and
ros2/rosorin_base/rosorin_base/vslam_odom.pyline 19 (setdefault), repo HEAD bfb61d8.
docs/research/doing_it_itself_reference.md(pattern, brain tests, design),docs/status.md2026-10-05 late
night (built, fault-injection results, not tested, not built) and 2026-10-05 end of day (claw servo, switch
renamed),docs/research/whole_project_review.mdfindings 12 and 19 (no arm check),docs/lessons.md
2026-10-05 (read back after a test; Fast DDS host ids). Command logcmdlog_robot.md: first round
2026-10-05T20:23:33, install 20:24:34, false camera restart 20:27:04, fault injection 20:27:38 and 20:29:11,
battery test 20:30:24, voice test 20:31:29, investigator 20:31:54, final round 20:32:45; DISABLED read and moved
2026-10-07T01:38:43 and 01:39:10. Read-only lookups on the robot 2026-10-07 20:31 UTC (latest.json,
actions.jsonl, timer,~/selfcarelisting).
← The robot API · Contents · Wheel skills - explore, practise, come, go home →