The robot's own mind · Chapter 22 · Time: 2 hours · Level: Intermediate · Status: Partly test-built
rosorin-api serves one token-protected HTTP port (8296) on the robot. Through it Buddy on bigbuddy reads the robot's state, sees its camera, hears its microphone, speaks through its speaker and starts its wheel tasks.
Buddy runs on bigbuddy; the robot's body runs on the Jetson. The robot API is the one door between them. It is a
small HTTP server inside a ROS 2 node (buddy_link/robot_api.py, run by rosorin-api.service). Connections go one
way only: bigbuddy connects to the robot and the robot never connects out. That way bigbuddy opens no ports, and
the robot has nothing to find. Every request carries a secret token that only the robot and bigbuddy hold.
New idea: an HTTP API
HTTP is the protocol your browser uses. A client opens a TCP connection to a server and sends a request: a
method (GETto read,POSTto send something or ask for an action), a path (/status), headers
(name/value lines such asX-Robot-Token: ...orContent-Type: application/json) and, for a POST, a body.
The server answers with a status code and its own body. The codes this API uses: 200 OK, 400 bad request,
401 no or wrong token, 404 unknown path, 409 busy or refused, 500 an error inside the server, 503 a device is
missing. An "API" is just the list of paths a program answers and what each one does.
New idea: a web server with a ROS node inside
robot_api.pyis two things in one process. A ROS node (robot_api) subscribes to topics and keeps only the
latest message of each in memory; one background thread runs the ROS executor so those callbacks keep firing.
AThreadingHTTPServeranswers each HTTP request in its own thread and reads the stored messages. Nothing is
fetched from ROS when a request arrives:/statereturns whatever the driver last published. Service calls
(the arm) go out withcall_asyncand the HTTP thread waits for the result.
Every request needs the header X-Robot-Token. Without it, or with a wrong value, every path answers 401
{"error": "missing or wrong X-Robot-Token"}. Any other exception inside a handler answers 500 with the error text.
| Method | Path | What it does | Special answers |
|---|---|---|---|
| GET | /state |
The board driver's last state message (wheels, e-stop, arm, gamepad mode), as published on /rosorin_board_driver/state. |
|
| GET | /status |
One-glance summary: battery_v, plugged, wheels, estop, arm, mind, mind_why, task, last (last = newest line of the newest ~/explore/<run>/log.jsonl). |
|
| GET | /objects |
Objects from the newest ~/vision/scans/*/scan.json, with bearing, distance, height, score, views. ?all=1 adds the dropped ones. |
{"scan": null, "objects": []} if no scan |
| GET | /frame.jpg |
The newest camera frame, JPEG quality 85. | 500 no camera frame (is rosorin-camera running?) |
| GET | /stream.mjpg |
Live camera as MJPEG, quality 70, at the camera's rate, while a client is connected. | |
| GET | /seen |
The mind's newest detections (/rosorin_mind/seen). |
{"dets": [], "t": null} before the first |
| GET | /roomscan, /roomscan.jpg |
The mind's newest room look-around: its meta.json, or its 3x2 mosaic. |
404 no room scan yet |
| GET | /task |
{active, last}: is any rosorin-task@ unit running, and the last run line. |
|
| GET | /say |
The next sentence the robot wants spoken (chapter 23). | {"say": null} when nothing waits |
| GET | /selfcare |
Summary of the newest self-care round (chapter 23). | {"time": null} before the first |
| GET | /audio |
The robot's ears: raw 16-bit mono 16 kHz audio from the microphone array, endless, while connected. | 409 if another client has it; 503 if the array is missing |
| POST | /look |
Aims the camera with the arm, returns a JPEG and an X-Aim header. Body: {"label": ...}, {"bearing_deg", "dist_m", "height_m"} or {"pan", "wrist"}. |
409 busy or refused; 404 label not in the newest scan |
| POST | /task/come, /task/home, /task/sethome |
Starts rosorin-task@<mode> (chapter 24). |
409 already driving; 404 unknown task |
| POST | /stop |
Stops rosorin-task@come, @home, @sethome and rosorin-explore. |
|
| POST | /play |
The robot's mouth: a WAV body is normalized and played; the call returns when playback ends. | 400 if the body is not a WAV; 503 no speaker |
| POST | /say/ack |
{"id": ...}: Buddy has spoken that sentence. |
|
| POST | /heard |
{"event": ..., "text": ...} with event wake, speech, inspect, learn or scan_room is published on /buddy/heard for the mind. No motion here. |
400 for any other event (its text says "event must be wake or speech", an old message) |
Who calls what, on bigbuddy: buddy_voice.py (Buddy's voice loop, chapter 26) uses /play, /say, /say/ack and
/heard; voice/robot_eyes.py (Buddy's robot tools) uses /heard, /seen, /frame.jpg, /task/*, /task,
/stop and /roomscan; robot-mic.service (chapter 27) holds /audio open all the time; the face server
(chapter 28) relays /status, /frame.jpg and /stream.mjpg to the touchscreen panel.
Safety
The token is permission to move the robot:POST /lookmoves the arm andPOST /task/...starts the wheels.
The API is plain HTTP on the LAN with no encryption, so anything that can read traffic on your network can read
the token. Keep the token file at mode 0600, never paste its contents into a chat, a commit or a log, and never
echothe variable that holds it.
The file is 459 lines. Read it in the repo at buddy_link/robot_api.py (github.com/burgerbarn/rosorin-pro, branch
rebuild, on your laptop at ~/CCode/rosorin-pro/buddy_link/robot_api.py). It is laid out top to bottom like this:
| Lines | Part | Job |
|---|---|---|
| 1-25 | Docstring | The endpoint list above, and the safety rule for the arm. |
| 26-62 | Imports and constants | Puts ~/selfcare and ~/vision on the import path; port, locks, PulseAudio environment, token path. |
| 65-111 | normalize_wav() |
The loudness chain for /play. |
| 114-119 | token() |
Creates the token on first start, then reads it. |
| 122-142 | aim_for(), latest_scan(), obj_out() |
Geometry for /look and /objects. |
| 145-226 | class Robot(Node) |
Subscriptions, the arm service clients, look() / _look(). |
| 229-248 | unit_active(), task_active(), last_run_line() |
Ask systemd and the run logs. |
| 251-441 | make_handler() |
The HTTP handler class: send, authed, do_GET, stream_audio, do_POST. |
| 444-459 | main() |
Token, ROS, executor thread, HTTP server. |
Two of the robot's own folders are put on Python's import path, so the API depends on them being on the robot:
~/selfcare/say.py (the speech queue, chapter 23) and ~/vision/mind_pause.py (handing the arm over from the
mind, chapter 21).
On your laptop:
import glob, json, math, os, secrets, subprocess, threading, time
import sys
sys.path.insert(0, os.path.expanduser('~/selfcare')) # the robot's self-care: speech queue + latest round
import say as selfcare_say
On your laptop:
import sys
sys.path.insert(0, os.path.expanduser('~/vision'))
from mind_pause import mind_paused # the robot mind hands the arm over for a look
PORT = int(os.environ.get('ROBOT_API_PORT', '8296'))
AUDIO_LOCK, PLAY_LOCK = threading.Lock(), threading.Lock()
PULSE_ENV = dict(os.environ, XDG_RUNTIME_DIR=os.environ.get('XDG_RUNTIME_DIR', f'/run/user/{os.getuid()}'),
PULSE_SERVER=os.environ.get('PULSE_SERVER', f'unix:/run/user/{os.getuid()}/pulse/native'))
TOKEN_FILE = os.path.expanduser('~/.config/rosorin/api_token')
SCANS = os.path.expanduser('~/vision/scans')
DRV = '/rosorin_board_driver'
HOME_PAN, WRIST_DEFAULT, CAM_Z = 500, 178, 0.33
PULSE_ENV matters: the service runs as burgerbarn without a login session, so it tells parec and paplay
where burgerbarn's PulseAudio socket is (/run/user/1000/pulse/native on the robot, uid 1000). That socket only
exists because linger is on for burgerbarn (chapter 6).
On your laptop:
class Robot(Node):
def __init__(self):
super().__init__('robot_api')
self.rgb, self.state = None, {}
self.create_subscription(Image, '/aurora/rgb/image_raw', lambda m: setattr(self, 'rgb', m), 2)
self.create_subscription(String, f'{DRV}/state', lambda m: self.state.update(json.loads(m.data)), 5)
self.cli = {n: self.create_client(T, f'{DRV}/{n}') for n, T in
(('arm/enable', SetBool), ('arm/recover', Trigger))}
self.cmd_lock = threading.Lock()
self.heard_pub = self.create_publisher(String, '/buddy/heard', 10)
self.battery_v, self.mind = None, {} # for the touchscreen robot panel (GET /status)
self.create_subscription(BatteryState, '/battery', lambda m: setattr(self, 'battery_v', round(m.voltage, 2)), 5)
self.create_subscription(String, '/rosorin_mind/mode', lambda m: self.mind.update(json.loads(m.data)),
QoSProfile(depth=1, durability=DurabilityPolicy.TRANSIENT_LOCAL))
self.seen = None # latest /rosorin_mind/seen (mind's detections)
self.create_subscription(String, '/rosorin_mind/seen', lambda m: setattr(self, 'seen', json.loads(m.data)), 2)
Each subscription callback only stores the message. The mind's mode topic is TRANSIENT_LOCAL ("latched"): a new
subscriber gets the last message at once, so /status knows the mind's mode right after the API starts.
main() ties it together: token first, then ROS, the executor in a daemon thread, then the HTTP server on every
interface.
On your laptop:
def main():
tok = token()
rclpy.init()
robot = Robot()
ex = SingleThreadedExecutor(); ex.add_node(robot)
threading.Thread(target=ex.spin, daemon=True).start()
srv = ThreadingHTTPServer(('0.0.0.0', PORT), make_handler(robot, tok))
robot.get_logger().info(f'robot API on :{PORT}')
try:
srv.serve_forever()
finally:
rclpy.shutdown()
The API does not drive. It asks systemd to start one of the task units from chapter 24 and returns at once
(--no-block). It refuses if any task or an exploration is already running. This needs burgerbarn's passwordless
sudo from chapter 6.
On your laptop:
if self.path.split('?')[0].startswith('/task/'):
mode = self.path.split('?')[0].split('/')[-1]
if mode not in TASKS:
return self.send(404, {'started': False, 'error': f'unknown task {mode}'})
if task_active() or unit_active('rosorin-explore'):
return self.send(409, {'started': False, 'error': 'already driving'})
r = subprocess.run(['sudo', 'systemctl', 'start', '--no-block', f'rosorin-task@{mode}'], capture_output=True, text=True)
return self.send(200 if r.returncode == 0 else 500, {'started': r.returncode == 0, 'error': r.stderr.strip()[:200]})
if self.path.split('?')[0] == '/stop':
r = subprocess.run(['sudo', 'systemctl', 'stop', '--no-block', 'rosorin-task@come', 'rosorin-task@home',
'rosorin-task@sethome', 'rosorin-explore'],
capture_output=True, text=True)
return self.send(200, {'ok': r.returncode == 0, 'error': r.stderr.strip()[:200]})
Stopping a unit is enough to stop the wheels: every wheel unit disables the wheels in its ExecStopPost
(chapter 24). /stop does not list rosorin-practice; the mind's own stop words do stop it (chapter 24).
/look moves only the arm, never the wheels. It pauses the mind so two programs never command the arm at once,
refuses while the wheels are enabled or an e-stop is latched, creates the arm/goal publisher only for the length
of the move (the driver e-stops when it sees more than one publisher on that topic), waits for a camera frame taken
0.4 s after the arm arrived, and disables the arm afterwards so the driver's auto-tuck takes it home after 2 minutes.
On your laptop:
def look(self, pan, wrist):
if not self.cmd_lock.acquire(blocking=False):
raise RuntimeError('busy: another arm command is running')
try:
with mind_paused('Buddy look'):
return self._look(pan, wrist)
finally:
self.cmd_lock.release()
def _look(self, pan, wrist):
s = self.state
if s.get('enabled') or s.get('estop'):
raise RuntimeError(f'refused: wheels enabled or e-stop ({s.get("estop")})')
if not s.get('arm_holding'):
self.call('arm/recover', Trigger.Request())
r = self.call('arm/enable', SetBool.Request(data=True))
if not r.success:
raise RuntimeError('arm enable refused: ' + r.message)
pub = self.create_publisher(JointState, 'arm/goal', 10)
try:
self.wait(lambda: pub.get_subscription_count() >= 1, 3, 'driver not subscribed to arm/goal')
goal = {1: pan, 4: wrist}
pub.publish(JointState(name=[str(k) for k in goal], position=[float(v) for v in goal.values()]))
def arrived():
st = self.state; p = st.get('arm_pose') or {}
return not st.get('arm_moving') and all(abs(p.get(str(k), -9) - v) <= 1 for k, v in goal.items())
self.wait(lambda: arrived() or self.state.get('estop'), 20, 'arm move')
if self.state.get('estop'):
raise RuntimeError('e-stop during the move')
t_ok = self.get_clock().now().nanoseconds + 400_000_000 # settle, then a fresh frame
self.wait(lambda: self.rgb is not None and
self.rgb.header.stamp.sec * 1e9 + self.rgb.header.stamp.nanosec > t_ok, 5, 'fresh frame')
return self.jpeg()
finally:
self.destroy_publisher(pub)
try:
self.call('arm/enable', SetBool.Request(data=False))
except Exception:
pass
Not test-built: /look
The record of 2026-09-29 says "/look not yet exercised", and no successful/lookanswer is in the command
logs since. Buddy's tools do not call it. Two more things are unknown:aim_for()still uses the wrist model
measured on 2026-09-29 (178 pulses = -10.1 degrees, clamp 150..280), while the mind has shifted its rest pose by
-65 pulses since 2026-10-05, so whether abearing_degorlabellook points the camera at the right place
now is not known; and/objectsreads~/vision/scans/*/scan.jsonfrom the scripted scan that was ruled out on
2026-10-02, so its data may be days old (the answer'sage_stells you).
On your laptop:
def token():
if not os.path.exists(TOKEN_FILE):
os.makedirs(os.path.dirname(TOKEN_FILE), exist_ok=True)
fd = os.open(TOKEN_FILE, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
os.write(fd, secrets.token_urlsafe(24).encode()); os.close(fd)
return open(TOKEN_FILE).read().strip()
On the first start the API makes 24 random bytes, encodes them as 32 URL-safe characters, and writes them to
~/.config/rosorin/api_token. os.open(..., O_CREAT | O_EXCL, 0o600) creates the file already at mode 0600, so it
is never readable by others, not even for a moment. Later starts read the same file, so the token survives restarts
and reboots. To make a new token, stop the service, delete the file, start the service, and copy the new file to
bigbuddy again.
Every request is checked with a constant-time comparison:
On your laptop:
def authed(self):
if secrets.compare_digest(self.headers.get('X-Robot-Token', ''), tok):
return True
self.send(401, {'error': 'missing or wrong X-Robot-Token'}); return False
secrets.compare_digest takes the same time whether the first or the last character is wrong, so the time of an
answer gives nothing away.
New idea: raw audio (PCM)
A microphone signal in a computer is a list of numbers, one per sample. 16 kHz means 16 000 samples per
second;s16lemeans each sample is a signed 16-bit integer, little-endian (2 bytes). With two channels
the samples alternate: left, right, left, right. So one second of 16 kHz mono s16le is exactly 32 000 bytes,
and 16 kHz stereo is 64 000. "Raw" means there is no file header: the bytes are only samples, and the receiver
must be told the format (audio/L16; rate=16000; channels=1here).
The microphone array (the reSpeaker XVF3800, chapter 27) gives two channels. The right one is the chip's own
speech-recognition output of the beam it picked, read from the device on 2026-10-05. The API reads the array
through burgerbarn's PulseAudio (not the raw ALSA card), so nothing else loses the card while Buddy listens.
On your laptop:
def stream_audio(self):
"""Raw s16le mono 16 kHz: the right USB channel of the reSpeaker XVF3800 (ASR output of the auto-selected
beam, read from the device 2026-10-05), taken from the robot's own PulseAudio so nothing else loses the
card. One listener at a time; parec ends when the listener goes."""
if not AUDIO_LOCK.acquire(blocking=False):
return self.send(409, {'error': 'another client has the audio stream'})
p = None
try:
src = subprocess.run(['pactl', 'list', 'short', 'sources'], capture_output=True, text=True, env=PULSE_ENV,
timeout=5).stdout
src = next((l.split('\t')[1] for l in src.splitlines() if 'alsa_input.usb-Seeed' in l), None)
if src is None:
return self.send(503, {'error': 'reSpeaker array not found'})
p = subprocess.Popen(['parec', f'--device={src}', '--rate=16000', '--channels=2', '--format=s16le',
'--raw', '--latency-msec=40'], stdout=subprocess.PIPE, env=PULSE_ENV)
self.send_response(200)
self.send_header('Content-Type', 'audio/L16; rate=16000; channels=1')
self.send_header('Cache-Control', 'no-cache'); self.end_headers()
while True:
buf = p.stdout.read(2560) # 640 stereo frames = 40 ms
if not buf:
break
mono = np.frombuffer(buf[:len(buf) // 4 * 4], np.int16)[1::2].tobytes()
self.wfile.write(mono); self.wfile.flush()
except (BrokenPipeError, ConnectionResetError):
return
finally:
if p is not None:
p.kill()
AUDIO_LOCK.release()
Read it line by line: one lock means one listener; pactl finds the array's PulseAudio source by name; parec
records stereo; each 2560-byte read is 640 stereo frames (640 x 2 channels x 2 bytes) = 40 ms; [1::2] keeps every
second 16-bit sample, which is the right channel; the result is written straight into the HTTP response. When the
client disconnects the write fails, parec is killed and the lock is freed.
New idea: dBFS, gain and a limiter
dBFS ("decibels relative to full scale") measures a digital level against the largest number a sample can
hold: 0 dBFS is the maximum, -6 dBFS is half of it, -20 dBFS a tenth. Gain multiplies every sample: +6 dB
doubles, -6 dB halves. A sample pushed past full scale is cut off flat (clipping), which sounds harsh. A
limiter prevents that by bending the loudest peaks down so they approach a ceiling but never cross it. A
compressor lifts the quiet parts more than the loud parts, so speech sounds evenly loud.
The array's own amplifier is small, so every clip Buddy sends is made as loud as it can be without clipping, the way
a desktop audio effects chain would, but in about 30 lines of numpy and with no desktop:
TONE_PEAK_DBFS = -20 dBFS (theSPEECH_TARGET_DBFS = -6 dBFS (measured on the loudSPEECH_MAX_BOOST_DB = 24 dB, and never so far that aSPEECH_GAIN_DB = 12 dB (the owner: "high gain").ceiling * tanh(a / ceiling) folds peaks into -1 dBFS instead of clipping them.SPEECH_OUT_DB = -2.75 dB (owner 2026-10-07: "voice down 10 %").Each number can be overridden with an environment variable (ROBOT_SPEECH_GAIN_DB and so on). The PulseAudio sink
stays at 100 %: on this array the sink volume is the ALSA control that chapter 27's level script resets, so step 6 is
the only volume knob for Buddy's voice.
On your laptop:
SPEECH_TARGET_DBFS = float(os.environ.get('ROBOT_SPEECH_TARGET_DBFS', '-6')) # loudness target (RMS of the loud parts)
SPEECH_MAX_BOOST_DB = float(os.environ.get('ROBOT_SPEECH_MAX_BOOST_DB', '24'))
TONE_PEAK_DBFS = float(os.environ.get('ROBOT_TONE_PEAK_DBFS', '-20')) # short clips (Buddy's chime / cue): peak level, owner 2026-10-06 "WAY down"
SPEECH_GAIN_DB = float(os.environ.get('ROBOT_SPEECH_GAIN_DB', '12')) # fixed gain stage before the limiter (owner: "high gain")
SPEECH_OUT_DB = float(os.environ.get('ROBOT_SPEECH_OUT_DB', '-2.75')) # after the limiter (owner 2026-10-07: voice down 10 % = PulseAudio's 90 %)
def normalize_wav(wav):
"""What an EasyEffects chain would do, in 30 lines, headless (owner 2026-10-06): the reSpeaker's small amp is the
ceiling, so the clip is made as loud as it can be without clipping - peak to -1 dBFS, then a slow compressor that
lifts the quieter parts of speech toward SPEECH_TARGET_DBFS (up to SPEECH_MAX_BOOST_DB), then a hard ceiling at
-1 dBFS. Returns (wav, gain applied in dB to the loud parts). Tones/chimes pass through with peak normalization."""
import io, wave
try:
w = wave.open(io.BytesIO(wav)); params = w.getparams(); frames = w.readframes(w.getnframes()); w.close()
if params.sampwidth != 2:
return wav, 0.0
a = np.frombuffer(frames, np.int16).astype(np.float32) / 32768.0
if len(a) < params.framerate // 10 or float(np.abs(a).max()) == 0.0:
return wav, 0.0
if len(a) < params.framerate * params.nchannels * 1.2: # a chime or cue, not speech: set its peak and stop
a = a * (10 ** (TONE_PEAK_DBFS / 20) / float(np.abs(a).max()))
out = io.BytesIO(); w2 = wave.open(out, 'wb'); w2.setparams(params)
w2.writeframes((a * 32767).astype(np.int16).tobytes()); w2.close()
return out.getvalue(), TONE_PEAK_DBFS
ceiling = 10 ** (-1 / 20)
a = a * (ceiling / float(np.abs(a).max())) # 1. peak at -1 dBFS
hop = max(1, params.framerate * params.nchannels // 50) # 2. 20 ms frames: level of the loud parts
n = len(a) // hop * hop
rms = np.sqrt((a[:n].reshape(-1, hop) ** 2).mean(1) + 1e-12)
loud = float(np.percentile(rms[rms > 0.01], 80)) if (rms > 0.01).any() else float(rms.max())
want = 10 ** (SPEECH_TARGET_DBFS / 20)
boost = min(10 ** (SPEECH_MAX_BOOST_DB / 20), max(1.0, want / loud))
# per-frame gain: full boost where the frame is quiet, less where it would exceed the ceiling; smoothed
g = np.minimum(boost, ceiling / np.maximum(rms * 3.0, 1e-6)) # 3 x RMS ~ speech peaks
g = np.convolve(np.r_[g[0], g, g[-1]], np.ones(5) / 5, mode='same')[1:-1]
gain = np.repeat(np.maximum(1.0, g), hop)
gain = np.r_[gain, np.full(len(a) - len(gain), gain[-1])] if len(gain) < len(a) else gain[:len(a)]
a = a * gain * 10 ** (SPEECH_GAIN_DB / 20) # 3. the gain stage (owner 2026-10-06: not conservative)
a = ceiling * np.tanh(a / ceiling) # 4. soft limiter: peaks fold into the ceiling, no hard clip
a = a * 10 ** (SPEECH_OUT_DB / 20) # 5. output level (the sink stays at 100 %: it is the ALSA control)
boost *= 10 ** ((SPEECH_GAIN_DB + SPEECH_OUT_DB) / 20)
out = io.BytesIO(); w2 = wave.open(out, 'wb'); w2.setparams(params)
w2.writeframes((a * 32767).astype(np.int16).tobytes()); w2.close()
return out.getvalue(), round(20 * math.log10(boost), 1)
except Exception:
return wav, 0.0
Any clip it cannot read (not 16-bit, under 0.1 s, silent, or an exception) is played unchanged.
The handler then plays the clip with paplay, once per output, at the same time: to the first USB speaker that is
not the array (the kit's WonderEcho box, if one is plugged in), and to the array itself. Playing through the array
gives the array's own echo canceller a copy of what comes out of the speaker, so it can take Buddy's voice back out
of the microphone signal. PLAY_LOCK makes clips play one after another, and the call returns only when playback
has ended.
On your laptop:
if self.path.split('?')[0] == '/play': # Buddy's voice out of the robot (2026-10-06)
n = int(self.headers.get('Content-Length', '0'))
wav = self.rfile.read(n) if n else b''
if not wav.startswith(b'RIFF'):
return self.send(400, {'error': 'WAV body expected'})
wav, gain_db = normalize_wav(wav) # the array's output is a headphone-class driver: use all of it
with PLAY_LOCK:
sinks = subprocess.run(['pactl', 'list', 'short', 'sinks'], capture_output=True, text=True,
env=PULSE_ENV, timeout=5).stdout
names = [l.split('\t')[1] for l in sinks.splitlines() if '\t' in l]
array = next((x for x in names if 'alsa_output.usb-Seeed' in x), None)
# the loud speaker (owner 2026-10-06): the kit's WonderEcho box (USB audio, 3 W amp) as speaker only;
# the same clip also goes to the reSpeaker array so its echo canceller has the reference
loud = next((x for x in names if 'alsa_output.usb-' in x and 'Seeed' not in x), None)
outs = [x for x in (loud, array) if x]
if not outs:
return self.send(503, {'error': 'no speaker found'})
procs = [subprocess.Popen(['paplay', f'--device={x}'], stdin=subprocess.PIPE, env=PULSE_ENV) for x in outs]
for pr in procs:
try:
pr.stdin.write(wav); pr.stdin.close()
except BrokenPipeError:
pass
rcs = [pr.wait(timeout=120) for pr in procs]
return self.send(200, {'played': all(rc == 0 for rc in rcs), 'bytes': n, 'gain_db': gain_db,
'speakers': [x.split('.')[1][:40] for x in outs]})
The docs differ from the code in two places; the code is what runs. docs/hardware.md says "Voice gain stage
11 dB" and a commit message says 10 dB, while the code and the robot's live copy (identical to the repo, per the
2026-10-07 survey) say 12 dB. docs/hardware.md also says the API sets the array's mic gain 135 at every start; the
code has no such step, and mic gain is set by scripts/ops/audio_levels.sh (chapter 27).
The API needs, already working: rosorin-base (chapter 11), rosorin-camera (chapter 15), the vision venv with
numpy and OpenCV (chapter 20), ~/vision/mind_pause.py (chapter 21), passwordless sudo and linger for burgerbarn
(chapter 6). /audio and /play also need the array (chapter 27): without it /audio answers 503, /play answers
503 unless some other USB speaker is plugged in, and everything else works.
The unit file, systemd/rosorin-api.service:
On your laptop:
[Unit]
Description=ROSOrin robot API for Buddy (state, objects, frame, look - arm only) on :8296
After=network-online.target rosorin-base.service rosorin-camera.service
Wants=network-online.target
[Service]
User=burgerbarn
Environment=ROS_DOMAIN_ID=0
ExecStart=/bin/bash -c 'source /opt/ros/humble/setup.bash && source /home/burgerbarn/ros2_ws/install/setup.bash && source /home/burgerbarn/vision/venv/bin/activate && exec python /home/burgerbarn/buddy_link/robot_api.py'
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
It sources ROS (for rclpy and the message types), the robot's workspace, and the vision venv (for numpy and cv2),
then execs Python so systemd's signals reach the Python process directly.
The install script, scripts/install_api_service.sh:
On your laptop:
#!/bin/bash
# Install + enable rosorin-api.service (robot API for Buddy, :8296, token in ~/.config/rosorin/api_token).
# Rollback: sudo bash ~/setup/rollback_api_service.sh
set -euo pipefail
[ "$(id -u)" = 0 ] || { echo "run with sudo"; exit 1; }
install -m 0644 /home/burgerbarn/setup/rosorin-api.service /etc/systemd/system/rosorin-api.service
systemctl daemon-reload
systemctl enable --now rosorin-api.service
sleep 8; systemctl --no-pager --lines=4 status rosorin-api.service | head -6
Copy the code, the self-care folder (the API imports its speech queue; chapter 23 explains it) and the unit with
its scripts to the robot. Nothing in the repo copies files into ~/setup; do it by hand:
On your laptop:
cd ~/CCode/rosorin-pro
python3 -m py_compile buddy_link/robot_api.py
ssh rosorin-wifi 'mkdir -p ~/buddy_link ~/setup ~/selfcare'
scp -q buddy_link/robot_api.py rosorin-wifi:buddy_link/
rsync -a --exclude __pycache__ selfcare/ rosorin-wifi:selfcare/
scp -q systemd/rosorin-api.service scripts/install_api_service.sh scripts/rollback_api_service.sh rosorin-wifi:setup/
Install and start it:
On the robot:
sudo bash ~/setup/install_api_service.sh
journalctl -u rosorin-api -b --no-pager -o cat | grep "robot API on" | tail -1
ls -la ~/.config/rosorin/api_token | cut -c1-10
Check
The service is enabled and running, the API logs its port, and the token file exists at mode 0600. From the
first install, 2026-09-29:Loaded: loaded (/etc/systemd/system/rosorin-api.service; enabled; vendor preset: enabled) Active: active (running) since Tue 2026-09-29 21:01:14 UTC; 8s ago [INFO] [1790715676.215096803] [robot_api]: robot API on :8296 -rw-------
If it fails
ModuleNotFoundError: No module named 'say'or'mind_pause'injournalctl -u rosorin-api: the API
imports~/selfcare/say.pyand~/vision/mind_pause.pyat start. Copy those folders (this follows from the
code; it is not a recorded failure).no camera frame (is rosorin-camera running?)from/frame.jpg: the API has received no image yet.
Checksystemctl is-active rosorin-cameraand chapter 15./task/...answers 500 with a sudo message: burgerbarn has no passwordless sudo (chapter 6). The API
returns the first 200 characters of sudo's error.- A restart hangs in
deactivatingfor up to 90 s: an open/audiostream. Seen on 2026-10-06 17:09
(systemctl is-activeprinteddeactivatingand the log showed a traceback from the ROS executor thread).
The next step fixes it.
bigbuddy's robot-mic.service keeps /audio open all the time. On 2026-10-06 a systemctl restart rosorin-api
then sat in deactivating: the stop did not finish, and systemd waits 90 s by default before it kills a service
that does not stop. The fix is a drop-in file that cuts that wait to 5 s. It is not in the repo; create it by hand
on the robot (these are the commands that were run on 2026-10-06 17:10 UTC):
On the robot:
sudo mkdir -p /etc/systemd/system/rosorin-api.service.d
printf "[Service]\n# the ears stream (/audio) keeps a connection open; do not wait 90 s on stop (2026-10-06)\nTimeoutStopSec=5\n" | sudo tee /etc/systemd/system/rosorin-api.service.d/10-fast-stop.conf >/dev/null
sudo systemctl daemon-reload
sudo systemctl restart rosorin-api
cat /etc/systemd/system/rosorin-api.service.d/10-fast-stop.conf
systemctl show rosorin-api -p TimeoutStopUSec -p KillMode | tr "\n" " "; echo
New idea: drop-ins
A file in/etc/systemd/system/<unit>.service.d/*.confadds to or overrides lines of the unit without editing
the unit file.systemctl cat rosorin-apishows the unit with all its drop-ins. Self-care (chapter 23) uses the
same mechanism for its own fixes.
Check
The file is there and systemd uses it (file read from the robot 2026-10-07;systemctl showfrom 2026-10-06):[Service] # the ears stream (/audio) keeps a connection open; do not wait 90 s on stop (2026-10-06) TimeoutStopSec=5 TimeoutStopUSec=5s KillMode=control-group
The deploy script for the API goes one step further. scripts/ops/deploy.sh api (chapter 29) copies the file,
compiles it, and then kills the service with SIGKILL before restarting it, because "an open /audio stream hangs a
plain stop":
On your laptop:
api) scp -q buddy_link/robot_api.py $R:buddy_link/; ssh $R 'python3 -m py_compile ~/buddy_link/robot_api.py && sudo systemctl kill -s SIGKILL rosorin-api 2>/dev/null; sudo systemctl restart rosorin-api' ;;
SIGKILL cannot be caught or delayed, so the old process is gone at once and the restart is immediate. bigbuddy's
robot-mic.service reconnects by itself (its loop runs curl again after 2 s).
bigbuddy keeps two copies of the same token. robot-mic.service reads ~/voice/robot_api_token (chapter 27);
buddy_voice.py and robot_eyes.py read ~/.config/rosorin/api_token (chapter 26). Both must be mode 0600. Copy
from your laptop so the value never appears on a screen (these are the commands used on 2026-09-29 and 2026-10-01):
On your laptop:
ssh bigbuddy 'mkdir -p ~/voice ~/.config/rosorin'
scp -3 -q rosorin-wifi:.config/rosorin/api_token bigbuddy:voice/robot_api_token
ssh bigbuddy 'chmod 600 ~/voice/robot_api_token'
ssh rosorin-wifi 'cat ~/.config/rosorin/api_token' | ssh bigbuddy 'bash -c "mkdir -p ~/.config/rosorin && umask 077 && cat > ~/.config/rosorin/api_token && chmod 600 ~/.config/rosorin/api_token && echo token installed"'
scp -3 copies between two remote hosts through your laptop. The last line pipes the file through ssh; only
token installed is printed.
bigbuddy's login shell is fish. Type bash first so every bigbuddy command in this chapter works as written.
On bigbuddy:
bash
stat -c '%a %s %n' ~/voice/robot_api_token ~/.config/rosorin/api_token
cmp ~/voice/robot_api_token ~/.config/rosorin/api_token && echo same
Check
Both files are mode 600, 32 bytes, and identical. The bigbuddy survey of 2026-10-07 recorded exactly that
(0600, 32 B, same content bycmp); the lines look like this:600 32 /home/burgerbarn/voice/robot_api_token 600 32 /home/burgerbarn/.config/rosorin/api_token same
If it fails
token installednever printed: the ssh from your laptop to one of the machines failed; check
ssh rosorin-wifi hostnameandssh bigbuddy hostname.- Buddy gets 401 after you made a new token on the robot: bigbuddy still has the old one. Copy both files
again and restartrobot-micandbuddy-voice(chapters 26-27).
Run these in bash on bigbuddy. The first line reads the token into a shell variable without printing it. The
robot answers on its Wi-Fi address, 192.168.1.108, which is the address Buddy uses.
On bigbuddy:
T=$(cat ~/.config/rosorin/api_token); R=http://192.168.1.108:8296
echo "no key -> $(curl -s -o /dev/null -w '%{http_code}' $R/state)"
curl -s -m 10 -H "X-Robot-Token: $T" $R/state | python3 -c "import json,sys; s=json.load(sys.stdin); print('state: wheels', s.get('enabled'), 'estop', s.get('estop'), 'arm_holding', s.get('arm_holding'), 'pad', s.get('pad_mode'))"
curl -s -m 10 -H "X-Robot-Token: $T" -o /tmp/robot_frame.jpg -w "frame.jpg -> HTTP %{http_code}, %{size_download} bytes\n" $R/frame.jpg
Check
No token gives 401; with the token you get the driver state and a camera frame. Output of the first test,
2026-09-29:no key -> 401 state: wheels False estop None arm_holding False pad locked frame.jpg -> HTTP 200, 71378 bytesOpen
/tmp/robot_frame.jpgon bigbuddy's screen to see what the camera sees.
-H "X-Robot-Token: $T" puts the token on curl's command line, where another user of bigbuddy could see it in the
process list for the moment the call runs. robot-mic.service avoids that by writing the header to a 0600 file
and passing -H @file. For a test on your own machine the variable is fine.
On bigbuddy:
curl -s -m 5 -H "X-Robot-Token: $T" $R/status; echo
curl -s -m 5 -H "X-Robot-Token: $T" $R/seen | head -c 300; echo
curl -s -m 5 -H "X-Robot-Token: $T" $R/say; echo
curl -s -m 5 -H "X-Robot-Token: $T" $R/selfcare | head -c 300; echo
Check
/statusanswers with one JSON line. A recorded answer (2026-10-02 03:12, robot unplugged, mind looking for
the owner):{"battery_v": 12.12, "plugged": false, "wheels": false, "estop": null, "arm": true, "mind": "curious", "mind_why": "lost you", "task": false, "last": {"reason": "localization overlay < 0.65", "t": 1790896446.48, "ev": "abort"}}
/seen,/sayand/selfcareanswer 200 with JSON. Their exact output with a token was not recorded; the
shapes come from the code:/sayis{"say": null}when nothing is waiting,/selfcarestarts with
{"time": "...", "healthy": ...}once self-care has run (chapter 23).
/audio serves one listener. If robot-mic.service from chapter 27 is already running on bigbuddy, it holds the
stream and your test gets 409 {"error": "another client has the audio stream"} (shape from the code). Stop it for
the test, measure 4 s of audio, start it again. Buddy is deaf for those seconds.
On bigbuddy:
systemctl --user stop robot-mic 2>/dev/null
n=$(timeout 4 curl -sN -H "X-Robot-Token: $T" $R/audio | wc -c); echo "bytes from /audio in ~4 s: $n (expect ~128000 = 32 kB/s)"
systemctl --user start robot-mic 2>/dev/null
Check
About 32 000 bytes per second of audio arrive. The first measurement, 2026-10-06 16:41 UTC:bytes from /audio in ~4 s: 105128 (expect ~128000 = 32 kB/s)
If it fails
{"error": "reSpeaker array not found"}(503): the array is not plugged in, or burgerbarn's PulseAudio
does not see it. The status note of 2026-10-06 evening records this answer while the array was on bigbuddy.
The same 503 comes when burgerbarn's PulseAudio is not running at all (linger off, chapter 6):pactlthen
lists nothing (from the code). Chapter 27 sets the array up.{"error": "another client has the audio stream"}(409):robot-micon bigbuddy still holds it; the
stopabove did not run in this shell, or something else is connected.
Make a 1-second test tone with Python's standard library and post it. The tone is built the same way as the
logged test clips (a 660 Hz sine at 0.2 of full scale), 1 s long so it has the same size as the chime that was
played on 2026-10-07.
On bigbuddy:
python3 - <<'EOF'
import math, struct, wave
sr = 16000
w = wave.open('/tmp/tone.wav', 'wb'); w.setnchannels(1); w.setsampwidth(2); w.setframerate(sr)
w.writeframes(b''.join(struct.pack('<h', int(0.2 * 32767 * math.sin(2 * math.pi * 660 * i / sr))) for i in range(sr)))
w.close()
EOF
curl -s -m 15 -H "X-Robot-Token: $T" -H "Content-Type: audio/wav" --data-binary @/tmp/tone.wav $R/play; echo
Check
You hear a quiet beep from the robot's speaker, and the answer says it played, at the chime level (-20 dBFS
peak), on the array. The answer for the 1 s chime (32 044 bytes) on 2026-10-07 01:20 UTC:{"played": true, "bytes": 32044, "gain_db": -20.0, "speakers": ["usb-Seeed_Studio_reSpeaker_XVF3800_4-Mic"]}A clip of 1.2 s or longer goes through the speech chain instead and
gain_dbreports the boost given to the
loud parts.
If it fails
- The call hangs and then fails;
paplaylogs "Failed to drain stream: Timeout": after a robot reboot on
2026-10-06 the array's playback stalled while capture still worked. A software USB unbind/bind made it worse
(full speed, 12 Mb/s); only a physical re-plug fixed it. Do not unbind the array.{"error": "no speaker found"}(503): no USB output in PulseAudio at all; see chapter 27.{"error": "WAV body expected"}(400): the body does not start withRIFF; check the file with
file /tmp/tone.wav.
On bigbuddy:
curl -s -m 3 -X POST -H "X-Robot-Token: $T" -H 'Content-Type: application/json' -d '{"event":"speech","text":"link test from bigbuddy"}' $R/heard; echo
On the robot:
grep '"heard"' ~/vision/mind/$(date -u +%Y%m%d)/events.jsonl | tail -1 | cut -c1-160
Check
The API answers{"ok": true}and the mind logs the event. From 2026-10-01 19:40 UTC:{"ok": true} {"t": 1790883614.109, "event": "heard", "mode": "track", "what": "speech", "text": "link test from bigbuddy"}Do not test with the words stop, stay, halt or freeze: the mind treats them as the owner's stop and holds
still for 30 minutes (chapter 21).
On bigbuddy:
timeout 5 curl -s -N -H "X-Robot-Token: $T" $R/stream.mjpg | grep -a -c -- "--frame" | awk '{printf "frames in 5 s: %s (%.1f fps)\n", $1, $1/5}'
Check
Roughly the camera's rate. This exact command was run against the face server's relay of this stream on
bigbuddy (2026-10-01), not against the robot directly:frames in 5 s: 64 (12.8 fps)
Leave /task and /look for later: /task/... starts the wheels (chapter 24) and /look moves the arm.
systemctl is-active rosorin-api prints active on the robot, and the API is still active after a reboot.curl without a token answers 401; with the token /status answers with the battery voltage and the mind's mode.systemctl show rosorin-api -p TimeoutStopUSec prints TimeoutStopUSec=5s./play comes out of the robot's speaker./audio deliver about 128 000 bytes.Where this comes from
buddy_link/robot_api.py(all code above, repo HEAD bfb61d8),systemd/rosorin-api.service,
scripts/install_api_service.sh,scripts/rollback_api_service.sh,scripts/ops/deploy.sh(api line),
vision/mind_pause.py,voice/robot_eyes.pyandvoice/buddy_voice.py(clients).docs/status.md
2026-10-06 entries (ears and mouth on the robot, 32 kB/s, chime from bigbuddy) and 2026-10-07 (voice down 10 %,
SPEECH_OUT_DB-2.75);docs/decisions.md2026-10-06 (direction Buddy to robot only);docs/hardware.md
2026-10-06/07 (playback controls, gain stage, sink = ALSA control);docs/lessons.md2026-10-06 night (playback
stall, re-plug). Command logcmdlog_robot.md: install 2026-09-29T21:01:05, first curl tests 2026-09-29T21:01:38,
token copy 2026-10-01T19:32:44,/heard2026-10-01T19:40:03,/status2026-10-01T23:36:38 and
2026-10-02T03:12:40,/audio2026-10-06T16:41, fast-stop drop-in 2026-10-06T17:10:26, chime 2026-10-07T01:20:27;
cmdlog_bigbuddy.mdstream relay 2026-10-01T23:47:00.sources/survey_live_robot.md(drop-in not in repo,
~/buddy_linkidentical to repo),sources/survey_bigbuddy.md(token files, robot-mic unit). Read-only lookup
on the robot 2026-10-07 (drop-in contents).