The brain PC · Chapter 26 · Time: 4 hours, plus 2-3 hours unattended if you train the wake word · Level: Advanced · Status: Partly test-built
Buddy's voice loop on bigbuddy - the "yo buddy" wake word, Whisper speech-to-text, Gemma with tools, Kokoro speech - installed as a user service, with the wake-word model copied from the repo or trained yourself.
Buddy is the voice assistant that lives on bigbuddy and talks through the robot. It waits for "yo buddy", records
what you say, turns it into text, asks the language model from chapter 25 (which can call tools: search the web,
move the robot, play music, switch lights), and speaks the answer. All of that runs on bigbuddy as one Python
program, voice/buddy_voice.py, started by a user service. The microphone and the speaker are on the robot; the
audio path between the two machines is chapter 27. This chapter builds the program and tests it without a
microphone.
Everything here was built on bigbuddy on 2026-09-28 and changed until 2026-10-06; the commands are the ones in the
command log. Two things are not test-built in the form given: recording your own wake-word clips through today's
microphone (the Komplete Audio 6 used on 2026-09-28 is no longer connected), and the speaker override used for a
test before chapter 27. Both are marked.
bash, not fish
bigbuddy's login shell is fish. Runbashbefore the@bigbuddyblocks (chapter 25).
New idea: a wake word
A wake word is a short phrase a small model listens for all the time, so the big models only run when you
address the robot. openWakeWord turns each 80 ms of audio into features (a mel spectrogram, then a pretrained
speech embedding of 96 numbers) and feeds the last 16 of those feature frames (the model input
(1, 16, 96)) into a tiny classifier trained only on your phrase. It outputs a score from 0 to 1 every 80 ms.
Above the threshold (0.5 here), Buddy wakes. The classifier for "yo buddy" has 50,403 weights and is a 215 KB
file; it runs on the CPU.
New idea: speech-to-text
Speech-to-text (STT) turns recorded speech into words. Buddy uses Whisper large-v3-turbo through
faster-whisper on the GPU. A voice-activity detector (Silero VAD) first cuts out the parts with no speech,
which stops Whisper from inventing words out of noise.
New idea: text-to-speech
Text-to-speech (TTS) turns the reply into audio. Buddy uses Kokoro, a small voice model (a 325 MB ONNX file)
run with onnxruntime on the CPU. Text is first turned into phonemes (sound symbols) by espeak-ng; Kokoro turns the
phonemes into 24 kHz audio in the chosen voice,am_onyx. It runs about 11 times faster than real time on
bigbuddy's CPU, which leaves the GPU to Whisper and the language model.
New idea: tools for a language model
A language model can only write text. "Tools" let it ask the program to do things: Buddy sends a list of
tools (name, description, parameters as JSON schema) with every question. If the model decides it needs one,
it answers with a structuredtool_callsentry such asweb_search {"query": "..."}instead of text. Buddy
runs the tool, sends the result back, and the model writes the final answer. Buddy allows at most two tool
rounds per question.
Measured on 2026-09-28 (voice/README.md): wake scores 0.78-0.90, STT 0.1-0.5 s, language model 0.2-0.5 s,
service RAM 2.3 GB. Through the robot's microphone the owner's wake word scored 0.83
(milestone 2026-10-06, docs/status.md).
| Path | What | Size (live) |
|---|---|---|
~/voice/venv |
runtime Python 3.14: Whisper, openWakeWord, Kokoro | 2.9 GB |
~/voice/tts-venv |
Python 3.14 with Kokoro only, used to synthesize training clips | 371 MB |
~/voice/oww-train |
Python 3.11 for wake-word training (only if you train) | 7.7 GB |
~/voice/models/kokoro/ |
kokoro-v1.0.onnx, voices-v1.0.bin |
354 MB |
~/voice/models/wakeword/ |
yo_buddy_v2.onnx, yo_buddy_v3.onnx |
215 KB each |
~/voice/wakeword/ |
training scripts, data/ (17 GB), work/ (2.5 GB), real/ (50 MB) |
|
~/voice/*.py, ~/voice/face/ |
Buddy's code, copied from the repo's voice/ |
|
~/voice/speed |
Kokoro speaking speed, live 1.30 |
Why three Python environments
The runtime uses Fedora's Python 3.14 and the current kokoro-onnx 0.4.7, which needs numpy 2.0.2 or newer.
openWakeWord's training code pins old packages (numpy below 2, speechbrain below 1, scipy below 1.15 for
acoustics), and numpy below 2 has no build for Python 3.14, so training gets its own Python 3.11. Training
needs Kokoro too, but Kokoro cannot live in the numpy<2 environment, so it gets a third, small one. The owner's
rule from 2026-09-28: few moving parts, each component a pinned venv plus a systemd user service plus docs
plus an uninstall (docs/decisions.md, "Sustainability rule").
These are the commands from the log (2026-09-28, 09:39 to 12:55 local), in order. The log used pip -q and
piped through tail; both are left out here so you see the output.
On bigbuddy:
mkdir -p ~/voice
python3 -m venv ~/voice/venv
~/voice/venv/bin/pip install --upgrade pip
~/voice/venv/bin/pip install faster-whisper
~/voice/venv/bin/pip install "nvidia-cublas-cu12" "nvidia-cudnn-cu12==9.*"
~/voice/venv/bin/pip install --no-deps openwakeword==0.6.0
~/voice/venv/bin/pip install kokoro-onnx soundfile scipy tqdm requests
~/voice/venv/bin/pip install scikit-learn
~/voice/venv/bin/pip install websockets
What each line is for:
faster-whisper brings ctranslate2 (the engine that runs Whisper on the GPU), onnxruntime, av andhuggingface_hub.nvidia-cublas-cu12 and nvidia-cudnn-cu12 are NVIDIA's CUDA math libraries as pip wheels. ctranslate2 needsnvidia-cuda-nvrtc-cu12 comes along as a dependencyopenwakeword==0.6.0 --no-deps: its declared dependencies include tflite-runtime, which has no build forscikit-learn is imported by openWakeWord's package even though Buddy never uses it.kokoro-onnx brings espeakng-loader (a bundled espeak-ng) and phonemizer-fork.websockets is for voice/media_screen.py, which drives the kiosk browser (chapter 28).Not pinned
Except openWakeWord, the recorded commands do not pin versions, so a later install gets newer packages.
The versions below are what runs on bigbuddy. Installing with==pins to these exact versions is the
obvious way to reproduce it; that was not done or tested.
Check
~/voice/venv/bin/pip list 2>/dev/null | grep -i -E '^(faster|ctranslate|nvidia|openwake|kokoro|onnxruntime|scikit|numpy|scipy|soundfile|requests|websockets|huggingface|phonem|espeak)'
on bigbuddy (2026-10-07):ctranslate2 4.8.2 espeakng-loader 0.2.4 faster-whisper 1.2.1 huggingface_hub 1.33.0 kokoro-onnx 0.4.7 numpy 2.5.3 nvidia-cublas-cu12 12.9.2.10 nvidia-cuda-nvrtc-cu12 12.9.86 nvidia-cudnn-cu12 9.26.0.51 onnxruntime 1.30.0 openwakeword 0.6.0 phonemizer-fork 3.3.1 requests 2.34.2 scikit-learn 1.9.1 scipy 1.18.1 soundfile 0.14.0 websockets 17.1
If it fails
- pip prints
openwakeword 0.6.0 requires tflite-runtime<3,>=2.8.0; platform_system == "Linux", which is not installed.This is expected after--no-depsand harmless.FAIL openwakeword.model ModuleNotFoundError No module named 'sklearn'(2026-09-28): install
scikit-learn.
The wake-word classifier sits on top of three shared models (mel spectrogram, speech embedding, voice activity)
that openWakeWord downloads into its own package folder. This is how they got into the runtime venv on 2026-09-28:
On bigbuddy:
~/voice/venv/bin/python -c "import openwakeword.utils as u; u.download_models(model_names=['__none__'])"
ls ~/voice/venv/lib/python3.14/site-packages/openwakeword/resources/models/
model_names=['__none__'] means "no pretrained wake words, only the base models".
Check
Thelsprints:embedding_model.onnx embedding_model.tflite melspectrogram.onnx melspectrogram.tflite silero_vad.onnx
ctranslate2 looks for the CUDA libraries on the normal library path; the pip wheels put them inside the venv. Point
LD_LIBRARY_PATH at them and load the model once. This is the test from 2026-09-28 without the transcription
part:
On bigbuddy:
V=~/voice/venv; SP=$($V/bin/python -c "import site;print(site.getsitepackages()[0])")
export LD_LIBRARY_PATH=$SP/nvidia/cublas/lib:$SP/nvidia/cudnn/lib
$V/bin/python -c 'import time; from faster_whisper import WhisperModel; t=time.time(); m=WhisperModel("large-v3-turbo", device="cuda", compute_type="float16"); print("model load %.1fs"%(time.time()-t))'
The first run downloads the weights into ~/.cache/huggingface/hub/models--mobiuslabsgmbh--faster-whisper-large-v3-turbo.
Check
It prints a load time. Recorded on 2026-09-28 (docs/hardware.md, bigbuddy voice host): "large-v3-turbo
fp16: load 19 s, 60 s audio in 13.4 s (first run), accurate". Later loads are faster because the weights
are cached.
If it fails
Without theLD_LIBRARY_PATHline Whisper cannot find cuBLAS and cuDNN on the GPU. That is the reason the
launcherrun_buddy_voice.sh(step 5) exists; the exact error text was not recorded.
The runtime venv already has Kokoro. The training scripts need it in an environment of its own (the reason is in
the box above), so make it now; it is small. From the log, 2026-09-28:
On bigbuddy:
python3 -m venv ~/voice/tts-venv && ~/voice/tts-venv/bin/pip install --upgrade pip
~/voice/tts-venv/bin/pip install kokoro-onnx soundfile
~/voice/tts-venv/bin/pip install scipy
Check
~/voice/tts-venv/bin/pip list 2>/dev/null | grep -i -E "kokoro|onnxruntime|numpy|phonemizer|espeak"
printed on 2026-09-28 (before scipy was added):espeakng-loader 0.2.4 kokoro-onnx 0.4.7 numpy 2.5.3 onnxruntime 1.30.0 phonemizer-fork 3.3.1
Why Kokoro
Chatterbox was tried first and rejected: one of its dependencies (spacy-pkuseg) fails to build on Python
3.14, and on older Python it pins torch 2.6, which does not support the RTX 5070 Ti's Blackwell GPU (sm_120)
(docs/decisions.md2026-09-28). Its 6.6 GB venv was deleted. The officialkokoropackage pins numpy
1.26.4, which has no build for Python 3.14;kokoro-onnx0.4.7 needs no torch at all.
Kokoro's weights and voices are two files from the kokoro-onnx project's release model-files-v1.0. The exact
download from 2026-09-28:
On bigbuddy:
mkdir -p ~/voice/models/kokoro && cd ~/voice/models/kokoro
for f in kokoro-v1.0.onnx voices-v1.0.bin; do [ -s $f ] || curl -fsSL -o $f https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/$f; done
ls -la; sha256sum *
Check
Output on 2026-09-28:-rw-r--r--. 1 burgerbarn burgerbarn 325532387 Sep 28 09:49 kokoro-v1.0.onnx -rw-r--r--. 1 burgerbarn burgerbarn 28214398 Sep 28 09:49 voices-v1.0.bin 7d5df8ecf7d4b1878015a32686053fd0eebe2bc377234608764cc0ef3636a6c5 kokoro-v1.0.onnx bca610b8308e8d99f32e6fe4197e7ec01679264efed0cac9140fe9c29f1fbf7d voices-v1.0.binThe sizes and both hashes must match.
The speed test from 2026-09-28. It writes two sample files you can play with paplay:
On bigbuddy:
cd ~/voice && ~/voice/tts-venv/bin/python - <<'PY'
import time, soundfile as sf
from kokoro_onnx import Kokoro
t=time.time(); k=Kokoro("models/kokoro/kokoro-v1.0.onnx","models/kokoro/voices-v1.0.bin"); print("load %.1fs"%(time.time()-t))
print("voices:", len(k.get_voices()), [v for v in k.get_voices() if v.startswith(("am_","af_","bm_","bf_"))][:40])
text="Yo! I'm Buddy. Tell me what you'd like me to look at, and I'll shine a light on it."
for v in ("am_michael","af_heart"):
t=time.time(); a,sr=k.create(text, voice=v, speed=1.0, lang="en-us"); dt=time.time()-t
sf.write(f"tts_{v}.wav",a,sr); print("%s: %.2fs audio in %.2fs (x%.1f realtime)"%(v,len(a)/sr,dt,len(a)/sr/dt))
PY
Check
Output on 2026-09-28 (voice list shortened here):load 0.3s voices: 54 ['af_alloy', 'af_aoede', 'af_bella', 'af_heart', ... 'bm_george', 'bm_lewis'] am_michael: 4.69s audio in 0.43s (x10.8 realtime) af_heart: 4.18s audio in 0.38s (x11.0 realtime)The owner listened to ten voices (am_michael, am_adam, bm_george, af_heart, am_liam, am_eric, am_echo,
am_fenrir, am_puck, am_onyx) and picked am_onyx; runner-up am_michael (docs/hardware.md).
kokoro-onnx 0.4.7 has a memory leak: during wake-word training (step 7) it grew each synthesis worker to 3.6 GB
in about 430 clips. Every
time it turns text into phonemes it calls phonemizer.phonemize(), which builds a brand new espeak backend: it
copies libespeak-ng.so into a fresh temporary folder (never deleted) and loads it again (never unloaded). That
is about 8 MB of RAM and 648 KB of /tmp per sentence. Buddy speaks all day, so the live voice loop needs the
fix too (docs/lessons.md, "bigbuddy OOM").
The fix replaces one method of Kokoro's tokenizer with a version that keeps one espeak backend per language.
Build it up:
The module grabs Kokoro's tokenizer module and a cache of backends:
On bigbuddy:
from kokoro_onnx import tokenizer as _tok
_backends = {}
One backend per language, created the first time it is needed, with the same options phonemize() uses:
On bigbuddy:
def _backend(lang):
b = _backends.get(lang)
if b is None:
from phonemizer.backend import EspeakBackend
b = _backends[lang] = EspeakBackend(lang, preserve_punctuation=True, with_stress=True)
return b
The replacement for Tokenizer.phonemize: normalize the text, phonemize with the cached backend, keep only
symbols Kokoro knows:
On bigbuddy:
def _phonemize(self, text, lang='en-us', norm=True):
if norm:
text = _tok.Tokenizer.normalize_text(text)
phonemes = _backend(lang).phonemize([text], strip=False)[0]
phonemes = ''.join(filter(lambda p: p in self.vocab, phonemes))
return phonemes.strip()
apply() swaps the method in, once:
On bigbuddy:
def apply():
if _tok.Tokenizer.phonemize is not _phonemize:
_tok.Tokenizer._orig_phonemize = _tok.Tokenizer.phonemize
_tok.Tokenizer.phonemize = _phonemize
The complete file is voice/kokoro_patch.py (an identical copy is voice/wakeword/kokoro_gen/kokoro_patch.py).
You will copy it to ~/voice/kokoro_patch.py in step 3:
On bigbuddy:
"""Fix a per-call leak in kokoro-onnx 0.4.7 + phonemizer-fork 3.3.1 (found 2026-09-28 on bigbuddy):
kokoro_onnx.tokenizer.Tokenizer.phonemize() calls phonemizer.phonemize(), which builds a NEW EspeakBackend
every call; each one copies libespeak-ng.so into a fresh tempdir (never removed) and dlopens it again
(never unloaded): ~8 MB RAM + 648 KB /tmp per call. Patch: one EspeakBackend per language, reused.
Output is identical to phonemizer.phonemize(text, lang, preserve_punctuation=True, with_stress=True).
Use: import kokoro_patch; kokoro_patch.apply() (before synthesizing)."""
from kokoro_onnx import tokenizer as _tok
_backends = {}
def _backend(lang):
b = _backends.get(lang)
if b is None:
from phonemizer.backend import EspeakBackend
b = _backends[lang] = EspeakBackend(lang, preserve_punctuation=True, with_stress=True)
return b
def _phonemize(self, text, lang='en-us', norm=True):
if norm:
text = _tok.Tokenizer.normalize_text(text)
phonemes = _backend(lang).phonemize([text], strip=False)[0]
phonemes = ''.join(filter(lambda p: p in self.vocab, phonemes))
return phonemes.strip()
def apply():
if _tok.Tokenizer.phonemize is not _phonemize:
_tok.Tokenizer._orig_phonemize = _tok.Tokenizer.phonemize
_tok.Tokenizer.phonemize = _phonemize
Check
On 2026-09-28 the patched and unpatched phonemes were compared for four texts in two accents
(identical: True; "yo buddy" isjˈoʊ bˈʌdiin en-us andjˈəʊ bˈʌdiin en-gb). Over 800 clips the
largest worker process then stayed between 570 and 614 MB, with one temporary folder the whole time.
buddy_voice.py imports the other modules in voice/ when it needs them: face_server (the face, chapter 28),
robot_eyes (the robot's camera and tasks, chapters 22 and 24), music, media_screen, home, evo and modes
(media, lights, amplifier, chapter 28), and kokoro_patch. Copy all of them, the launcher, and the face's web
files. From your laptop:
On your laptop:
cd ~/CCode/rosorin-pro/voice && scp *.py run_buddy_voice.sh bigbuddy:voice/ && scp -r face bigbuddy:voice/
The record copied files one at a time as they changed (scp -q buddy_voice.py run_buddy_voice.sh bigbuddy:voice/);
the result is the same. On 2026-10-07 every voice file on bigbuddy matched the repo by md5, except
face_server.py, which listens on all interfaces on bigbuddy and on 127.0.0.1 in the repo (chapter 28).
Buddy reads the robot API's token from ~/.config/rosorin/api_token on bigbuddy (mode 0600; the same content as
the robot's file of that name, chapter 22). Nothing in the repo copies it.
Not test-built: copying the token
How the file got to bigbuddy is not recorded ("Manual; there is no command", survey of the robot software).
One way that never prints it, from your laptop:
ssh rosorin-wifi 'cat ~/.config/rosorin/api_token' | ssh bigbuddy 'bash -c "umask 077; mkdir -p ~/.config/rosorin; cat > ~/.config/rosorin/api_token"'.
Check withssh bigbuddy 'stat -c "%a %s" ~/.config/rosorin/api_token':600 32on bigbuddy today.
voice/buddy_voice.py is 711 lines. Read it with this map:
| Lines | Part |
|---|---|
| 1-56 | settings from environment variables, the mouth (_out), speaking speed |
| 59-94 | links to the robot: robot_say_poll (the robot asks Buddy to say something, chapter 23), tell_robot (wake and speech events, chapter 22) |
| 96-204 | the system prompt and the 25 tools |
| 205-335 | run_tool (what each tool does), web_search |
| 341-386 | Mic: the microphone reader |
| 389-456 | Voice: Kokoro, the chime and the follow-up cue |
| 459-495 | Brain: the language model with tools |
| 498-519 | music ducking |
| 521-537 | load_stt, the phantom filter, transcribe |
| 540-595 | record_command, the echo guard, respond |
| 598-707 | main: test modes, then the listening loop |
Every setting is an environment variable with a default in the code. The service's drop-ins (step 6) change three
of them.
| Variable | Code default | Live on bigbuddy | Meaning |
|---|---|---|---|
BUDDY_MIC |
buddy_aec_source |
robot_mic |
PipeWire source Buddy listens to |
BUDDY_SPK |
buddy_aec_sink |
default | speaker used when the robot does not answer |
BUDDY_MOUTH |
robot |
default | robot = post the audio to the robot's /play; evo = play on BUDDY_SPK |
BUDDY_WAKE_MODEL |
~/voice/models/wakeword/yo_buddy_v3.onnx |
default | wake-word model |
BUDDY_WAKE_TH |
0.5 |
default | wake threshold |
BUDDY_LLM_URL |
http://localhost:1234/v1/chat/completions |
default | LM Studio |
BUDDY_LLM_MODEL |
google/gemma-4-12b |
default | model name |
BUDDY_VOICE |
am_onyx |
default | Kokoro voice |
BUDDY_SPEED |
1.0 (only if ~/voice/speed is missing) |
file says 1.30 |
speaking speed, clamped 0.7-1.5 |
BUDDY_ROBOT_URL |
http://192.168.1.108:8296 |
default | robot API over Wi-Fi |
BUDDY_SEARCH_URL |
https://mattysearch.matttesch.com/search |
default | SearXNG on server |
BUDDY_DUCK_VOL |
20 |
default | music volume while you talk to Buddy |
BUDDY_STT_MODEL |
large-v3-turbo |
default | Whisper model |
BUDDY_SPEECH_DB |
6 |
12 |
dB above the room's noise floor that counts as speech |
BUDDY_FOLLOWUP_S |
2 |
default | seconds of listening after an answer without the wake word |
BUDDY_ECHO_TAIL_S |
0.6 |
1.0 |
seconds of microphone ignored after Buddy makes a sound |
Docs differ
voice/README.mdstill describes the 2026-09-28 setup: Komplete Audio 6 microphone, HDMI speaker, a 5 s
follow-up window, 0.6 s echo tail, speech at floor + 6 dB, and Whisper hotwords "Buddy, faceplay". The code
and drop-ins above are what runs.
Mic starts parec (PulseAudio's recorder, served by PipeWire) on the chosen source, 16 kHz mono 16-bit, and
reads it in a background thread, 1280 samples (80 ms) at a time, through a 4th-order 100 Hz high-pass filter.
Each frame is stored with the time it arrived. ignore_until(t) makes frame() throw away every frame that
arrived before t:
On bigbuddy:
class Mic:
"""parec raw 16 kHz mono s16 + 100 Hz high-pass, read continuously by a background thread (so nothing queues
in parec/PipeWire while Buddy speaks). Frames carry arrival time; ignore_until() drops Buddy's own voice by
time. (A pipe flush left ~seconds of queued speech server-side -> Buddy answered itself, 2026-09-28.)"""
def __init__(self):
self.p = subprocess.Popen(['parec', f'--device={MIC}', '--rate=16000', '--channels=1',
'--format=s16le', '--raw', '--latency-msec=40'], stdout=subprocess.PIPE)
self.sos = ss.butter(4, 100, btype='high', fs=SR, output='sos')
self.zi = ss.sosfilt_zi(self.sos) * 0.0
self.q = queue.Queue(maxsize=int(30 * SR / FRAME)) # 30 s cap
self.skip_until = 0.0
threading.Thread(target=self._reader, daemon=True).start()
def _reader(self):
while True:
b = self.p.stdout.read(FRAME * 2)
if len(b) < FRAME * 2:
self.q.put((time.time(), None)); return
x = np.frombuffer(b, np.int16).astype(np.float32)
y, self.zi = ss.sosfilt(self.sos, x, zi=self.zi)
item = (time.time(), np.clip(y, -32768, 32767).astype(np.int16))
try:
self.q.put_nowait(item)
except queue.Full: # consumer stalled: drop the oldest frame
self.q.get_nowait(); self.q.put_nowait(item)
def frame(self):
while True:
t, f = self.q.get()
if f is None:
raise RuntimeError('mic stream ended')
if t >= self.skip_until:
return f
Buddy answered himself (2026-09-28)
The first version flushed the pipe after each reply. That only emptied the operating system's pipe (about
1.8 s); PipeWire had more of Buddy's own voice queued, which arrived later, was heard as a follow-up question,
and was answered, in a loop. The fix is the design above: read continuously, stamp each frame, drop by time
(end of playback plus the echo tail), and a similarity guard on follow-ups (below)
(docs/lessons.md, "Voice loop answered itself").
The listening loop in main() keeps a 20-second history of frame levels (250 frames) for the room's noise floor,
runs every frame through openWakeWord, and also wakes on a tap on the face (chapter 28):
On bigbuddy:
score = oww.predict(f)[key]
tapped = getattr(face, 'listen_req', None) is not None and face.listen_req.is_set()
if score < WAKE_TH and not tapped:
continue
if tapped:
face.listen_req.clear()
floor = float(np.percentile(floors, 20))
log(f'{"TAP (touchscreen)" if tapped else "WAKE"} score {score:.2f} (room floor {floor:.1f} dBFS)')
tell_robot('wake')
ducked = duck_music()
voice.chime()
mic.ignore_until(time.time() + ECHO_TAIL_S) # skip the chime: via the robot it lands ~0.5-0.75 s after /play returns (2026-10-06)
cmd, followup = record_command(mic.frame, floor), False
The room floor is the 20th percentile of recent levels: the quietest fifth of the last 20 s. tell_robot('wake')
posts to the robot's /heard in a background thread, so the robot turns to look for you (chapter 21); an
unreachable robot never delays Buddy.
On bigbuddy:
def record_command(next_frame, floor_db, start_timeout=6.0):
"""Wait up to start_timeout s for speech, stop after 0.8 s of silence or 12 s total."""
frames, started, silent, t0, peak = [], False, 0, time.time(), -120.0
face.state('listening')
while True:
f = next_frame()
lvl = db(f); peak = max(peak, lvl)
face.level((lvl - floor_db) / 30)
loud = lvl > floor_db + SPEECH_DB
if started:
frames.append(f)
silent = 0 if loud else silent + 1
if silent * 0.08 >= 0.8 or len(frames) * 0.08 >= 12:
log(f'command {len(frames) * 0.08:.1f}s, peak {peak:.1f} dBFS (floor {floor_db:.1f})')
return np.concatenate(frames)
elif loud:
started, frames = True, [f]
elif time.time() - t0 > start_timeout:
log(f'no speech: peak {peak:.1f} dBFS, needed > {floor_db + SPEECH_DB:.1f}')
return None
A frame counts as speech when it is SPEECH_DB above the floor. Recording starts at the first loud frame and stops
after ten quiet frames in a row (0.8 s) or 12 s. If nothing loud comes within 6 s, it gives up. A live example
from 2026-10-07: no speech: peak -17.1 dBFS, needed > -12.7.
On bigbuddy:
def load_stt():
from faster_whisper import WhisperModel
return WhisperModel(E('BUDDY_STT_MODEL', 'large-v3-turbo'), device='cuda', compute_type='float16')
# Whisper invents these from noise/silence (video-outro phrases); in the follow-up window they made Buddy answer
# "you're welcome" to nobody (2026-10-01)
PHANTOMS = re.compile(r"^(thank you( (so much|very much|for watching))?|thanks( for watching)?|i'?ll see you next time"
r"|see you next time|voil[aà]|bye( bye)?|you|okay|ok|\.+|so|um+|uh+)[.!?]*$", re.I)
def transcribe(stt, audio_i16):
segs, _ = stt.transcribe(audio_i16.astype(np.float32) / 32768, language='en', beam_size=1,
vad_filter=True, condition_on_previous_text=False, # Silero VAD: noise is never transcribed (owner 2026-10-06)
hotwords=None) # 2026-10-06: the hotword text came back verbatim on silence ("Buddy, faceplay" -> music nobody asked for)
segs = [s for s in segs if s.no_speech_prob < 0.6] # Whisper's own "this was not speech"
return ' '.join(s.text.strip() for s in segs).strip()
Three guards against Whisper inventing words: Silero VAD removes non-speech before Whisper sees it, segments
Whisper itself rates as probably not speech (no_speech_prob >= 0.6) are dropped, and in the follow-up window
known phantom phrases are ignored. beam_size=1 and English only keep it fast.
On bigbuddy:
def echo_of(text, reply):
"""True if a heard follow-up is mostly Buddy's own last reply (echo that slipped through)."""
a, b = re.sub(r'[^a-z ]', '', text.lower()), re.sub(r'[^a-z ]', '', (reply or '').lower())
return bool(a and b) and difflib.SequenceMatcher(None, a, b[:len(a) + 20]).ratio() > 0.6
def respond(stt, brain, voice, audio, followup=False):
"""Returns True if something was answered."""
face.state('thinking')
t = time.time(); text = transcribe(stt, audio); t_stt = time.time() - t
log(f'heard ({t_stt:.1f}s){" [follow-up]" if followup else ""}: {text!r}')
last = brain.history[-1]['content'] if brain.history else ''
if followup and echo_of(text, last):
log('ignored: echo of my own reply'); return False
if followup and PHANTOMS.match(text.strip()):
log(f'ignored: likely Whisper phantom {text!r}'); return False
if len(text) >= 2:
tell_robot('speech', text) # not Buddy's own echo
if len(text) < 2 or (followup and len(text.split()) < 2): # stray noise in the follow-up window: stay quiet
if not followup:
voice.say("Sorry, I didn't catch that.")
return False
face.state('thinking', text)
t = time.time(); reply = brain.ask(text); t_llm = time.time() - t
log(f'reply ({t_llm:.1f}s): {reply!r}')
voice.say(reply)
return True
In a follow-up (no wake word), Buddy stays silent for anything that looks like his own last reply (more than 60 %
similar), a phantom phrase, or a single word. After the wake word, an empty transcript gets the spoken "didn't
catch that" line from the code above.
On bigbuddy:
class Brain:
def __init__(self):
self.history = []
self.used_media = False
self.started_play = False # a tool started a track/movie this turn -> wake word only
def ask(self, text):
now = time.strftime('%A %d %B %Y, %H:%M %Z')
msgs = ([{'role': 'system', 'content': f'{SYSTEM} Current date and time: {now}. Your speaking speed is {speed():.2f}.'}] + self.history[-8:]
+ [{'role': 'user', 'content': text}])
reply, self.used_media, self.started_play = '', False, False
for _ in range(3): # at most 2 tool rounds, then answer
r = requests.post(LLM_URL, json={'model': LLM_MODEL, 'messages': msgs, 'max_tokens': 250,
'temperature': 0.7, 'reasoning_effort': 'none', 'tools': TOOLS},
timeout=90)
r.raise_for_status()
m = r.json()['choices'][0]['message']
calls = m.get('tool_calls') or []
if not calls:
reply = (m.get('content') or '').strip()
break
msgs.append({'role': 'assistant', 'content': m.get('content') or '', 'tool_calls': calls})
for c in calls:
name = c['function']['name']
try:
args = json.loads(c['function'].get('arguments') or '{}')
log(f'tool {name}: {args}')
res = run_tool(name, args)
self.used_media |= name in MEDIA_TOOLS
self.started_play |= name in ('play_music', 'faceplay', 'play_movie') or (
name == 'media_control' and args.get('action') in ('resume', 'next', 'previous'))
except Exception as e:
res = f'{name} failed: {type(e).__name__}: {e}'
log(f' -> {str(res)[:120]}')
msgs.append({'role': 'tool', 'tool_call_id': c['id'], 'content': res})
self.history += [{'role': 'user', 'content': text}, {'role': 'assistant', 'content': reply}]
return reply or "Sorry, I lost my train of thought."
Each question carries the system prompt (who Buddy is, that it talks out loud in one or two sentences, which tools
it has, the date and time), the last 8 messages, and the question. The loop allows two rounds of tool calls; a
tool that throws becomes a "failed" result the model can talk about instead of crashing the loop.
A tool is described by a small helper:
On bigbuddy:
def _tool(name, desc, props=None, required=()):
return {'type': 'function', 'function': {'name': name, 'description': desc, 'parameters': {
'type': 'object', 'properties': props or {}, 'required': list(required)}}}
for example _tool('web_search', 'Search the web (private SearXNG) for current information.', {'query': {'type': 'string', 'description': 'search query'}}, ['query']). The 25 tools and where each is built:
| Tools | Implemented in | Chapter |
|---|---|---|
robot_look_around, robot_see, robot_come, robot_go_home, robot_set_home, robot_stop, teach_robot |
voice/robot_eyes.py, calling the robot API |
22, 24 |
web_search |
web_search() here: SearXNG JSON on server, top 5 results |
- |
play_music, faceplay, play_album, media_volume, media_control, now_playing |
voice/music.py (mpv) |
28 |
play_movie, set_video_style, fullscreen, tv_frame, screen_mode |
voice/media_screen.py (kiosk Chromium) |
28 |
lights, light_status, tv |
voice/home.py (Home Assistant on server) |
28 |
amp |
voice/evo.py (Evo 150 amplifier) |
28 |
room_mode |
voice/modes.py |
28 |
voice_speed |
writes ~/voice/speed |
- |
Buddy refused to move (2026-10-01)
The system prompt still said Buddy could not move or look at anything after the robot tools were added, and
Gemma refused robot tasks out loud. When you add a tool, update the prompt in the same change
(docs/lessons.md, "Buddy kept saying you're welcome").
On bigbuddy:
def render(self, text):
parts, sr = [], 24000
for s in [s for s in re.split(r'(?<=[.!?])\s+', text.strip()) if s]:
a, sr = self.k.create(s, voice=VOICE, speed=speed(), lang='en-us')
parts += [a, np.zeros(int(0.12 * sr), np.float32)]
return np.concatenate([np.zeros(int(self.LEAD_S * sr), np.float32)] + parts), sr
The reply is split into sentences, each synthesized by Kokoro with 0.12 s of silence after it, behind 0.4 s of
lead silence (once needed because an HDMI speaker dropped the first 0.3 s after idle), and played as one WAV.
Voice.__init__ applies kokoro_patch before loading Kokoro. The finished WAV goes to _out():
On bigbuddy:
def _out(wav_bytes):
"""Buddy's voice: the robot's speaker (reSpeaker XVF3800, whose echo canceller then has its reference), else SPK."""
if MOUTH == 'robot':
try:
import robot_eyes
r = requests.post(robot_eyes.ROBOT_URL + '/play', data=wav_bytes, headers=robot_eyes._h(),
timeout=120)
if r.ok:
return
log(f'robot mouth: {r.status_code} {r.text[:80]} - falling back to {SPK}')
except Exception as e:
log(f'robot mouth unreachable ({str(e)[:60]}) - falling back to {SPK}')
subprocess.run(['paplay', f'--device={SPK}'], input=wav_bytes, check=False)
The robot's /play returns when playback ends (chapter 27), so the echo tail starts at the right moment.
main() has three ways to run without a microphone (voice/README.md):
| Command | Does |
|---|---|
run_buddy_voice.sh --say "text" |
Kokoro only, then exit |
run_buddy_voice.sh --ask "question" |
language model with tools, then speak, then exit |
run_buddy_voice.sh --wav file_16k.wav |
a 16 kHz recording through wake word, recording, Whisper, model and voice |
--say and --ask return before the wake-word model and Whisper are loaded.
voice/run_buddy_voice.sh sets up the CUDA library path and the microphone gain, then replaces itself with Python.
The venv and its site-packages folder, and the CUDA libraries from step 1:
On bigbuddy:
V=/home/burgerbarn/voice/venv
SP=$($V/bin/python -c "import site; print(site.getsitepackages()[0])")
export LD_LIBRARY_PATH=$SP/nvidia/cublas/lib:$SP/nvidia/cudnn/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}
A +12 dB gain on buddy_aec_source, the echo-cancelled microphone from chapter 27. PipeWire forgets source
volumes when it restarts, so it is set at every start; the loop retries for up to 10 s in case PipeWire is still
starting:
On bigbuddy:
for i in 1 2 3 4 5 6 7 8 9 10; do pactl set-source-volume buddy_aec_source 12dB 2>/dev/null && break; sleep 1; done
exec replaces the shell with Python, so systemd's stop signal reaches Buddy directly:
On bigbuddy:
exec $V/bin/python /home/burgerbarn/voice/buddy_voice.py "$@"
The complete file, voice/run_buddy_voice.sh (copied to ~/voice/run_buddy_voice.sh in step 3):
On bigbuddy:
#!/bin/bash
# Launcher for buddy_voice.py (systemd --user buddy-voice.service). CUDA libs for faster-whisper come from
# the venv's nvidia-cublas/cudnn wheels.
V=/home/burgerbarn/voice/venv
SP=$($V/bin/python -c "import site; print(site.getsitepackages()[0])")
export LD_LIBRARY_PATH=$SP/nvidia/cublas/lib:$SP/nvidia/cudnn/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}
# mic gain AFTER the echo canceller (owner 2026-10-05, XVF3800 speech beam has no AGC: speech peaked at -22 dBFS;
# with +12 dB about -8 dBFS, no clipping). PipeWire does not keep this across restarts, so it is set at every start.
# 2026-10-06: 0 dB - robot_mic already carries +12 dB; +12 here too clipped the mic 9% of the time (owner: "very unstable")
for i in 1 2 3 4 5 6 7 8 9 10; do pactl set-source-volume buddy_aec_source 12dB 2>/dev/null && break; sleep 1; done
exec $V/bin/python /home/burgerbarn/voice/buddy_voice.py "$@"
The comment and the code disagree, and live Buddy does not use this source
The comment's last line says 0 dB; the command sets 12 dB. The 0 dB change was reverted on 2026-10-06 with
"Mic boosts restored to the owner's earlier setting" (docs/lessons.md) and the comment stayed. Since the
60-mic.confdrop-in (step 6) Buddy listens torobot_micdirectly, so this gain changes nothing Buddy hears;
robot_micitself is at 0 dB live. Chapter 27 covers the levels.
Several files spell out /home/burgerbarn: this launcher, buddy-voice.service, yo_buddy.yaml and
retrain.sh. With another user name, change them all.
The owner's first spoken test, 2026-09-28, once LM Studio's server was enabled (chapter 25):
On bigbuddy:
cd ~/voice && ./run_buddy_voice.sh --ask "yo buddy, introduce yourself in one sentence"
Check
The last line of output on 2026-09-28:12:57:40 reply: 'Hello, I am Buddy, your helpful little home robot with wheels and a robotic arm.'With a tool, the same day:
./run_buddy_voice.sh --ask "What is the weather in Los Angeles right now?"
loggedtool web_search: 'current weather in Los Angeles September 28 2026'and then a two-sentence reply,
about 1 s later. Ifbuddy_aec_sourcedoes not exist yet (before chapter 27), the launcher waits 10 s
before Python starts.
Where the reply is heard depends on the mouth. By default Buddy posts it to the robot's /play; that needs the
robot API (chapter 22) and the token from step 3. If the robot does not answer, Buddy falls back to paplay on
buddy_aec_sink, which exists only after chapter 27. Before that, nothing is heard; the reply: line still
proves the brain works.
Not test-built: hearing it on bigbuddy before chapter 27
The code readsBUDDY_MOUTHandBUDDY_SPKfrom the environment, so
BUDDY_MOUTH=evo BUDDY_SPK=$(pactl get-default-sink) ./run_buddy_voice.sh --say "Testing, one two three."
should play on bigbuddy's default speaker.BUDDY_MOUTH=evowas used live (as a drop-in, by
scripts/ops/buddy_ears.sh); this exact command line was not run.
voice/buddy-voice.service (identical to the live ~/.config/systemd/user/buddy-voice.service):
On bigbuddy:
# ~/.config/systemd/user/buddy-voice.service on bigbuddy. Needs PipeWire (user session) and
# lmstudio-server.service (LM Studio API on :1234, loads the model on first request).
[Unit]
Description=Buddy voice loop (yo buddy -> Whisper -> Gemma/LM Studio -> Kokoro on HDMI)
After=pipewire.service pipewire-pulse.service lmstudio-server.service
Wants=lmstudio-server.service
[Service]
ExecStart=/home/burgerbarn/voice/run_buddy_voice.sh
Restart=on-failure
RestartSec=5
MemoryMax=8G
MemorySwapMax=0
[Install]
WantedBy=default.target
After= the audio server and LM Studio, and Wants= LM Studio so starting Buddy starts it too.Restart=on-failure with 5 s between tries: a crash (for example the microphone stream ending) restarts Buddy.MemoryMax=8G and MemorySwapMax=0: if Buddy ever leaks, systemd kills Buddy, not the desktop.WantedBy=default.target plus linger (chapter 25): Buddy starts at boot without a login.Install it. From your laptop:
On your laptop:
scp ~/CCode/rosorin-pro/voice/buddy-voice.service bigbuddy:.config/systemd/user/buddy-voice.service
A drop-in is a small file in buddy-voice.service.d/ that adds to the unit without editing it. bigbuddy has three,
all written on 2026-10-06 with these exact commands:
On bigbuddy:
mkdir -p ~/.config/systemd/user/buddy-voice.service.d
printf "[Service]\nEnvironment=BUDDY_ECHO_TAIL_S=1.0\n" > ~/.config/systemd/user/buddy-voice.service.d/40-echo-tail.conf
printf "[Service]\nEnvironment=BUDDY_SPEECH_DB=12\n" > ~/.config/systemd/user/buddy-voice.service.d/50-speech-db.conf
printf "[Service]\nEnvironment=BUDDY_MIC=robot_mic\n" > ~/.config/systemd/user/buddy-voice.service.d/60-mic.conf
systemctl --user daemon-reload
| Drop-in | Why |
|---|---|
40-echo-tail.conf, BUDDY_ECHO_TAIL_S=1.0 |
Through the robot the chime reaches the microphone 0.45-0.75 s after /play returns. With a fixed 0.25 s skip, recordings began with the chime: Whisper wrote "BANG", "BELL RINGS", "Beep", and the owner's first words were cut. First set to 1.2 s, then 1.0 s (commit bd434d1, 2026-10-06 18:25). |
50-speech-db.conf, BUDDY_SPEECH_DB=12 |
At 18:21 that evening Buddy heard "BELL RINGS" after a wake, then "Buddy, faceplay" in a near-silent follow-up (peak -21.7 dBFS): the Whisper hotword text came back by itself and started music nobody asked for. Fix in one commit (24d900d): VAD on, no hotwords, speech threshold 12 dB above the floor. |
60-mic.conf, BUDDY_MIC=robot_mic |
Buddy listens to the robot's microphone stream directly instead of the echo-cancelled buddy_aec_source (chapter 27). Written 18:29 with no commit and no note of the reason. A fact from the code: Buddy's voice goes to the robot's /play, not to buddy_aec_sink, so bigbuddy's echo canceller no longer gets Buddy's speech as its reference. |
Wake threshold: docs say 0.4, live is 0.5
docs/runbook.mdanddocs/hardware.mdsay a drop-in20-wake.confsetsBUDDY_WAKE_TH=0.4, and
scripts/ops/readback.shprints the threshold from that drop-in (today it prints nothing). That file was written on 2026-10-06 13:08 and moved to~/voice/reverted-20261006/at
18:06 when the owner had the audio reverted to the ears-and-mouth milestone (commit 448b466: "wake threshold
back to default 0.5"). Every start since logs@ 0.5. This guide follows the live value: no wake drop-in.
Do not enable the service yet
Buddy reads its microphone withparec --device=robot_mic. Until chapter 27 creates that source, the code's
reader gets end-of-stream at once, raisesmic stream ended, and systemd restarts the service every 5 s,
loading Whisper onto the GPU each time (from the code; not seen in the record). Enable it at the end of
chapter 27:systemctl --user enable --now buddy-voice.service
Check
systemctl --user show buddy-voice -p Environmentprints, live:Environment=BUDDY_ECHO_TAIL_S=1.0 BUDDY_SPEECH_DB=12 BUDDY_MIC=robot_micOnce it runs (chapter 27),
journalctl --user -u buddy-voice -o cat | grep ready: | tail -1shows the
line from 2026-10-07:10:42:05 ready: wake yo_buddy_v3 @ 0.5, STT on GPU, LLM google/gemma-4-12b, voice am_onyx, mic robot_micand every wake logs a line such as
10:32:06 WAKE score 0.93 (room floor -26.5 dBFS).
If it fails
heard (0.0s): ''after a wake. The VAD found no speech in the recording, and Buddy says he did not
catch that. Live on 2026-10-07:command 7.8s, peak -3.9 dBFS (floor -24.5)thenheard (0.0s): ''.
Look at the levels in chapter 27 before changing thresholds.- Buddy answers "you're welcome" to nobody. Whisper turned noise into "Thank you" or "I'll see you next
time" in the follow-up window (2026-10-01). ThePHANTOMSfilter andno_speech_probguard handle the
known phrases; add new ones toPHANTOMS.Sorry, my brain isn't answering right now.spoken after a wake:Brain.askraised, usually because
LM Studio is down or answered an error (chapter 25).- Music starts by itself. Check the log for a
[follow-up]line with a short phrase; that is the
"Buddy, faceplay" case above.
You need ~/voice/models/wakeword/yo_buddy_v3.onnx. Copy it from the repo (option A, five minutes), or train it
the way it was made (option B, an afternoon). Option A is what runs on bigbuddy: the repo file and the live file
have the same md5.
On your laptop:
ssh bigbuddy 'mkdir -p ~/voice/models/wakeword'
scp ~/CCode/rosorin-pro/voice/models/yo_buddy_v3.onnx ~/CCode/rosorin-pro/voice/models/yo_buddy_v2.onnx bigbuddy:voice/models/wakeword/
ssh bigbuddy 'sha256sum ~/voice/models/wakeword/*.onnx'
Check
5cccf306f205ecc0addf358d8f6c114f3ba72ea542f4fe92002da5a21d94fa74 /home/burgerbarn/voice/models/wakeword/yo_buddy_v2.onnx a197fb540a5932a064e9afc0f832e86c1adb7c550df002e0149260ba1951f9e6 /home/burgerbarn/voice/models/wakeword/yo_buddy_v3.onnxBoth files are 215120 bytes. v2 is the rollback: start Buddy with
BUDDY_WAKE_MODELpointing at it.
This is what happened on 2026-09-28, in three rounds:
| Model | Trained on | Synthetic recall @0.5 | False activations, 10.7 h of general audio | Owner's real clips @0.5 |
|---|---|---|---|---|
| v1 | 20k + 2k Kokoro "yo buddy", 20k + 2k sound-alikes, ACAV100M negatives, AudioSet background, room echoes | 43.2 % | 0.09 per hour | 6 of 23 |
| v2 | v1 + 18 owner clips ×120 + 66 two-second chunks of the owner talking ×15 | 45.2 % | 0.09 per hour | 4 of 5 held out (0.89, 0.92, 0.13, 0.90, 0.76) |
| v3 | v2 + music and movie sound as heard through the echo canceller | 45.6 % | 0.00 per hour | 4 of 5 held out (0.91, 0.95, 0.30, 0.90, 0.83) |
v1 was synthetic only, and it missed most of the owner's real "yo buddy"s. v2 fixed that with a minute and a half
of the owner's voice. v3 removed the remaining false wakes and scored 0 triggers on 3.1 minutes of held-out music,
film and talk.
What each part needs:
| Part | Needs |
|---|---|
| v1 (synthetic) | this chapter, 17 GB of disk for data, a few hours |
| v2 (your voice) | a working microphone: after chapter 27 that is robot_mic |
| v3 (media negatives) | chapters 27 and 28 (the echo-cancelled source and the jukebox and video players) |
Memory and GPU
Clip generation runs on the CPU in parallel and can exhaust bigbuddy's 32 GB of RAM; training uses the GPU,
which Whisper and the language model already fill to about 12 of 16 GB. Both have taken bigbuddy down once.
Use the memory-capped launch andretrain.shexactly as given.
Python 3.11 from Fedora (dnf transaction 42; sudo needs your password):
On bigbuddy:
sudo dnf install -y python3.11 python3.11-devel
Then the venv and its packages, exactly as installed on 2026-09-28 (torch from PyTorch's CUDA 12.8 index, which
supports the RTX 5070 Ti; onnxscript was added after the first export failed):
On bigbuddy:
python3.11 -m venv ~/voice/oww-train && ~/voice/oww-train/bin/pip install --upgrade pip
V=~/voice/oww-train
$V/bin/pip install "torch>=2.7" "torchaudio>=2.7" --index-url https://download.pytorch.org/whl/cu128
$V/bin/pip install "numpy<2" scipy pyyaml tqdm mutagen pronouncing "speechbrain<1" audiomentations torch-audiomentations acoustics torchinfo "torchmetrics<1" onnx onnxruntime datasets soundfile kokoro-onnx
$V/bin/pip install --no-deps openwakeword==0.6.0
$V/bin/pip install "scipy>=1.13,<1.15"
$V/bin/pip install onnxscript
Check
The import test of 2026-09-28 printed, among others,ok torch 2.11.0+cu128,ok torchaudio 2.11.0+cu128,ok speechbrain 0.5.16,ok audiomentations 0.43.1,ok torch_audiomentations 0.12.0;
after the scipy pin,scipy 1.14.1 acoustics ok,openwakeword.data ok,utils ok; after onnxscript,
onnxscript 0.7.2 onnx 1.23.0. Thepkg_resources is deprecatedandtorchvision is not available
warnings are harmless.
If it fails
acousticsfails to import with scipy 1.15 or newer: pinscipy>=1.13,<1.15(the last line above).- Without
onnxscriptthe training finishes and then fails at export with
ModuleNotFoundError: No module named 'onnxscript'(2026-09-28): torch 2.11's ONNX export needs it.
On your laptop:
scp -r ~/CCode/rosorin-pro/voice/wakeword bigbuddy:voice/
That gives ~/voice/wakeword/ with prepare_data.py, yo_buddy.yaml, train.sh, train_wrapper.py,
add_real.py, eval_real.py, record_media_negatives.sh, retrain.sh and kokoro_gen/.
openWakeWord learns "yo buddy" against a large amount of audio that is not "yo buddy". voice/wakeword/prepare_data.py
fetches the openWakeWord author's published data into ~/voice/wakeword/data (about 17 GB):
models: openWakeWord's base models for the training venv.rir: MIT's room impulse responses (270 WAVs after conversion to 16 kHz mono). Convolving a clean clip withbackground: two shards of AudioSet (everyday sounds), converted to 16 kHz WAVs and mixed under the clips.features: precomputed openWakeWord features of 2,000 hours of general audio (ACAV100M, 17.28 GB): thevalidation_set_features.npy (0.18 GB, about 10.7 hours) to count false activations.On bigbuddy:
"""Download/convert training data for openWakeWord (runs on bigbuddy, venv ~/voice/oww-train).
Sources (openWakeWord author's published data + AudioSet): all go to ~/voice/wakeword/data (delete after training).
- davidscripka/openwakeword_features: ACAV100M negative features (17.3 GB) + validation_set_features.npy (0.18 GB)
- davidscripka/MIT_environmental_impulse_responses: room impulse responses (272 wavs)
- agkphysics/AudioSet bal_train shards 00-01 (~1.4 GB): background audio for augmentation
- openwakeword base models (melspectrogram / embedding onnx)"""
import io, os, sys
import numpy as np, soundfile as sf
from scipy.signal import resample_poly
from huggingface_hub import hf_hub_download
D = os.path.expanduser('~/voice/wakeword/data')
os.makedirs(D, exist_ok=True)
def to16k(a, sr):
if a.ndim > 1:
a = a.mean(axis=1)
if sr != 16000:
g = np.gcd(int(sr), 16000)
a = resample_poly(a, 16000 // g, int(sr) // g)
return (np.clip(a, -1, 1) * 32767).astype(np.int16)
step = sys.argv[1] if len(sys.argv) > 1 else 'all'
if step in ('all', 'models'):
import openwakeword.utils as u
u.download_models(model_names=['__none__']) # base feature models only
print('base models ok', flush=True)
if step in ('all', 'rir'):
out = os.path.join(D, 'rir'); os.makedirs(out, exist_ok=True)
from huggingface_hub import snapshot_download
p = snapshot_download('davidscripka/MIT_environmental_impulse_responses', repo_type='dataset',
local_dir=os.path.join(D, 'rir_src'))
n = 0
for root, _, files in os.walk(p):
for f in files:
if f.lower().endswith('.wav'):
a, sr = sf.read(os.path.join(root, f), dtype='float32')
sf.write(os.path.join(out, f), to16k(a, sr), 16000, subtype='PCM_16'); n += 1
print('rir wavs', n, flush=True)
if step in ('all', 'background'):
import pyarrow.parquet as pq
out = os.path.join(D, 'background'); os.makedirs(out, exist_ok=True)
n = 0
for shard in ('00', '01'):
p = hf_hub_download('agkphysics/AudioSet', f'data/bal_train/{shard}.parquet', repo_type='dataset',
local_dir=os.path.join(D, 'audioset_src'))
t = pq.read_table(p, columns=['audio'])
for row in t.column('audio').to_pylist():
try:
a, sr = sf.read(io.BytesIO(row['bytes']), dtype='float32')
except Exception:
continue
sf.write(os.path.join(out, f'as_{shard}_{n:05d}.wav'), to16k(a, sr), 16000, subtype='PCM_16'); n += 1
os.remove(p) # keep only the converted wavs
print('background wavs', n, flush=True)
if step in ('all', 'features'):
for f in ('validation_set_features.npy', 'openwakeword_features_ACAV100M_2000_hrs_16bit.npy'):
p = hf_hub_download('davidscripka/openwakeword_features', f, repo_type='dataset', local_dir=D)
print('feature file', p, os.path.getsize(p) / 1e9, 'GB', flush=True)
print('DONE', step)
You do not run it by hand: train.sh runs prepare_data.py all when the data is missing.
Check
The lines it printed intrain.logon 2026-09-28:base models ok,rir wavs 270,background wavs 1000,
feature file /home/burgerbarn/voice/wakeword/data/validation_set_features.npy 0.18483(then the 17.28 GB
file),DONE all.du -sh ~/voice/wakeword/datais 17G.
If it fails
Warning: You are sending unauthenticated requests to the HF Hubis harmless.Temporary failure in name resolution ... Retrying in 1s [Retry 1/5]: the network dropped. On 2026-09-28
that was bigbuddy suspending in the middle of the download (chapter 25); the download retried after the
wake. Keep suspend off.
openWakeWord's train.py makes its "yo buddy" clips with Piper, through a function generate_samples() it
imports from the folder named by piper_sample_generator_path. voice/wakeword/kokoro_gen/ provides a function
with that name and signature that uses Kokoro instead. Because Kokoro cannot run in the numpy<2 training venv, the
function writes the job list to a JSON file and runs kokoro_worker.py in ~/voice/tts-venv.
voice/wakeword/kokoro_gen/generate_samples.py:
On bigbuddy:
"""Drop-in replacement for piper-sample-generator's generate_samples(), used by openwakeword/train.py
(config piper_sample_generator_path -> this directory). The training venv (Py3.11, numpy<2) cannot run the
current kokoro-onnx (numpy>=2), so synthesis runs in the TTS venv (~/voice/tts-venv) via kokoro_worker.py.
Writes 16 kHz mono int16 wavs (what openWakeWord expects)."""
import json, os, random, subprocess, tempfile, uuid
TTS_PY = os.path.expanduser('~/voice/tts-venv/bin/python')
WORKER = os.path.join(os.path.dirname(os.path.abspath(__file__)), 'kokoro_worker.py')
def generate_samples(text, max_samples, output_dir, length_scales=(1.0,), file_names=None,
batch_size=None, noise_scales=None, noise_scale_ws=None, auto_reduce_batch_size=None, **kw):
texts = [text] if isinstance(text, str) else list(text)
os.makedirs(output_dir, exist_ok=True)
names = file_names or [uuid.uuid4().hex + '.wav' for _ in range(max_samples)]
jobs = [[texts[i % len(texts)], length_scales[i % len(length_scales)], os.path.join(output_dir, names[i]),
random.getrandbits(32)] for i in range(max_samples)]
random.shuffle(jobs)
with tempfile.NamedTemporaryFile('w', suffix='.json', delete=False) as f:
json.dump(jobs, f)
try:
subprocess.run([TTS_PY, WORKER, f.name], check=True)
finally:
os.unlink(f.name)
voice/wakeword/kokoro_gen/kokoro_worker.py does the synthesis in a pool of processes. Read it for what each
part protects against:
_init() runs once in every worker. It gives each worker a private temp folder on disk (the espeak copies/tmp: 10,000 clips made 6.3 GB), limits onnxruntime to KOKORO_THREADS threads, turnskokoro_patch._one() makes one clip: two random English voices blended half the time, speed jittered ±10 %, punctuation__main__ runs _init() once in the parent first so a broken setup fails at once, then a Pool ofKOKORO_PROCS workers recycled every KOKORO_RECYCLE tasks. It fails if fewer than 90 % of clips wereOn bigbuddy:
"""Kokoro synthesis worker (runs in ~/voice/tts-venv). Input: JSON list of [text, length_scale, path, seed].
English voices + random 2-voice blends, speed from length_scale with jitter, punctuation variants, 16 kHz int16."""
import json, os, random, shutil, sys, tempfile
from multiprocessing import Pool
import numpy as np, soundfile as sf
from scipy.signal import resample_poly
MODEL = os.path.expanduser('~/voice/models/kokoro/kokoro-v1.0.onnx')
VOICES = os.path.expanduser('~/voice/models/kokoro/voices-v1.0.bin')
_k = None
_names = None
_tmp = None
def _init():
# Each worker holds its own model: keep threads/memory small (14 default workers OOM-killed bigbuddy's
# desktop, 2026-09-28). KOKORO_PROCS workers x KOKORO_THREADS onnxruntime threads.
global _k, _names, _tmp
# phonemizer-fork copies libespeak-ng.so (648 KB) into a NEW tempdir on every phonemize call and never
# removes it: 10k clips = 6.3 GB in RAM-backed /tmp (2026-09-28). Private on-disk TMPDIR, emptied per clip.
_tmp = tempfile.mkdtemp(prefix=f'kokoro-{os.getpid()}-', dir=os.path.expanduser('~/voice/wakeword/tmp'))
os.environ['TMPDIR'] = _tmp
tempfile.tempdir = _tmp
import onnxruntime as ort
from kokoro_onnx import Kokoro
so = ort.SessionOptions()
so.intra_op_num_threads = int(os.environ.get('KOKORO_THREADS', '2'))
so.inter_op_num_threads = 1
so.enable_cpu_mem_arena = False
so.enable_mem_pattern = False # per-shape buffers grew RSS to OOM (16G scope) over ~2k clips
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import kokoro_patch
kokoro_patch.apply() # reuse one espeak backend (per-call leak fix)
_k = Kokoro.from_session(ort.InferenceSession(MODEL, so, providers=['CPUExecutionProvider']), VOICES)
_names = [v for v in _k.get_voices() if v[:2] in ('af', 'am', 'bf', 'bm')]
def _one(job):
text, length_scale, path, seed = job
rnd = random.Random(seed)
a_name, b_name = rnd.sample(_names, 2)
w = rnd.random() if rnd.random() < 0.5 else 1.0 # half pure voices, half blends
style = w * _k.get_voice_style(a_name) + (1 - w) * _k.get_voice_style(b_name)
lang = 'en-gb' if a_name.startswith('b') else 'en-us'
speed = float(np.clip((1.0 / length_scale) * rnd.uniform(0.9, 1.1), 0.6, 1.6))
t = rnd.choice([text, text + '.', text + '!', text + '?', text.capitalize() + '!'])
try:
a, sr = _k.create(t, voice=style, speed=speed, lang=lang)
except Exception as e:
print('skip', repr(t), e, file=sys.stderr)
return 0
for d in os.listdir(_tmp): # drop espeak lib copies (already dlopen'ed)
shutil.rmtree(os.path.join(_tmp, d), ignore_errors=True)
a = resample_poly(a, 2, 3) # 24 kHz -> 16 kHz
a = a / max(1e-6, np.abs(a).max()) * rnd.uniform(0.3, 0.9)
sf.write(path, (a * 32767).astype(np.int16), 16000, subtype='PCM_16')
return 1
if __name__ == '__main__':
jobs = json.load(open(sys.argv[1]))
os.makedirs(os.path.expanduser('~/voice/wakeword/tmp'), exist_ok=True)
_init() # fail fast here, not in every pool worker
n = 0
with Pool(int(os.environ.get('KOKORO_PROCS', '4')), initializer=_init,
maxtasksperchild=int(os.environ.get('KOKORO_RECYCLE', '200'))) as p: # bound any growth
for i, r in enumerate(p.imap_unordered(_one, jobs, chunksize=8)):
n += r
if i % 1000 == 0:
print(f'kokoro: {i}/{len(jobs)}', flush=True)
print(f'kokoro: wrote {n}/{len(jobs)}', flush=True)
if n < 0.9 * len(jobs):
sys.exit(f'kokoro: only {n}/{len(jobs)} clips written')
Test the generator with 40 clips, as on 2026-09-28. It runs in the training venv, exactly as train.py will
call it:
On bigbuddy:
rm -rf /tmp/kgen_test; cd ~/voice/wakeword/kokoro_gen
timeout 300 ~/voice/oww-train/bin/python -c "
import time; from generate_samples import generate_samples
t=time.time(); generate_samples(text=['yo buddy'], max_samples=40, output_dir='/tmp/kgen_test', length_scales=[0.75,1.0,1.25]); print('40 clips in %.1fs'%(time.time()-t))
" 2>&1 | tail -3
~/voice/tts-venv/bin/python -c "
import soundfile as sf, glob, numpy as np
fs=sorted(glob.glob('/tmp/kgen_test/*.wav')); d=[sf.info(f).duration for f in fs]; i=sf.info(fs[0])
print(len(fs),'files', i.samplerate,'Hz', i.subtype, 'dur min/med/max %.2f/%.2f/%.2f s'%(min(d),np.median(d),max(d)))"
Check
Output on 2026-09-28:kokoro: 0/40 kokoro: wrote 40/40 40 clips in 11.7s 40 files 16000 Hz PCM_16 dur min/med/max 0.49/0.95/1.58 sPlay a few with
paplay /tmp/kgen_test/<name>.wav.
If it fails
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 18fromkokoro_onnx/config.py
get_voice_names: Kokoro ran inside the training venv, whose older kokoro-onnx (numpy<2) cannot read
voices-v1.0.bin. That is why the worker runs in~/voice/tts-venv(2026-09-28).ModuleNotFoundError: No module named 'scipy'fromkokoro_worker.py: install scipy in the tts-venv.
voice/wakeword/yo_buddy.yaml (keys as openWakeWord 0.6.0's train.py reads them):
On bigbuddy:
# openWakeWord 0.6.0 train.py config for "yo buddy" (keys per openwakeword/train.py).
# Positive/adversarial clips synthesized by Kokoro (kokoro_gen/generate_samples.py), not Piper.
model_name: "yo_buddy"
target_phrase: ["yo buddy"]
custom_negative_phrases: ["hey buddy", "yo body", "you buddy", "yo bunny", "yeah buddy", "oh buddy", "buddy",
"yo", "yo dude", "go buddy", "no buddy", "my buddy", "yo bro", "hi buddy", "yo baby"]
n_samples: 20000
n_samples_val: 2000
tts_batch_size: 50
augmentation_batch_size: 16
piper_sample_generator_path: "/home/burgerbarn/voice/wakeword/kokoro_gen"
output_dir: "/home/burgerbarn/voice/wakeword/work"
rir_paths: ["/home/burgerbarn/voice/wakeword/data/rir"]
# + real echo-cancelled music/movie residual from bigbuddy (2026-09-28), weighted x10
background_paths: ["/home/burgerbarn/voice/wakeword/data/background", "/home/burgerbarn/voice/wakeword/data/background_aec"]
background_paths_duplication_rate: [1, 10]
false_positive_validation_data_path: "/home/burgerbarn/voice/wakeword/data/validation_set_features.npy"
augmentation_rounds: 1
feature_data_files:
"ACAV100M_sample": "/home/burgerbarn/voice/wakeword/data/openwakeword_features_ACAV100M_2000_hrs_16bit.npy"
batch_n_per_class:
"ACAV100M_sample": 1024
"adversarial_negative": 50
"positive": 50
model_type: "dnn"
layer_size: 32
steps: 50000
max_negative_weight: 1500
target_false_positives_per_hour: 0.2
n_samples: 20000 positive and 20,000 adversarial training clips, 2,000 of each for testing.custom_negative_phrases: sound-alikes the model must learn to reject ("hey buddy", "yo body", "yo baby").background_paths with duplication_rate: noise mixed under the clips; the v3 media residual counts ten times.batch_n_per_class: each training batch has 1024 general-audio negatives, 50 sound-alikes and 50 positives.model_type: dnn, layer_size: 32, steps: 50000: the small classifier and how long to train it.max_negative_weight and target_false_positives_per_hour: how hard training pushes false wakes down.The first run needs the v1 background lines
The repo file is the v3 config: it listsdata/background_aec, which only exists after the v3 media step.
For a first, synthetic-only run, use the two lines as they were for v1 and v2 (from the 2026-09-28 log):background_paths: ["/home/burgerbarn/voice/wakeword/data/background"] background_paths_duplication_rate: [1]What openWakeWord does with a missing background folder was not tested.
The first full run reached the augmentation step and stopped with
ImportError: TorchCodec is required for load_with_torchcodec. Please install torchcodec to use this function.
torchaudio 2.9 and newer send load() through torchcodec (which needs a system FFmpeg) and dropped info().
Training only reads local WAV files, so the wrapper replaces those two functions with soundfile, then runs
openWakeWord's own train.py as if it were the main program.
Replace torchaudio.load with a soundfile read that honours the arguments train.py uses:
On bigbuddy:
def _load(path, frame_offset=0, num_frames=-1, normalize=True, channels_first=True, format=None, **kw):
stop = None if num_frames is None or num_frames < 0 else frame_offset + num_frames
dtype = 'float32' if normalize else 'int16'
data, sr = sf.read(path, start=frame_offset, stop=stop, dtype=dtype, always_2d=True)
t = torch.from_numpy(data.T.copy() if channels_first else data.copy())
return t, sr
Replace torchaudio.info with an object that has the fields train.py reads:
On bigbuddy:
def _info(path, format=None, **kw):
i = sf.info(path)
return SimpleNamespace(sample_rate=i.samplerate, num_frames=i.frames, num_channels=i.channels,
bits_per_sample=16 if 'PCM_16' in i.subtype else 0, encoding=i.subtype)
Patch them in before openWakeWord is imported, then run train.py with the wrapper's own arguments:
On bigbuddy:
torchaudio.load = _load
torchaudio.info = _info
import openwakeword
train_py = os.path.join(os.path.dirname(openwakeword.__file__), 'train.py')
sys.argv = [train_py] + sys.argv[1:]
runpy.run_path(train_py, run_name='__main__')
The complete voice/wakeword/train_wrapper.py:
On bigbuddy:
"""Run openwakeword/train.py with torchaudio.load/info backed by soundfile.
torchaudio >= 2.9 routes load() through torchcodec (+ system FFmpeg) and dropped info(); training only
reads local WAV files, which soundfile handles. Usage: train_wrapper.py <train.py args...>"""
import os, runpy, sys
from types import SimpleNamespace
import soundfile as sf
import torch
import torchaudio
def _load(path, frame_offset=0, num_frames=-1, normalize=True, channels_first=True, format=None, **kw):
stop = None if num_frames is None or num_frames < 0 else frame_offset + num_frames
dtype = 'float32' if normalize else 'int16'
data, sr = sf.read(path, start=frame_offset, stop=stop, dtype=dtype, always_2d=True)
t = torch.from_numpy(data.T.copy() if channels_first else data.copy())
return t, sr
def _info(path, format=None, **kw):
i = sf.info(path)
return SimpleNamespace(sample_rate=i.samplerate, num_frames=i.frames, num_channels=i.channels,
bits_per_sample=16 if 'PCM_16' in i.subtype else 0, encoding=i.subtype)
torchaudio.load = _load
torchaudio.info = _info
import openwakeword
train_py = os.path.join(os.path.dirname(openwakeword.__file__), 'train.py')
sys.argv = [train_py] + sys.argv[1:]
runpy.run_path(train_py, run_name='__main__')
voice/wakeword/train.sh runs the three stages of train.py: make clips, augment them (add room echo and
background, compute features), train the classifier.
On bigbuddy:
#!/bin/bash
# Train the "yo buddy" wake word on bigbuddy. Unattended; log: ~/voice/wakeword/train.log
# Order: data -> generate clips (Kokoro) -> augment -> train. Output: work/yo_buddy.onnx
set -euo pipefail
V=~/voice/oww-train/bin/python; W=~/voice/wakeword; cd $W
[ -f data/validation_set_features.npy ] && [ -d data/background ] && [ -d data/rir ] || $V prepare_data.py all
TP=$W/train_wrapper.py # torchaudio.load/info via soundfile (torchaudio>=2.9 needs torchcodec+FFmpeg)
$V $TP --training_config yo_buddy.yaml --generate_clips
$V $TP --training_config yo_buddy.yaml --augment_clips
$V $TP --training_config yo_buddy.yaml --train_model
ls -la work/*.onnx
echo TRAIN_DONE
The launch line from voice/README.md, used for every run since the OOM:
On bigbuddy:
cd ~/voice/wakeword && KOKORO_PROCS=6 KOKORO_THREADS=2 setsid nohup systemd-run --user --scope -q \
--unit=yo-buddy-train -p MemoryMax=16G -p MemorySwapMax=0 nice -n 19 bash train.sh > train.log 2>&1 < /dev/null &
Read it from the inside out:
bash train.sh is the job.nice -n 19 gives it the lowest CPU priority, so the desktop and Buddy stay responsive.systemd-run --user --scope --unit=yo-buddy-train -p MemoryMax=16G -p MemorySwapMax=0 puts the job and everyKOKORO_PROCS=6 KOKORO_THREADS=2: six synthesis workers with two threads each, about 0.6 GB per worker.setsid nohup ... & with all output in train.log detaches it from your ssh session.Check
systemctl --user status yo-buddy-train.scope --no-pager | grep -E "Active|Memory|Tasks"30 s after a start
on 2026-09-28:Active: active (running) since Mon 2026-09-28 10:57:07 PDT; 30s ago Tasks: 90 (limit: 37824) Memory: 7G (max: 16G, swap max: 0B, available: 8.9G, peak: 7G)With the patched workers memory stayed flat:
anon 4.2 GB, total 4.6-4.7 GB, at 1, 2 and 3 minutes while about
900 clips a minute were written. Follow progress with
tr '\r' '\n' < ~/voice/wakeword/train.log | grep -E "kokoro:|Computing features|Training" | tail -3:
kokoro: 7000/8938during synthesis,Computing features: 59%|...| 740/1250during augmentation, and for
the final stageTraining: 28%|...| 13854/50000 [00:42<01:39, 361.50it/s](50,000 steps take about 2-3
minutes on the GPU). The run ends withTRAIN_DONE.
If it fails: the 2026-09-28 out-of-memory
The first run usedKOKORO_PROCS=14with onnxruntime's defaults (16 threads per worker and a memory arena).
Fourteen workers exhausted 30 GB of RAM in minutes; the kernel's out-of-memory killer took down the owner's
KDE session at 10:50 and bigbuddy needed a hard restart. Then, with the cap in place, workers still grew to
3.6 GB each by 430 clips: the espeak leak (step 2). Fixed by six workers, two threads, arena and memory
pattern off, the private temp folder,kokoro_patch, recycling workers, and the 16 GB scope
(docs/lessons.md, "bigbuddy OOM from wake-word clip generation"). Two more facts from that day: Python
3.14'smultiprocessingstarts workers with "forkserver", so workers do not share a model loaded in the
parent; andPool(maxtasksperchild=...)counts chunks, not clips.
If it fails: other stops on 2026-09-28
ImportError: TorchCodec is required for load_with_torchcodec: run throughtrain_wrapper.py, as
train.shdoes.ModuleNotFoundError: No module named 'onnxscript'at the end: install it in the training venv.- A traceback ending in
RuntimeError: ... No Previous Version of LayerNormalization exis...from
onnx/version_converter.py, after the ONNX file was written: openWakeWord then tries to convert to
TFLite, which needsonnx_tf. TFLite is not used; the ONNX model is complete.retrain.shtolerates
exactly this failure and no other.- Re-running skips existing clips (
WARNING:root:Skipping generation of positive clips for training, as ~20000 already exist). To redo features, deletework/yo_buddy/*_features_train.npyfirst.
work/yo_buddy.onnx (14,532 bytes) plus work/yo_buddy.onnx.data (200,704 bytes): torch 2.11 writes the
weights to a separate file. Input x of shape [1, 16, 96], 50,403 parameters.
v1 was lost
On 2026-09-28 the synthetic model was saved withcp work/yo_buddy.onnx work/yo_buddy_v1_synthetic.onnx.
The copy still pointed atyo_buddy.onnx.data, which the next training run overwrote, so v1 now loads v2's
weights. Always pack the model into one file before keeping it.
The packing step, as run for v2 on 2026-09-28 (it also proves the packed file gives identical output):
On bigbuddy:
cd ~/voice/wakeword; mkdir -p ~/voice/models/wakeword
~/voice/oww-train/bin/python - <<'PY' 2>&1 | grep -v -i warning
import onnx, onnxruntime as ort, numpy as np, hashlib
m=onnx.load("work/yo_buddy.onnx") # pulls in yo_buddy.onnx.data
out="/home/burgerbarn/voice/models/wakeword/yo_buddy_v2.onnx"
onnx.save_model(m, out, save_as_external_data=False)
x=np.random.rand(1,16,96).astype(np.float32)
a=ort.InferenceSession("work/yo_buddy.onnx").run(None,{"x":x})[0]; b=ort.InferenceSession(out).run(None,{"x":x})[0]
print("single-file identical:", np.allclose(a,b), "size", len(open(out,"rb").read()), "sha256", hashlib.sha256(open(out,"rb").read()).hexdigest()[:16])
PY
Check
For v2:single-file identical: True size 215120 sha256 5cccf306f205ecc0. Your hash will differ: every
training run gives different weights. For a synthetic-only first model, changeyo_buddy_v2inout=to a
name of your own.
v1 recognised only 6 of the owner's 23 real "yo buddy"s at 0.5. Real recordings fixed it. On 2026-09-28 the
owner said "yo buddy" 23 times over 97 s, with pauses and varied loudness, then talked normally for 90 s without
saying it. The recordings went to ~/voice/wakeword/real/owner_pos_1.wav and owner_neg_1.wav. The record
command then (48 kHz mono, from the Komplete Audio 6):
On bigbuddy:
mkdir -p ~/voice/wakeword/real
SRC=alsa_input.usb-Native_Instruments_Komplete_Audio_6_00A241CB-00.analog-mono-in-a
setsid nohup timeout 120 pw-record --target $SRC --rate 48000 --channels 1 --format s16 ~/voice/wakeword/real/owner_pos_1.wav > ~/voice/wakeword/real/rec1.log 2>&1 < /dev/null &
and the same with timeout 90 and owner_neg_1.wav for the talking. Stop early with pkill -INT pw-record.
Not test-built: recording through today's microphone
The Komplete Audio 6 is gone; bigbuddy's only microphone is the robot's, as the PipeWire sourcerobot_mic
(chapter 27). Recording with--target robot_micinstead of the old source has not been done. Stand where
you will talk to the robot, and vary your distance (voice/README.md: "more owner clips (omni, varied
distance)").
add_real.py expects the talking file at 16 kHz with a 100 Hz high-pass, named owner_neg_1_16k.wav. On
2026-09-28 it was made from the 48 kHz recording with this snippet (the 1,3 resample assumes 48 kHz input):
On bigbuddy:
cd ~/voice/wakeword/real
~/voice/oww-train/bin/python - <<'PY'
import numpy as np, soundfile as sf, scipy.signal as ss
a,sr=sf.read("owner_neg_1.wav", dtype="float32")
hp=ss.butter(4,100,btype="high",fs=sr,output="sos"); y=ss.resample_poly(ss.sosfilt(hp,a),1,3).astype(np.float32)
sf.write("owner_neg_1_16k.wav",(y*32767).clip(-32768,32767).astype(np.int16),16000,subtype="PCM_16")
PY
voice/wakeword/add_real.py turns the recordings into training data:
real/owner_pos_*.wav into single utterances by energy: 20 ms frames, "on" when 12 dB abovereal/pos_test for testing; the rest are copied 120 times into the positiveowner_neg_1_16k.wav and any neg_media_*.wav) is split: the first 75 % becomesreal/neg_test_<name>.wav for testing.On bigbuddy:
"""Add the owner's real recordings to the openWakeWord training set (run on bigbuddy, oww-train venv).
Positives: every real/owner_pos_*.wav (one "yo buddy" per utterance, cut by energy) -> clips in real/pos_clips;
every 4th clip held out -> real/pos_test; the rest copied POS_REPS times into work/yo_buddy/positive_train.
Negatives: real/owner_neg_1_16k.wav (talking) and real/neg_media_*.wav (echo-cancelled music/movie, 2026-09-28):
first 75 % of each cut into 2 s chunks (1 s hop) copied NEG_REPS times into negative_train; last 25 % of each
-> real/neg_test_<name>.wav. Re-running replaces earlier owner_* / real_* files."""
import glob, os, shutil
import numpy as np, soundfile as sf
import scipy.signal as ss
W = os.path.expanduser('~/voice/wakeword')
POS_REPS, NEG_REPS = 120, 15
pt, nt = f'{W}/work/yo_buddy/positive_train', f'{W}/work/yo_buddy/negative_train'
for d in (pt, nt):
for f in glob.glob(f'{d}/owner_*.wav') + glob.glob(f'{d}/real_*.wav'):
os.remove(f)
for d in ('pos_clips', 'pos_test'):
shutil.rmtree(f'{W}/real/{d}', ignore_errors=True); os.makedirs(f'{W}/real/{d}')
def to16k(path):
a, sr = sf.read(path, dtype='float32')
if a.ndim > 1:
a = a.mean(axis=1)
if sr != 16000:
a = ss.resample_poly(ss.sosfilt(ss.butter(4, 100, btype='high', fs=sr, output='sos'), a), 16000, sr)
return a.astype(np.float32)
def utterances(y, sr=16000):
f = int(0.02 * sr)
db = 20 * np.log10(np.array([np.sqrt(np.mean(y[i:i + f] ** 2)) for i in range(0, len(y) - f, f)]) + 1e-9)
on = db > np.percentile(db, 20) + 12
segs, i = [], 0
while i < len(on):
if on[i]:
j = i
while j < len(on) and (on[j] or (j + 15 < len(on) and on[j:j + 15].any())):
j += 1
segs.append((i, j)); i = j
else:
i += 1
return [(s * f, e * f) for s, e in segs if 0.3 <= (e - s) * 0.02 <= 2.2]
k = 0
for rec in sorted(glob.glob(f'{W}/real/owner_pos_*.wav')):
if rec.endswith('_16k.wav'):
continue
y = to16k(rec)
for s, e in utterances(y):
c = y[max(0, s - 2400):min(len(y), e + 2400)]
c = c / max(1e-6, np.abs(c).max()) * 0.7
sf.write(f'{W}/real/pos_clips/owner_{k:03d}.wav', (c * 32767).astype(np.int16), 16000, subtype='PCM_16')
k += 1
clips = sorted(glob.glob(f'{W}/real/pos_clips/owner_*.wav'))
n_tr = 0
for i, f in enumerate(clips):
if i % 4 == 3:
shutil.copy(f, f'{W}/real/pos_test/'); continue
for r in range(POS_REPS):
shutil.copy(f, f'{pt}/owner_{i:03d}_{r:03d}.wav')
n_tr += 1
neg_sources = [f'{W}/real/owner_neg_1_16k.wav'] + sorted(glob.glob(f'{W}/real/neg_media_*.wav'))
n_neg = 0
for src in neg_sources:
if not os.path.exists(src):
continue
name = os.path.basename(src).replace('.wav', '')
a, sr = sf.read(src, dtype='int16')
cut = int(len(a) * 0.75)
sf.write(f'{W}/real/neg_test_{name}.wav', a[cut:], sr, subtype='PCM_16')
for s in range(0, cut - 2 * sr, sr):
for r in range(NEG_REPS):
sf.write(f'{nt}/real_{name}_{s // sr:04d}_{r:02d}.wav', a[s:s + 2 * sr], sr, subtype='PCM_16')
n_neg += 1
print(f'positives: {len(clips)} clips from {len(glob.glob(f"{W}/real/owner_pos_*.wav"))} recordings, '
f'{n_tr} train x{POS_REPS}, {len(clips) - n_tr} held out; negatives: {n_neg} chunks x{NEG_REPS} from '
f'{len(neg_sources)} sources')
To measure a model on single clips, voice/wakeword/eval_real.py pads each clip with a second of silence, streams
it through openWakeWord in 80 ms frames exactly as Buddy does, and takes the highest score:
On bigbuddy:
"""Score a wake-word model on real recordings of individual utterances (one wav per utterance, 16 kHz).
Each clip is padded with 1 s silence before/after and streamed through openWakeWord in 80 ms frames;
the clip's score = max model output. Usage: eval_real.py model.onnx clip_dir [clip_dir...]"""
import glob, sys
import numpy as np, soundfile as sf
from openwakeword.model import Model
m = Model(wakeword_models=[sys.argv[1]], inference_framework='onnx')
key = list(m.models.keys())[0]
for d in sys.argv[2:]:
scores = []
for f in sorted(glob.glob(d + '/*.wav')):
a, sr = sf.read(f, dtype='int16')
a = np.concatenate([np.zeros(16000, np.int16), a, np.zeros(16000, np.int16)])
m.reset()
scores.append(max(m.predict(a[i:i + 1280])[key] for i in range(0, len(a) - 1280, 1280)))
s = np.array(scores)
print(f'{d}: n={len(s)} detected@0.5 {int((s >= 0.5).sum())}/{len(s)} @0.3 {int((s >= 0.3).sum())}/{len(s)} '
f'scores {np.round(np.sort(s), 2).tolist()}')
Retrain with your recordings using retrain.sh (next section): bash retrain.sh v2. Then score it:
On bigbuddy:
cd ~/voice/wakeword && ~/voice/oww-train/bin/python eval_real.py ~/voice/models/wakeword/yo_buddy_v2.onnx real/pos_test
Check
For the owner's v2 on 2026-09-28:real/pos_test: n=5 detected@0.5 4/5 @0.3 4/5 scores [0.12999999523162842, 0.7599999904632568, 0.8899999856948853, 0.8999999761581421, 0.9200000166893005]and 0 triggers on the 22.5 s of held-out talking (max score 0.00). On 2026-09-28 v2 was built with the same
steps run by hand (add_real.py, delete the two feature files,--augment_clips,--train_model, pack);
retrain.shwas written later that day and does all of them.
Buddy shares the room with music and films. v3 adds what Buddy's microphone hears while media plays, after echo
cancellation, as negatives and as extra background. It needs chapters 27 and 28.
voice/wakeword/record_media_negatives.sh stops Buddy, starts music, records buddy_aec_source for 300 s, starts
a film, records 120 s, stops the media and starts Buddy again:
On bigbuddy:
#!/bin/bash
# Record what Buddy hears (echo-cancelled mic) while media plays, as wake-word NEGATIVES.
# Pauses buddy-voice during recording so it can't react. Output: ~/voice/wakeword/real/neg_media_<tag>.wav (16 kHz)
set -uo pipefail
MUSIC_S=${1:-300}; MOVIE_S=${2:-120}; D=~/voice/wakeword/real; mkdir -p $D; cd ~/voice
V=~/voice/venv/bin/python; T=$(date +%H%M)
systemctl --user stop buddy-voice.service
rec() { timeout $2 parec --device=buddy_aec_source --rate=16000 --channels=1 --format=s16le --raw > /tmp/rec_$1.raw
$V -c "import numpy as np, soundfile as sf; a=np.fromfile('/tmp/rec_$1.raw', np.int16); sf.write('$D/neg_media_$1.wav', a, 16000, subtype='PCM_16'); print('$1', len(a)/16000, 's')"; }
$V media_screen.py music >/dev/null 2>&1; sleep 5; rec music_$T $MUSIC_S
$V media_screen.py movie >/dev/null 2>&1; sleep 10; rec movie_$T $MOVIE_S
$V media_screen.py control stop >/dev/null 2>&1
systemctl --user start buddy-voice.service
echo RECORD_DONE
On 2026-09-28 a second, louder set was recorded with the amplifier at 40 % (music_loud 118.0 s, rms -43.8 dBFS,
movie_loud 118.0 s, rms -50.3 dBFS), giving four files: neg_media_music_1414.wav, neg_media_music_loud.wav,
neg_media_movie_raiders.wav, neg_media_movie_loud.wav.
Docs differ
Since60-mic.confBuddy listens torobot_mic, but this script still recordsbuddy_aec_source. To record
what Buddy hears today, the device would berobot_mic; that was not done.
The same recordings, cut into 10-second pieces with a 5-second hop from the first 75 % of each file, became the
data/background_aec set that yo_buddy.yaml mixes in ten times over. This step is not a repo script; it was run
from the command line on 2026-09-28:
On bigbuddy:
cd ~/voice/wakeword; V=~/voice/oww-train/bin/python
mkdir -p data/background_aec; rm -f data/background_aec/*
$V - <<'PY'
import glob, numpy as np, soundfile as sf
n = 0
for f in sorted(glob.glob('real/neg_media_*.wav')):
a, sr = sf.read(f, dtype='int16'); cut = int(len(a) * 0.75) # same 75 % train split as add_real.py
for s in range(0, cut - 10 * sr, 5 * sr):
sf.write(f'data/background_aec/{f.split("/")[-1][:-4]}_{s // sr:04d}.wav', a[s:s + 10 * sr], sr, subtype='PCM_16'); n += 1
print('background_aec clips (10 s):', n)
PY
Check
On 2026-09-28:background_aec clips (10 s): 91.
voice/wakeword/retrain.sh reruns augmentation and training with whatever is in real/, and exports a packed
model. Its first version had a bug that this version exists to prevent (below). Build it up.
Arguments and the optional "skip to training" mode:
On bigbuddy:
set -euo pipefail
TAG=${1:?tag}; V=~/voice/oww-train/bin/python; W=~/voice/wakeword; cd $W
if [ "${2:-}" != "--train-only" ]; then
$V add_real.py
rm -f work/yo_buddy/positive_features_train.npy work/yo_buddy/negative_features_train.npy
$V train_wrapper.py --training_config yo_buddy.yaml --augment_clips
fi
Before the GPU stage: note the time, stop Buddy (Whisper and the language model already hold 12 of 16 GB), and
make sure Buddy comes back however the script ends:
On bigbuddy:
MARK=$(date +%s)
systemctl --user stop buddy-voice.service
trap 'systemctl --user start buddy-voice.service' EXIT
Train. Accept a failure only if it is the known TFLite conversion error; then refuse to export unless the model
file is newer than MARK:
On bigbuddy:
$V train_wrapper.py --training_config yo_buddy.yaml --train_model > train_step.log 2>&1 || \
grep -q -E "onnx_tf|convert_onnx_to_tflite" train_step.log || { tail -5 train_step.log; exit 1; }
[ "$(stat -c %Y work/yo_buddy.onnx)" -ge "$MARK" ] || { echo "no new model written"; tail -5 train_step.log; exit 1; }
Pack and export:
On bigbuddy:
$V -c "
import onnx; m = onnx.load('work/yo_buddy.onnx'); onnx.save_model(m, '/home/burgerbarn/voice/models/wakeword/yo_buddy_$TAG.onnx', save_as_external_data=False); print('exported yo_buddy_$TAG.onnx')"
echo RETRAIN_DONE
The complete voice/wakeword/retrain.sh:
On bigbuddy:
#!/bin/bash
# Retrain "yo buddy" with the current real recordings (add_real.py) and export ~/voice/models/wakeword/yo_buddy_<tag>.onnx
# Usage: retrain.sh <tag> [--train-only]. Pauses buddy-voice during GPU training (Whisper + LM Studio + training
# exceeded 16 GB -> CUDA OOM, 2026-09-28). Exports ONLY if training wrote a new model (the known TFLite-export
# failure after the ONNX export is tolerated; nothing else is).
set -euo pipefail
TAG=${1:?tag}; V=~/voice/oww-train/bin/python; W=~/voice/wakeword; cd $W
if [ "${2:-}" != "--train-only" ]; then
$V add_real.py
rm -f work/yo_buddy/positive_features_train.npy work/yo_buddy/negative_features_train.npy
$V train_wrapper.py --training_config yo_buddy.yaml --augment_clips
fi
MARK=$(date +%s)
systemctl --user stop buddy-voice.service
trap 'systemctl --user start buddy-voice.service' EXIT
$V train_wrapper.py --training_config yo_buddy.yaml --train_model > train_step.log 2>&1 || \
grep -q -E "onnx_tf|convert_onnx_to_tflite" train_step.log || { tail -5 train_step.log; exit 1; }
[ "$(stat -c %Y work/yo_buddy.onnx)" -ge "$MARK" ] || { echo "no new model written"; tail -5 train_step.log; exit 1; }
$V -c "
import onnx; m = onnx.load('work/yo_buddy.onnx'); onnx.save_model(m, '/home/burgerbarn/voice/models/wakeword/yo_buddy_$TAG.onnx', save_as_external_data=False); print('exported yo_buddy_$TAG.onnx')"
echo RETRAIN_DONE
Run it under the same memory cap as train.sh:
On bigbuddy:
cd ~/voice/wakeword; systemctl --user reset-failed yo-buddy-train.scope 2>/dev/null
setsid nohup systemd-run --user --scope -q --unit=yo-buddy-train -p MemoryMax=16G -p MemorySwapMax=0 nice -n 19 bash retrain.sh v3 > retrain.log 2>&1 < /dev/null &
Check
The first line ofretrain.logfor v3 on 2026-09-28:positives: 23 clips from 2 recordings, 18 train x120, 5 held out; negatives: 549 chunks x15 from 5 sourcesand the last two lines
exported yo_buddy_v3.onnxandRETRAIN_DONE. Then
sha256sum ~/voice/models/wakeword/*.onnx | cut -c1-16,65-must show a new hash for the new model; on
2026-09-28:5cccf306f205ecc0for v2 anda197fb540a5932a0for v3.
The retrain that exported the old model (2026-09-28)
The firstretrain.shended its train step with|| trueto get past the TFLite error, and Buddy was
restarted while it trained. The train step failed (the record names CUDA out of memory, with Whisper, the
language model and training sharing 16 GB),|| truehid it, and the script packed the previous
work/yo_buddy.onnxas v3:sha256sumshowed5cccf306f205ecc0for both v2 and v3, and the model file was
still dated 12:53. The version above stops Buddy during training, starts him again on exit, tolerates only
the TFLite failure, and refuses to export a model older than the run. It was rerun with--train-onlyand
produced the real v3.
retrain.sh needs the service unit
retrain.shrunssystemctl --user stop buddy-voice.serviceunderset -e. If the unit file from step 6 is
not installed, that command fails and the script ends before training (from the script; not seen in the
record). Its exit trap starts Buddy again; before chapter 27 is done, stop him afterwards with
systemctl --user stop buddy-voice.service.
Compare the models on everything held out. The full comparison of 2026-09-28 printed:
On bigbuddy:
v2: synthetic recall 45.2% sound-alike FA 0.25% general audio 0.09/h | Matt held-out 4/5 [0.89, 0.92, 0.13, 0.9, 0.76] | media/talk FT 0 (max 0.00)
v3: synthetic recall 45.6% sound-alike FA 0.25% general audio 0.00/h | Matt held-out 4/5 [0.91, 0.95, 0.3, 0.9, 0.83] | media/talk FT 0 (max 0.04)
Buddy switched to v3 the same day (commit 3b5ee3c) and logged ready: wake yo_buddy_v3 @ 0.5.
What was tried and dropped
The owner also recorded "yo buddy" over music. It was discarded: Whisper's word timings over music were too
unreliable to cut the clips, and mixing the media residual into the background covers "yo buddy over music"
(voice/README.md, yo_buddy_v3).
When training is finished, the 17 GB in ~/voice/wakeword/data can be deleted; the owner kept it on bigbuddy
for later fine-tuning.
~/voice/venv/bin/pip list shows faster-whisper, openwakeword 0.6.0, kokoro-onnx and the NVIDIA wheels, andsha256sum ~/voice/models/kokoro/* matches 7d5df8ec... and bca610b8....sha256sum ~/voice/models/wakeword/yo_buddy_v3.onnx prints a197fb54... (or your own model's hash)../run_buddy_voice.sh --ask "yo buddy, introduce yourself in one sentence" logs a reply: line.~/.config/systemd/user/buddy-voice.service and the three drop-ins exist, andsystemctl --user show buddy-voice -p Environment shows BUDDY_MIC=robot_mic. The service is enabled at theWhere this comes from
sources/survey_bigbuddy.md(sections 1, 2 Phase 1 and 5, 3 hops 7-11),voice/README.md,
voice/buddy_voice.py,voice/run_buddy_voice.sh,voice/buddy-voice.service,voice/kokoro_patch.py,
voice/wakeword/(all scripts andyo_buddy.yaml),voice/models/;sources/cmdlog_bigbuddy.md
2026-09-28 16:39-21:57 UTC (venv and pip commands, Kokoro download and speed test, oww-train installs,
generator test, OOM and relaunches, torchcodec and onnxscript failures, v1/v2/v3 evaluation, the retrain
bug) and 2026-10-07 00:58-01:29 UTC (= 2026-10-06 evening local: drop-ins, revert of20-wake.conf);
docs/lessons.md("bigbuddy OOM", "Voice loop answered itself", 2026-10-01 "Buddy kept saying you're
welcome", 2026-10-06 mic boosts);docs/decisions.md2026-09-28 (compute split, sustainability rule) and
2026-10-06 (Buddy's presence moves to the robot);docs/hardware.md"bigbuddy voice host";docs/status.md
milestone 2026-10-06; commits a806139, 2305f23, ce9fd03, 33f2155, 3b5ee3c, f4904b6, 24d900d, bd434d1,
448b466; read-only checks on bigbuddy 2026-10-07 (md5 of all voice files against the repo, drop-ins, pip
lists,journalctl --user -u buddy-voice).