# Prototype-phase findings Outcome of the two probes run on 2026-06-11 against the Pi at `192.168.50.115`. Both passed their spec criteria. This document captures what we learned, what to carry forward, and what to revisit when designing the main assistant. ## What we proved - **Recording + playback works.** `python-sounddevice` + `soundfile` at 16 kHz mono PCM16 round-trips cleanly through the USB Speaker Phone on the Pi. Level meter responds to speech; playback is recognisable. - **Wakeword detection works.** openwakeword's stock `alexa` model loads in ~0.6 s, detects reliably (user reports "runs perfectly"), false-positive rate is acceptable, CPU is well below saturation on a single core. ## Environment facts to remember - Pi 4B, 4 GB, Debian 13 (trixie), Python 3.13.5, kernel `6.12.75+rpt-rpi-v8`. - PipeWire is the audio server, but PortAudio sees only the raw ALSA hostapi — no PipeWire/Pulse hostapi. - The USB Speaker Phone (Anhui LISTENAI, USB `f201:3220`) is exposed as ALSA card 3. **Native sample rate is 48 kHz only.** Raw `hw:3,0` rejects 16 kHz with `PaErrorCode -9997`. The ALSA `plughw:CARD=Phone,DEV=0` device exists and does the resampling. - The same USB device is *both* mic and speaker. There is no separate output device. - Pi `sudo` for user `pi` requires a password (not passwordless). ## Code patterns to carry forward ### 1. Force PortAudio through ALSA's `plug` plugin Set this **before importing `sounddevice`** in every test: ```python import os os.environ.setdefault("PA_ALSA_PLUGHW", "1") ``` Best done in one place in the main assistant — e.g. an `assistant/audio.py` that sets the env var and re-exports `sounddevice`. The ordering is fragile if every entry point has to remember it. ### 2. USB device discovery `find_usb_device()` is duplicated byte-identically in both tests: ```python def find_usb_device() -> int: for idx, dev in enumerate(sd.query_devices()): name = dev["name"].lower() if "usb" in name and dev["max_input_channels"] >= 1: return idx print("USB audio device not found. Devices:", file=sys.stderr) print(sd.query_devices(), file=sys.stderr) sys.exit(1) ``` Lift to a shared module in the main assistant. ### 3. openwakeword needs `--no-deps` on Python 3.13 `openwakeword==0.6.0` declares a hard dep on `tflite-runtime>=2.8.0,<3 ; platform_system == "Linux"`. There is no `tflite-runtime` wheel for Python 3.13 and pip refuses to install. Workaround: `bin/deploy` does two pip-install passes — a normal `requirements.txt` (sounddevice, numpy, onnxruntime, scipy, scikit-learn, tqdm, requests) and a `requirements-nodeps.txt` (openwakeword) installed with `--no-deps`. Keep this split for the main assistant until either the Pi drops to 3.12 or openwakeword stops hard-requiring tflite-runtime. ### 4. `openwakeword.utils.download_models(target_directory=None)` crashes The function does `os.path.exists(target_directory)` with no None-handling. Call `download_models()` with no kwargs — the default is the package's `resources/models` dir, which is what you want. ## Things to fix or redesign in the main assistant ### A. Beep / TTS blocks the input stream In Test 2, `sd.play(beep) + sd.wait()` runs in the same thread that pulls from the InputStream. During the ~200 ms beep, the InputStream's ring buffer fills up; the next `stream.read()` returns stale audio and may flag input overflow. In the main assistant the same pattern would **drop the first ~200 ms of the user's command** right after the wakeword fires — the worst possible moment. Pattern for the main assistant: - Callback-driven InputStream that pushes frames onto a `queue.Queue`. - A consumer thread (or asyncio task) reads from the queue and feeds the wakeword model / Realtime stream. - A separate, persistent OutputStream opened once at startup that the TTS / beep path pushes to. ### B. The speaker is the mic The USB Speaker Phone hears its own output. TTS playback will re-trigger the wakeword on its own speech. Options: - **Gate the detector during playback.** Set a flag while TTS is playing; skip `model.predict()` while it's set. Simple and reliable. - **Use hardware AEC if the device has it.** The Anhui LISTENAI USB speakerphones often have on-board AEC — worth probing (try `arecord` while `aplay` is running, see if the input is clean). - **Software AEC** (`webrtc-audio-processing`). Adds complexity; only if hardware AEC isn't enough. Recommendation: gate the detector during playback as the v1 behaviour. Revisit AEC only if barge-in becomes a requirement. ### C. Custom wakeword (target: "hey assistant" or similar) Test 2 used the stock `alexa` model as a probe. The final wakeword needs to be trained. openwakeword has a Colab notebook that generates synthetic training data and trains a model in ~1 hour. Plan this as a separate task before the main assistant ships. ### D. Single-core CPU budget The Pi 4 has 4 cores and the wakeword model uses well under 50% of one. The main assistant will add the Realtime WebSocket (small) plus opus encode/decode for the audio it sends and receives (non-trivial). Budget for: 1 core for audio I/O + wakeword, 1 core for Realtime / codec, headroom on the other two. ### E. `bin/bootstrap` and the sudo password `bin/bootstrap` currently does `echo "$PI_PASS" | sudo -S` to authenticate sudo on the Pi. That puts the password into the remote process's `/proc//cmdline` briefly. Alternatives, in order of safety: 1. **Enable passwordless sudo for `pi` on the Pi.** Cleanest. Requires explicit security decision. Add a single file `/etc/sudoers.d/010-pi-nopasswd` containing `pi ALL=(ALL) NOPASSWD: ALL`. 2. **Run bootstrap once manually, never script it.** Avoids the issue entirely. 3. **Switch to `expect` instead of sshpass + sudo -S.** More code, similar trust model. Pre-decide for the main assistant; the deploy/install scripts will hit this same wall. ### F. No requirements lockfile For prototypes, pinning major versions in `requirements.txt` is fine. For the long-lived main assistant, switch to `pip-compile` (or `uv`) to produce a lockfile so the Pi doesn't drift over time. ### G. `~/assistant/` layout vs. rsync `--delete` `bin/deploy` uses `rsync -az --delete` against `~/assistant/`. That's correct for a code dir but will nuke anything else placed there. For the main assistant, split the layout: - `~/assistant/code/` — rsync target with `--delete`. - `~/assistant/state/` — recordings, logs, custom wakeword model, runtime data. Never rsynced. ### H. systemd autostart Not in scope for prototypes. The main assistant needs a user systemd service so it survives reboots and restarts on crash. ## Cosmetic / known-warnings - onnxruntime prints two warnings on startup: ``` Failed to detect devices under "/sys/class/drm/card0" Failed to detect devices under "/sys/class/drm/card1" ``` It's probing for GPU vendor info; the Pi's vc4 framebuffer doesn't expose a `device/vendor` sysfs file. Harmless. Suppressible via `os.environ["ORT_DISABLE_GPU_DETECTION"] = "1"` (or similar onnxruntime log-level setting) if it bothers you. ## Open questions for the next design session 1. Wakeword phrase for the custom model — "hey assistant", "assistant", "computer", something else? 2. Conversation lifecycle — wakeword → open Realtime session → close when? On silence (VAD)? On user-spoken endword? On timeout? 3. Barge-in (user interrupting the assistant mid-reply) — in scope or not? 4. Where the OpenAI API key lives. (Env var? Pi-only file? Pulled from a secrets store at boot?) 5. Local-only fallback when the network is down — required behaviour or out of scope? 6. Multi-language? Or English only for v1? ## C# probe outcome (2026-06-12 — Test 1 ported to C#) The Python Test 1 (record 5 s → write WAV → play back) was re-implemented in C# / .NET 9 to validate that the Pi client can be written in the same language as the backend. The full probe lives at `tests/01-record-play-cs/` with the deploy script at `bin/probe-cs`. All three pass criteria from the original spec held on hardware. ### What we proved - **PortAudioSharp2 1.0.6** is a usable .NET binding for PortAudio. Conventional API surface (`PortAudio.Initialize`, `Stream` ctor taking input/output `StreamParameters`, callback delegate with `ref StreamCallbackTimeInfo`) — matched our assumptions, no adjustments needed. - The NuGet package **bundles `libportaudio.so` for linux-arm64** in its runtime folder, so a self-contained publish doesn't need to rely on the Pi's system `libportaudio2`. (Both are present; the bundled one wins.) - A single `dotnet publish -c Release -r linux-arm64 --self-contained` produces a ~70 MB directory that runs as-is on the Pi. **No .NET install on the Pi.** Deploy is `dotnet publish → scp -r → ssh ./Probe`, wrapped in `bin/probe-cs`. - The `PA_ALSA_PLUGHW=1` trick from the original findings is still required in C# — but with a real twist (see next section). - Recording, RMS metering, WAV writing, and playback all behave identically to the Python prototype. Pass criteria 1 (recognisable playback), 2 (meter responsive), 3 (no over/underflow) all met. ### Key gotcha — .NET env vars do NOT propagate to libc `getenv()` on Linux `Environment.SetEnvironmentVariable("PA_ALSA_PLUGHW", "1")` updates .NET's process-table mirror but does **not** call libc `setenv`. libportaudio reads its config via `getenv("PA_ALSA_PLUGHW")` and sees nothing. ALSA then opens the USB Speaker Phone as `hw:3,0` (native 48 kHz only), and the 16 kHz stream fails with `paInvalidSampleRate`. Python's `os.environ[x] = y` calls `setenv` synchronously, which is why the Python probe never tripped on this. **Fix:** P/Invoke libc `setenv` directly, alongside the managed call: ```csharp [DllImport("libc", EntryPoint = "setenv")] static extern int setenv(string name, string value, int overwrite); Environment.SetEnvironmentVariable("PA_ALSA_PLUGHW", "1"); // for .NET-side readers setenv("PA_ALSA_PLUGHW", "1", 1); // for the C runtime view PortAudio reads ``` **Implication for future C# probes and the eventual Pi client:** any env knob a C library reads (PortAudio, onnxruntime, any P/Invoked native) must be set via libc `setenv`, not just `Environment.SetEnvironmentVariable`. Bake a small `Libc.setenv` helper into the eventual client. ### Small .NET ergonomics - `true` is required in the csproj for the callback's `Buffer.MemoryCopy` over `short*`. - `ImplicitUsings` pulls in `System.IO.Stream`, which collides with `PortAudioSharp.Stream`. Disambiguate with `using Stream = PortAudioSharp.Stream;`. - Top-level statements don't auto-pick up the project's own namespace. Add `using RecordPlayProbe;` (or whatever the project's namespace is) so the entry-point file can reach internal helpers. ### Open questions deferred to Test 2 - **No equivalent to openwakeword exists in .NET.** The choices are (a) port the openwakeword inference pipeline to C# using `Microsoft.ML.OnnxRuntime` (mel extractor → Google speech embedder → keyword classifier, three ONNX files), (b) use Picovoice Porcupine's .NET SDK (commercial, key-gated), (c) Vosk + keyword spotting (heavier). Decision deferred to the Test 2 brainstorm. - The `Libc.setenv` workaround will likely be needed again for any onnxruntime env knob on the Pi. ### Reference paths (C# probe) - Spec: `docs/superpowers/specs/2026-06-12-record-play-csharp-probe-design.md` - Plan: `docs/superpowers/plans/2026-06-12-record-play-csharp-probe.md` - Code: `tests/01-record-play-cs/` - Deploy: `bin/probe-cs` ## C# wakeword probe outcome (2026-06-12 — Test 2 ported to C#) The Python Test 2 (openwakeword "alexa" listener with beep on detect) was re-implemented in C# / .NET 9. openwakeword's streaming inference pipeline — three ONNX models chained: mel-spectrogram → Google speech embedding → "alexa" classifier — was ported to `Microsoft.ML.OnnxRuntime`. Probe lives at `tests/02-wakeword-cs/`; deploy script `bin/probe-cs-2`. All three pass criteria from the original spec held on hardware. ### What we proved - **Microsoft.ML.OnnxRuntime 1.26.0** runs on linux-arm64 from a self-contained `dotnet publish`. The NuGet package **bundles** `libonnxruntime.so` and `libonnxruntime_providers_shared.so` for the RID — no separate native package, no fallback needed. - The three openwakeword ONNX models (`melspectrogram.onnx`, `embedding_model.onnx`, `alexa.onnx`, vendored from openwakeword 0.6.0 on the Pi, SHA-256-pinned in commit `849ad47`) load and chain correctly when fed real audio. Predict() loop returned plausible scores from the first run on hardware. - **Two PortAudio streams on one USB device works.** The persistent input stream stays open while a per-detection fire-and-forget output stream plays the beep. `stream.Start()` on the second stream didn't throw, the input callback kept ticking during the 200 ms beep — no input overflow during playback. This fixes the bug `findings.md` § A flagged in the Python probe (where `sd.wait()` stalled input for the beep duration). - Hardware test on the Pi: detection fires reliably on "alexa" at normal volume from ~1 m, audible beep on each. CPU peak **31%** of one core, typical **17%** — well under the 50% budget. ### Key gotchas - **The mel-spectrogram model's output shape was wrong in our initial port.** The plan's verbatim code assumed `(1, n_frames, 1, 32)` based on a memory-derived shape. The actual shape is `(1, 1, n_frames, 32)` — `n_frames` lives at `Dimensions[2]`, not `Dimensions[1]`. Caught during Task 3 implementation by the InputMetadata-vs-derived-shape comparison done in Task 1's sanity spike. Fix is documented in `WakewordModel.cs`. The lesson: always print the actual model metadata on first contact with a vendored ONNX file; do not infer shapes from upstream code without verification. - **ORT 1.26.0's GPU-detection logger initializes BEFORE managed code can set env vars.** `Libc.setenv("ORT_LOGGING_LEVEL", "3", 1)` alone does NOT suppress the `Failed to detect devices under "/sys/class/drm/card0"` warnings — by the time .NET code runs `setenv`, ORT's C++ `device_discovery` already logged. Workaround: explicit `OrtEnv.CreateInstanceWithOptions` with `logLevel = ORT_LOGGING_LEVEL_ERROR` BEFORE constructing any `InferenceSession`. The `Libc.setenv` call stays in place for consistency with the PortAudio pattern but does no work for ORT itself on this version. Bake the `OrtEnv` initialization into the eventual client. - **Fire-and-forget output streams need careful lifetime management.** Initial implementation had `Task.Run` waiting on a `ManualResetEventSlim` BEFORE `stream.Start()` ran. If `Start()` had thrown (the exact failure mode the dual-stream test exists to surface), the cleanup task would have blocked forever on an event that would never fire — silently leaking the stream and hanging the diagnostic. Fix: run `Start()` synchronously under try/catch first, dispose + rethrow on failure, only then launch the cleanup task. Also added a 2 s timeout on `done.Wait()` — PortAudio silently swallows callback exceptions, so without the timeout a faulting callback would hang cleanup the same way. - **Vendored alexa classifier has a static batch dim.** `alexa.onnx` declares input shape `[1, 16, 96]` (not `[-1, 16, 96]` like a typical dynamic-batch model). Other two models use dynamic batch. The shape-assertion code intentionally ignores dim 0 so static-`1` and dynamic-`-1` both pass; the input tensor is created with batch=1 either way. - **Workstation Ctrl-C needs `ssh -t`.** `bin/probe-cs-2` uses `ssh -t` (force pseudo-TTY) so workstation Ctrl-C propagates as SIGINT to the remote `Probe`, triggering the `Console.CancelKeyPress` graceful-shutdown path. Test 1 didn't need this because the probe self-terminated after 5 s; Test 2 listens until interrupted. - **Inference latency was not measured.** Frame size is 1280 samples (80 ms); if `WakewordModel.Predict` ever exceeds 80 ms, the bounded `BlockingCollection` (capacity 16) will start dropping frames with `[status] consumer behind` warnings. No such warnings were observed during the hardware test, but we did not instrument `Predict` timing. Add this when promoting to the main assistant. ### Open questions for the main assistant - **Custom wakeword model.** Stock "alexa" works as a probe, but the shipped assistant needs a custom phrase (e.g. "hey assistant"). openwakeword has a Colab notebook for synthetic training; budget ~1 hour to produce a model that this same C# pipeline can load (it's a drop-in `.onnx` replacement for `alexa.onnx`). - **Speaker-is-mic self-trigger.** The 880 Hz 200 ms beep used here is far enough from speech-band that it doesn't re-trigger the classifier. Real TTS playback will. Spec § B recommends detector gating during playback as the v1 fix; revisit AEC only if barge-in becomes a requirement. - **Inference timing instrumentation.** Add per-`Predict` latency logging in the main assistant so consumer-behind warnings are diagnosable instead of opaque. ### Reference paths (C# wakeword probe) - Spec: `docs/superpowers/specs/2026-06-12-test-2-wakeword-csharp-design.md` - Plan: `docs/superpowers/plans/2026-06-12-test-2-wakeword-csharp.md` - Code: `tests/02-wakeword-cs/` - Deploy: `bin/probe-cs-2` ## C# full-cycle outcome (2026-06-12 — Test 3) Test 3 stitched the C# Test 1 (record + play) and Test 2 (wakeword + beep) into one state machine: idle → "alexa" → beep → record 5 s → playback → idle. Probe lives at `tests/03-full-cycle-cs/`; deploy script `bin/probe-cs-3`. All seven spec pass criteria held on hardware on the first run-after-fix (see "Key gotchas" below). ### What we proved - **The single-input-stream + state-machine + per-cycle-output-streams shape works on this hardware.** One PortAudio input stream open for the process lifetime feeds a `BlockingCollection` (bounded 16); a single consumer thread dispatches each 80 ms frame through IDLE / BEEPING / RECORDING / PLAYBACK. No locks, no `Interlocked` — `state` is consumer-thread-local; the only cross-thread datum is a `volatile bool` in a `PlaybackDoneFlag`. - **Detector gating during PLAYBACK prevents self-trigger.** The recorded user audio comes back out the speaker and re-enters the input stream; because the consumer's PLAYBACK branch discards frames (never calls `model.Predict`), the playback never re-triggers detection. - **`WakewordModel.Reset()` on PLAYBACK→IDLE re-arms the warmup guard correctly.** Five lines (`_rawRingFill = 0; _melRing.Clear(); _embRing.Clear(); _framesSeen = 0`) restore the model to the same state it has just after construction; the existing `WarmupFrames = 16` then suppresses the next ~1.3 s of `Predict` calls until the rings refill with fresh post-playback audio. Next "alexa" after playback ends is detected reliably and quickly (within target). - **3+ cycles back-to-back, no degradation.** Detection rate holds, CPU stays in IDLE-state range between cycles, no stuck states. - **No `[status]` or `[error]` warnings during normal operation.** Input doesn't overflow during PLAYBACK (the consumer keeps draining the queue into the discard branch), consumer doesn't fall behind, no fire-and-forget timeouts after the fix below. - **Ctrl-C from the workstation cleanly exits the probe.** Same `ssh -t` + `Console.CancelKeyPress` + `cts.Cancel()` pattern as Test 2. ### Key gotchas - **The fire-and-forget cleanup task's `done.Wait` timeout must scale to the buffer length, not be a flat constant.** Test 2's `FireAndForgetBeep` used `done.Wait(TimeSpan.FromSeconds(2))` to guard against silently-faulted PortAudio callbacks (PortAudio swallows callback exceptions). For a 200 ms beep, 2 s is generous; the cleanup task's wait always returns "set", then sleeps 200 ms and disposes. Test 3's `FireAndForgetPlayback` initially inherited the same flat 2 s timeout — but the recording playback runs for ~5 s, so the timeout fired mid-playback, the cleanup task called `stream.Stop()` 200 ms later, and the audio cut at ~2.2 s. **First hardware run produced "[error] recording playback timed out — callback may have faulted" on every cycle.** Fix: compute `expectedSeconds = buf.Length / SampleRate` and use `TimeSpan.FromSeconds(expectedSeconds + 2.0)` as the cleanup wait. Commit `a7d408e`. - **Partial deploys are silently wrong: a self-contained .NET publish uses `Probe` (native launcher) + `Probe.dll` (managed code).** Copying only `Probe` after a rebuild leaves the old `.dll` on the target, which the new launcher loads — so the binary on the Pi reflects code from the previous build. Always wipe-and-replace the whole publish directory (the pattern already in `bin/probe-cs-3`: `rm -rf ~/probe-cs-3 && mkdir ~/probe-cs-3 && scp -r $PUBLISH_DIR/. pi:~/probe-cs-3/`). Don't do incremental `scp Probe`. - **`PlaybackDoneFlag` should be created per cycle, not reused.** A theoretical late `Set()` from a delayed cleanup `Task.Run` (cycle N) could land on the shared flag while the consumer is in cycle N+1's PLAYBACK, causing a spurious immediate PLAYBACK→IDLE that skips playback. Allocating a fresh `PlaybackDoneFlag()` inside the RECORDING→PLAYBACK transition (and reading with `playbackDone!.Consume()` in PLAYBACK) sends the stale `Set()` to a now-unreachable object. Commit `070cc07`. Not observed in practice but cheap to defend. - **Log the cleanup-task's `done.Wait` timeout return.** Without the log, a silently-faulted PortAudio callback produces "PLAYBACK→IDLE with no audio" with no diagnostic — exactly the failure mode the timeout was supposed to defend against. Both `FireAndForgetBeep` and `FireAndForgetPlayback` now log `[error] ... timed out — callback may have faulted` on `!done.Wait(timeout)`. Commit `070cc07`. ### Implications for the main assistant - The whole shape — persistent input stream, single-consumer state machine, per-cycle output streams, fire-and-forget cleanup with timeout scaled to buffer length — is good to lift directly into the main Pi client. The main assistant adds: Opus encoding on the recorded buffer before sending to the Realtime WebSocket, an Opus decoder feeding the output stream during TTS playback, and replacing the local 5 s recording window with VAD-driven endpointing. None of those replace the state-machine shape; they slot in. - `WakewordModel` should be promoted out of `tests/` into a real source tree (e.g. `src/Audio/Wakeword/`) when the main assistant starts. Test 3's copy is the live version (it has `Reset()`); Test 2's is frozen. - The "always wipe-and-replace the whole publish directory" deploy pattern needs to be the default in the main assistant's deploy script too. A partial `scp` of `Probe` is a footgun every time the .dll changes (i.e. every code change). ### Reference paths (C# full-cycle probe) - Spec: `docs/superpowers/specs/2026-06-12-test-3-full-cycle-csharp-design.md` - Plan: `docs/superpowers/plans/2026-06-12-test-3-full-cycle-csharp.md` - Code: `tests/03-full-cycle-cs/` - Deploy: `bin/probe-cs-3` ## Reference paths - Spec: `docs/superpowers/specs/2026-06-11-voice-assistant-prototypes-design.md` - Plan: `docs/superpowers/plans/2026-06-11-voice-assistant-prototypes.md` - Test 1: `tests/01-record-play/main.py` - Test 2: `tests/02-wakeword/main.py` - Deploy helper: `bin/deploy` - Pi bootstrap: `bin/bootstrap` - Pi access: `pi@192.168.50.115`, password `assistant` (see `CLAUDE.md`)