diff --git a/docs/superpowers/plans/2026-06-12-test-2-wakeword-csharp.md b/docs/superpowers/plans/2026-06-12-test-2-wakeword-csharp.md new file mode 100644 index 0000000..7050914 --- /dev/null +++ b/docs/superpowers/plans/2026-06-12-test-2-wakeword-csharp.md @@ -0,0 +1,969 @@ +# Test 2 wakeword C# probe — implementation plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Re-implement Test 2 (listen for the "alexa" wakeword via openwakeword's stock model → beep on detection) in C# on the Raspberry Pi, by porting `openwakeword/model.py`'s streaming inference pipeline to C# / `Microsoft.ML.OnnxRuntime`. Pass the same hardware criteria as the Python version. + +**Architecture:** Single-folder throwaway probe under `tests/02-wakeword-cs/`. `WakewordModel.cs` chains three vendored ONNX models (mel-spectrogram → Google speech embedding → alexa keyword classifier) with the same streaming buffer geometry as the Python implementation. `Program.cs` drives a PortAudio input stream into a bounded `BlockingCollection`; the main thread pulls frames, runs `WakewordModel.Predict`, and dispatches a fire-and-forget output stream for the detection beep so input ingestion never stalls. `bin/probe-cs-2` builds a self-contained linux-arm64 binary and ships it to the Pi. + +**Tech Stack:** .NET 9, C#, PortAudioSharp2 (NuGet), Microsoft.ML.OnnxRuntime (NuGet), bash, sshpass. Target host: Pi at `192.168.50.115`, user `pi`, password `assistant` (see `CLAUDE.md`). The Pi gets no .NET install — the binary is self-contained. + +--- + +## File structure + +| Path | Created in task | Purpose | +| --- | --- | --- | +| `tests/02-wakeword-cs/Probe.csproj` | Task 1 | net9.0 console project; linux-arm64 RID; PortAudioSharp2 + Microsoft.ML.OnnxRuntime refs; bundles `models/*.onnx` to output. | +| `tests/02-wakeword-cs/models/melspectrogram.onnx` | Task 1 | Vendored from the Pi's openwakeword install (~14 KB). | +| `tests/02-wakeword-cs/models/embedding_model.onnx` | Task 1 | Vendored from the Pi (~10 MB), Google's speech embedding model. | +| `tests/02-wakeword-cs/models/alexa.onnx` | Task 1 | Vendored from the Pi (~1.5 MB), stock alexa classifier. | +| `tests/02-wakeword-cs/Program.cs` | Task 1 (spike), rewritten in Task 4 & extended in Task 5 | Entry point: env vars → PortAudio init → device pick → input stream → main inference loop → optional beep. | +| `tests/02-wakeword-cs/WakewordModel.cs` | Task 3 | Loads the three sessions; `Predict(short[] frame1280) → float`. | +| `bin/probe-cs-2` | Task 2 | `dotnet publish` → `scp` → `ssh -t` flow. | + +`tests/02-wakeword/` (Python), `tests/01-record-play-cs/` (Test 1 C#), and `bin/probe-cs` are untouched. + +--- + +## Task 1: Scaffold project + vendor ONNX models + ORT sanity spike on the Pi + +**Goal:** Stand up the .NET project with both NuGet refs, vendor the three openwakeword ONNX models into the repo, and prove on the Pi that `Microsoft.ML.OnnxRuntime` loads all three from a self-contained linux-arm64 publish. Capture each model's declared input / output shape (we'll use these as ground-truth assertions in Task 3). + +**Files:** +- Create: `tests/02-wakeword-cs/Probe.csproj` +- Create: `tests/02-wakeword-cs/models/melspectrogram.onnx` (vendored) +- Create: `tests/02-wakeword-cs/models/embedding_model.onnx` (vendored) +- Create: `tests/02-wakeword-cs/models/alexa.onnx` (vendored) +- Create: `tests/02-wakeword-cs/Program.cs` (spike — replaced in Task 4) + +- [ ] **Step 1: Pin the Microsoft.ML.OnnxRuntime version** + +Use Context7 to fetch the current Microsoft.ML.OnnxRuntime NuGet docs and identify the latest stable version (or check `https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime`): + +```sh +curl -s https://api.nuget.org/v3-flatcontainer/microsoft.ml.onnxruntime/index.json | jq -r '.versions[]' | grep -v -E '(rc|preview|alpha|beta)' | tail -1 +``` + +Note the exact version (call it ``). Pin this in the csproj — do NOT use a floating `*`. + +Same drill for PortAudioSharp2 (already pinned at 1.0.6 in Test 1; use that exact version unless a newer non-prerelease has shipped): + +```sh +curl -s https://api.nuget.org/v3-flatcontainer/portaudiosharp2/index.json | jq -r '.versions[-1]' +``` + +- [ ] **Step 2: Vendor the three ONNX files from the Pi** + +```sh +mkdir -p tests/02-wakeword-cs/models +sshpass -p assistant scp \ + pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/melspectrogram.onnx \ + pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/embedding_model.onnx \ + pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/alexa.onnx \ + tests/02-wakeword-cs/models/ +ls -l tests/02-wakeword-cs/models/ +sha256sum tests/02-wakeword-cs/models/*.onnx +``` + +Expected: three files, roughly 14 KB / 10 MB / 1.5 MB. Record the SHA-256 hashes in your commit message in Step 8 — that pins the exact model files used. + +If `scp` fails because the Pi's openwakeword venv lives elsewhere, run: + +```sh +sshpass -p assistant ssh pi@192.168.50.115 'find ~ -path "*openwakeword/resources/models/*.onnx" 2>/dev/null' +``` + +and substitute the correct directory. + +- [ ] **Step 3: Write `tests/02-wakeword-cs/Probe.csproj`** + +```xml + + + Exe + net9.0 + WakewordProbe + Probe + enable + enable + linux-arm64 + true + false + true + true + + + + + + + + PreserveNewest + + + +``` + +Replace `` with the version from Step 1. + +- [ ] **Step 4: Write the spike `tests/02-wakeword-cs/Program.cs`** + +This temporary Program.cs loads each model and prints its declared shapes — exactly what Task 3 needs to assert against. It is replaced in Task 4. + +```csharp +using Microsoft.ML.OnnxRuntime; + +string modelsDir = Path.Combine(AppContext.BaseDirectory, "models"); +string[] modelFiles = { "melspectrogram.onnx", "embedding_model.onnx", "alexa.onnx" }; + +var opts = new SessionOptions +{ + IntraOpNumThreads = 1, + InterOpNumThreads = 1, + LogSeverityLevel = OrtLoggingLevel.ORT_LOGGING_LEVEL_ERROR, +}; + +foreach (var f in modelFiles) +{ + string path = Path.Combine(modelsDir, f); + Console.WriteLine($"== {f} =="); + using var session = new InferenceSession(path, opts); + foreach (var kv in session.InputMetadata) + { + var dims = string.Join(",", kv.Value.Dimensions); + Console.WriteLine($" input '{kv.Key}': shape=[{dims}] dtype={kv.Value.ElementType.Name}"); + } + foreach (var kv in session.OutputMetadata) + { + var dims = string.Join(",", kv.Value.Dimensions); + Console.WriteLine($" output '{kv.Key}': shape=[{dims}] dtype={kv.Value.ElementType.Name}"); + } +} +Console.WriteLine("All three models loaded."); +``` + +- [ ] **Step 5: Build and publish locally** + +```sh +dotnet publish tests/02-wakeword-cs -c Release -r linux-arm64 --self-contained \ + -o /tmp/probe-cs-2-out +``` + +Expected: build succeeds. Confirm the publish dir contains the binary, the three model files, and the bundled ONNX Runtime native: + +```sh +ls /tmp/probe-cs-2-out | grep -E 'libonnxruntime|libportaudio|Probe' +ls /tmp/probe-cs-2-out/models/ +``` + +Expected: `Probe`, `libonnxruntime.so` (or `.so.`), `libportaudio.so` (or similar), plus the three `.onnx` files under `models/`. If `libonnxruntime.so` is missing, the `Microsoft.ML.OnnxRuntime` package for your version may have split the native runtime — install the matching native package (`Microsoft.ML.OnnxRuntime.Native.` if it exists for ``) or downgrade `` to a version where the linux-arm64 native is included in the main package. Do NOT proceed to Step 6 if the native is missing. + +- [ ] **Step 6: Push the spike to the Pi and run it** + +```sh +sshpass -p assistant ssh pi@192.168.50.115 'rm -rf ~/probe-cs-2 && mkdir ~/probe-cs-2' +sshpass -p assistant scp -r /tmp/probe-cs-2-out/. pi@192.168.50.115:~/probe-cs-2/ +sshpass -p assistant ssh pi@192.168.50.115 'cd ~/probe-cs-2 && chmod +x Probe && ./Probe' +``` + +Expected output (numbers and names matching the openwakeword Python source from `/tmp/oww-utils.py`): + +``` +== melspectrogram.onnx == + input '': shape=[-1,-1] dtype=Single (1D audio, variable length) + output '': shape=[1,-1,1,32] dtype=Single (mel: 32 bins, variable n_frames) +== embedding_model.onnx == + input 'input_1': shape=[-1,76,32,1] dtype=Single (76 mel frames × 32 bins × 1 channel, variable batch) + output '': shape=[-1,1,1,96] dtype=Single (96-d embedding per batch) +== alexa.onnx == + input '': shape=[-1,16,96] dtype=Single (16 embeddings × 96 dims, variable batch) + output '': shape=[-1,1] dtype=Single (binary classifier score) +All three models loaded. +``` + +Negative dims (`-1`) mean "dynamic / batch dimension" — that's fine, we'll pass `1` at inference time. + +**Save the actual printed output to your scratchpad** — the input/output names (`'input_1'` etc.) and exact dimension lists become the assertion values in Task 3 Step 4. The dimension numbers (76, 32, 16, 96) should match the Python source; if they don't, stop and investigate before proceeding. + +- [ ] **Step 7: Mark `bin/probe-cs-2` placeholder NOT created yet** + +Task 2 creates the deploy script. For Task 1 we use the inline scp commands above. (Task 1 cannot depend on Task 2.) + +- [ ] **Step 8: Commit** + +```sh +git add tests/02-wakeword-cs/Probe.csproj tests/02-wakeword-cs/Program.cs tests/02-wakeword-cs/models/ +git commit -m "Scaffold C# wakeword probe: vendored ONNX models + ORT sanity spike + +Models sourced from openwakeword==0.6.0 on the Pi. +SHA-256: + melspectrogram.onnx + embedding_model.onnx + alexa.onnx " +``` + +--- + +## Task 2: Wrap the deploy flow as `bin/probe-cs-2` + +**Goal:** Capture the build / scp / ssh sequence as a single script so subsequent tasks just run `bin/probe-cs-2`. Mirrors `bin/probe-cs` with new paths and adds `ssh -t` so workstation Ctrl-C propagates as SIGINT to the long-running probe. + +**Files:** +- Create: `bin/probe-cs-2` + +- [ ] **Step 1: Write `bin/probe-cs-2`** + +```sh +#!/usr/bin/env bash +set -euo pipefail + +PI_HOST="${PI_HOST:-192.168.50.115}" +PI_USER="${PI_USER:-pi}" +PI_PASS="${PI_PASS:-assistant}" +PUBLISH_DIR="${PUBLISH_DIR:-/tmp/probe-cs-2-out}" + +repo_root="$(cd "$(dirname "$0")/.." && pwd)" +cd "$repo_root" + +echo ">> dotnet publish (linux-arm64, self-contained)" +dotnet publish tests/02-wakeword-cs \ + -c Release -r linux-arm64 --self-contained \ + -o "$PUBLISH_DIR" + +echo ">> scp to $PI_USER@$PI_HOST:~/probe-cs-2/ (wipe-and-replace)" +sshpass -p "$PI_PASS" ssh -o StrictHostKeyChecking=accept-new \ + "$PI_USER@$PI_HOST" 'rm -rf ~/probe-cs-2 && mkdir ~/probe-cs-2' +sshpass -p "$PI_PASS" scp -r "$PUBLISH_DIR/." \ + "$PI_USER@$PI_HOST:~/probe-cs-2/" + +echo ">> ssh + run on Pi (Ctrl-C from this terminal stops the probe)" +sshpass -p "$PI_PASS" ssh -t -o StrictHostKeyChecking=accept-new \ + "$PI_USER@$PI_HOST" 'cd ~/probe-cs-2 && chmod +x Probe && ./Probe' +``` + +The `ssh -t` (force pseudo-tty) makes Ctrl-C from the workstation propagate as SIGINT to the remote `./Probe` process — needed because Task 4's main loop runs until Ctrl-C. Test 1's `bin/probe-cs` didn't need it because the probe self-terminated after 5 s of recording + playback. + +- [ ] **Step 2: Mark it executable** + +```sh +chmod +x bin/probe-cs-2 +``` + +- [ ] **Step 3: Smoke-test against Task 1's spike** + +```sh +bin/probe-cs-2 +``` + +Expected: same shape-dump output as Task 1 Step 6, ending with `All three models loaded.` Ctrl-C is not needed because the spike exits on its own. If it errors, fix the script — do not work around in later tasks. + +- [ ] **Step 4: Commit** + +```sh +git add bin/probe-cs-2 +git commit -m "Add bin/probe-cs-2: build + scp + run the C# wakeword probe on the Pi" +``` + +--- + +## Task 3: `WakewordModel.cs` — port openwakeword's streaming pipeline + +**Goal:** A `WakewordModel` class that wraps the three ONNX sessions and exposes a single `Predict(short[] frame1280) → float` method. Internal state matches `openwakeword/AudioFeatures._streaming_features()` and `Model.predict()` from the openwakeword 0.6.0 source on the Pi (already fetched to `/tmp/model.py` and `/tmp/oww-utils.py` during plan-writing — re-fetch if needed: `sshpass -p assistant scp pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/{model,utils}.py /tmp/`). + +**Files:** +- Create: `tests/02-wakeword-cs/WakewordModel.cs` +- Modify: `tests/02-wakeword-cs/Program.cs` (replace spike with a load-and-print test) + +### Pipeline geometry (derived from openwakeword/utils.py `AudioFeatures`) + +Per call to `Predict(short[] frame1280)`: + +1. **Raw-audio buffer.** Append the 1280 new samples to a ring of recent samples. Keep at least the last `1280 + 480 = 1760` samples (the +480 = `160 * 3` is the mel model's 30 ms warm-up context that openwakeword passes in via the `-n_samples-160*3:` slice in `_streaming_melspectrogram`). +2. **Mel stage.** Slice the last 1760 samples → run `melspectrogram.onnx` with input shape `(1, 1760)` dtype Single → output is `(1, n_frames, 1, 32)`; squeeze trailing dims to get `(n_frames, 32)` where `n_frames` ≈ 8 (`ceil(1760/160 - 3) = ceil(8) = 8`). Apply the openwakeword post-transform `x = x / 10 + 2` element-wise. Append the new mel frames to the mel-frame ring. Trim ring to the last 970 frames (= `10 * 97`, openwakeword's `melspectrogram_max_len`). +3. **Embedding stage.** Slice the last 76 mel frames from the ring → shape `(1, 76, 32, 1)`. Run `embedding_model.onnx` (input name `input_1`) → output shape `(1, 1, 1, 96)`; squeeze to a `float[96]` embedding. Append to the embedding ring. Trim ring to the last 120 entries (openwakeword's `feature_buffer_max_len`). +4. **Classifier stage.** Slice the last 16 embeddings from the ring → shape `(1, 16, 96)`. Run `alexa.onnx` → output shape `(1, 1)`; return `output[0, 0]` as the score. +5. **Initialization gate.** Return `0f` for the first 16 calls — until the embedding ring has 16 real entries, classifier output is meaningless. (openwakeword pre-populates with random-audio embeddings then zeros the first 5 scores; we do the simpler equivalent: skip outright until we have 16 real embeddings.) + +- [ ] **Step 1: Document the geometry in `WakewordModel.cs`** + +Create `tests/02-wakeword-cs/WakewordModel.cs` and start with constants and the class skeleton: + +```csharp +using Microsoft.ML.OnnxRuntime; +using Microsoft.ML.OnnxRuntime.Tensors; + +namespace WakewordProbe; + +// Streaming wakeword inference port of openwakeword 0.6.0's predict pipeline. +// References (read alongside this file): +// openwakeword/utils.py:AudioFeatures._streaming_features (mel + embedding stages) +// openwakeword/utils.py:AudioFeatures._streaming_melspectrogram +// openwakeword/utils.py:AudioFeatures._get_embeddings +// openwakeword/model.py:Model.predict (classifier stage) +internal sealed class WakewordModel : IDisposable +{ + // Audio I/O geometry — per-call input contract. + public const int FrameSamples = 1280; // 80 ms @ 16 kHz; each Predict() call. + public const int SampleRate = 16_000; + + // Mel model: input is last (FrameSamples + MelContextSamples) raw samples + // -- the +480 is openwakeword's `-n_samples-160*3:` slice (3 hops of context). + public const int MelContextSamples = 480; // 160 * 3 + public const int MelInputSamples = FrameSamples + MelContextSamples; // 1760 + public const int MelBins = 32; // melspectrogram model output dim + public const int MelBufferMaxFrames = 970; // 10 * 97 (openwakeword `melspectrogram_max_len`) + + // Embedding model: 76-mel-frame window in, 96-d embedding out. + public const int EmbeddingWindowMelFrames = 76; + public const int EmbeddingDim = 96; + public const int EmbeddingBufferMax = 120; // openwakeword `feature_buffer_max_len` + public const string EmbeddingInputName = "input_1"; // openwakeword convention; assert at startup + + // Classifier: 16 embeddings in, scalar score out. + public const int ClassifierEmbeddings = 16; + + // Skip the first N Predict() calls — buffer fill-up window. + public const int WarmupFrames = ClassifierEmbeddings; // 16 frames ≈ 1.28 s + + private readonly InferenceSession _mel; + private readonly InferenceSession _emb; + private readonly InferenceSession _cls; + + private readonly string _melInputName; // discovered at startup + private readonly string _clsInputName; // discovered at startup + + private readonly short[] _rawRing = new short[MelInputSamples]; + private int _rawRingFill = 0; // samples buffered (≤ MelInputSamples) + + private readonly List _melRing = new(MelBufferMaxFrames); // each entry is a length-32 mel frame + private readonly List _embRing = new(EmbeddingBufferMax); // each entry is a length-96 embedding + + private int _framesSeen = 0; + + public WakewordModel(string melPath, string embeddingPath, string classifierPath) + { + var opts = new SessionOptions + { + IntraOpNumThreads = 1, + InterOpNumThreads = 1, + LogSeverityLevel = OrtLoggingLevel.ORT_LOGGING_LEVEL_ERROR, + }; + + _mel = new InferenceSession(melPath, opts); + _emb = new InferenceSession(embeddingPath, opts); + _cls = new InferenceSession(classifierPath, opts); + + // Discover input names + assert shape geometry. + _melInputName = _mel.InputMetadata.Keys.Single(); + + AssertEmbeddingShape(_emb); // input_1: [batch, 76, 32, 1], dtype float + _clsInputName = _cls.InputMetadata.Keys.Single(); + AssertClassifierShape(_cls); // [batch, 16, 96], dtype float + } + + // ... methods follow in Step 2 ... + + public void Dispose() + { + _mel.Dispose(); + _emb.Dispose(); + _cls.Dispose(); + } +} +``` + +- [ ] **Step 2: Implement the shape-assertion helpers** + +Append inside the `WakewordModel` class: + +```csharp + private static void AssertEmbeddingShape(InferenceSession sess) + { + if (!sess.InputMetadata.TryGetValue(EmbeddingInputName, out var meta)) + throw new InvalidOperationException( + $"embedding_model.onnx: expected input named '{EmbeddingInputName}', got [{string.Join(",", sess.InputMetadata.Keys)}]"); + var d = meta.Dimensions; + // Expected: [batch, 76, 32, 1] — batch may be -1 (dynamic). + if (d.Length != 4 || d[1] != EmbeddingWindowMelFrames || d[2] != MelBins || d[3] != 1) + throw new InvalidOperationException( + $"embedding_model.onnx: expected input shape [batch,{EmbeddingWindowMelFrames},{MelBins},1], got [{string.Join(",", d)}]"); + if (meta.ElementType != typeof(float)) + throw new InvalidOperationException($"embedding_model.onnx: expected Single input, got {meta.ElementType.Name}"); + } + + private static void AssertClassifierShape(InferenceSession sess) + { + var inputName = sess.InputMetadata.Keys.Single(); + var meta = sess.InputMetadata[inputName]; + var d = meta.Dimensions; + // Expected: [batch, 16, 96] — batch may be -1. + if (d.Length != 3 || d[1] != ClassifierEmbeddings || d[2] != EmbeddingDim) + throw new InvalidOperationException( + $"alexa.onnx: expected input shape [batch,{ClassifierEmbeddings},{EmbeddingDim}], got [{string.Join(",", d)}]"); + if (meta.ElementType != typeof(float)) + throw new InvalidOperationException($"alexa.onnx: expected Single input, got {meta.ElementType.Name}"); + } +``` + +- [ ] **Step 3: Implement `Predict`** + +Append inside the `WakewordModel` class: + +```csharp + public float Predict(short[] frame1280) + { + if (frame1280.Length != FrameSamples) + throw new ArgumentException($"Expected {FrameSamples} samples, got {frame1280.Length}"); + + // 1. Append 1280 new samples to the raw ring (shift older samples down if full). + if (_rawRingFill < MelInputSamples) + { + int copyToFront = Math.Min(MelInputSamples - _rawRingFill, FrameSamples); + Array.Copy(frame1280, 0, _rawRing, _rawRingFill, copyToFront); + _rawRingFill += copyToFront; + if (copyToFront < FrameSamples) + { + // Shouldn't happen on the very first call, but defensive. + int leftover = FrameSamples - copyToFront; + Array.Copy(_rawRing, leftover, _rawRing, 0, MelInputSamples - leftover); + Array.Copy(frame1280, copyToFront, _rawRing, MelInputSamples - leftover, leftover); + } + } + else + { + // Shift older samples left by FrameSamples, then append new at the tail. + Array.Copy(_rawRing, FrameSamples, _rawRing, 0, MelInputSamples - FrameSamples); + Array.Copy(frame1280, 0, _rawRing, MelInputSamples - FrameSamples, FrameSamples); + } + + // Skip everything until we have the full mel-context window primed. + if (_rawRingFill < MelInputSamples) + { + _framesSeen++; + return 0f; + } + + // 2. Mel stage: feed the full _rawRing as float32 (1, MelInputSamples) into mel model. + var melInputData = new float[MelInputSamples]; + for (int i = 0; i < MelInputSamples; i++) melInputData[i] = _rawRing[i]; // int16 → float32, NO normalisation + var melInputTensor = new DenseTensor(melInputData, new[] { 1, MelInputSamples }); + using var melResults = _mel.Run(new[] { + NamedOnnxValue.CreateFromTensor(_melInputName, melInputTensor) + }); + var melOut = melResults.First().AsTensor(); // shape (1, n_frames, 1, 32) + + // Apply openwakeword's `x / 10 + 2` transform and append each new frame to _melRing. + int nFrames = melOut.Dimensions[1]; + for (int f = 0; f < nFrames; f++) + { + var bin = new float[MelBins]; + for (int b = 0; b < MelBins; b++) + bin[b] = melOut[0, f, 0, b] / 10f + 2f; + _melRing.Add(bin); + } + if (_melRing.Count > MelBufferMaxFrames) + _melRing.RemoveRange(0, _melRing.Count - MelBufferMaxFrames); + + // 3. Embedding stage: need ≥ 76 mel frames; slice the last 76 → (1, 76, 32, 1). + if (_melRing.Count < EmbeddingWindowMelFrames) + { + _framesSeen++; + return 0f; + } + var embInputData = new float[EmbeddingWindowMelFrames * MelBins]; + int startMel = _melRing.Count - EmbeddingWindowMelFrames; + for (int f = 0; f < EmbeddingWindowMelFrames; f++) + Array.Copy(_melRing[startMel + f], 0, embInputData, f * MelBins, MelBins); + var embInputTensor = new DenseTensor(embInputData, new[] { 1, EmbeddingWindowMelFrames, MelBins, 1 }); + using var embResults = _emb.Run(new[] { + NamedOnnxValue.CreateFromTensor(EmbeddingInputName, embInputTensor) + }); + var embOut = embResults.First().AsTensor(); // shape (1, 1, 1, 96) + var newEmb = new float[EmbeddingDim]; + for (int i = 0; i < EmbeddingDim; i++) newEmb[i] = embOut[0, 0, 0, i]; + _embRing.Add(newEmb); + if (_embRing.Count > EmbeddingBufferMax) + _embRing.RemoveRange(0, _embRing.Count - EmbeddingBufferMax); + + _framesSeen++; + + // 4. Warm-up + classifier stage. + if (_embRing.Count < ClassifierEmbeddings || _framesSeen <= WarmupFrames) + return 0f; + + var clsInputData = new float[ClassifierEmbeddings * EmbeddingDim]; + int startEmb = _embRing.Count - ClassifierEmbeddings; + for (int i = 0; i < ClassifierEmbeddings; i++) + Array.Copy(_embRing[startEmb + i], 0, clsInputData, i * EmbeddingDim, EmbeddingDim); + var clsInputTensor = new DenseTensor(clsInputData, new[] { 1, ClassifierEmbeddings, EmbeddingDim }); + using var clsResults = _cls.Run(new[] { + NamedOnnxValue.CreateFromTensor(_clsInputName, clsInputTensor) + }); + var clsOut = clsResults.First().AsTensor(); // shape (1, 1) + return clsOut[0, 0]; + } +``` + +- [ ] **Step 4: Replace `Program.cs` spike with a load-only smoke test** + +Replace `tests/02-wakeword-cs/Program.cs` with: + +```csharp +using WakewordProbe; + +string modelsDir = Path.Combine(AppContext.BaseDirectory, "models"); +Console.WriteLine("Loading WakewordModel..."); +var t0 = DateTime.UtcNow; +using var model = new WakewordModel( + Path.Combine(modelsDir, "melspectrogram.onnx"), + Path.Combine(modelsDir, "embedding_model.onnx"), + Path.Combine(modelsDir, "alexa.onnx")); +Console.WriteLine($"Loaded in {(DateTime.UtcNow - t0).TotalSeconds:0.00}s."); + +// Smoke test: feed 50 frames of silence (1280 zero samples each). +// Expect every score to be 0f (warmup gate + no signal). +var silentFrame = new short[WakewordModel.FrameSamples]; +int nonZero = 0; +for (int i = 0; i < 50; i++) +{ + float s = model.Predict(silentFrame); + if (s != 0f) nonZero++; +} +Console.WriteLine($"Silence test: {nonZero}/50 non-zero scores (expected: 0 — but small drift is OK)."); +Console.WriteLine("Predict pipeline ran without throwing."); +``` + +The "non-zero on silence" count may be small but not literally 0 — the classifier sees ~zero input but its output is rarely exactly 0f. Anything under ~5 (and well below threshold 0.5) is fine. The point of the smoke test is: did `Predict` run end-to-end without throwing on shape mismatches? If it throws, the assertion bug is in the model load or pipeline geometry. + +- [ ] **Step 5: Build locally** + +```sh +dotnet build tests/02-wakeword-cs -c Release +``` + +Expected: success. Common compile errors: + +- `NamedOnnxValue` / `DenseTensor` / `InferenceSession` not found → confirm `using Microsoft.ML.OnnxRuntime;` and `using Microsoft.ML.OnnxRuntime.Tensors;` are at top. +- `'InputMetadata' has no Single()` → add `using System.Linq;` (with `ImplicitUsings` enabled this should already be there, but the spike didn't need it). +- API surface drift in `Microsoft.ML.OnnxRuntime` `` (the package occasionally renames `NamedOnnxValue.CreateFromTensor`) → consult Context7 `/microsoft/onnxruntime` docs for the pinned version. Do NOT silently switch to a different API; document the version-specific shape if it diverges. + +- [ ] **Step 6: Run on the Pi via `bin/probe-cs-2`** + +```sh +bin/probe-cs-2 +``` + +Expected output: + +``` +Loading WakewordModel... +Loaded in 0.XXs. +Silence test: /50 non-zero scores (expected: 0 — but small drift is OK). +Predict pipeline ran without throwing. +``` + +If the run aborts with an `InvalidOperationException` from one of the shape assertions, the model in the repo disagrees with what this plan expects — STOP. Either the vendored ONNX file is from a newer openwakeword version with a different shape (re-verify the SHA-256 hashes against Step 8 of Task 1), or our shape derivation from `oww-utils.py` was wrong. Investigate by re-running the Task 1 spike and comparing dimensions before adjusting `WakewordModel.cs`. + +- [ ] **Step 7: Commit** + +```sh +git add tests/02-wakeword-cs/WakewordModel.cs tests/02-wakeword-cs/Program.cs +git commit -m "WakewordModel: port openwakeword streaming pipeline (mel→emb→cls)" +``` + +--- + +## Task 4: `Program.cs` — live input stream + main inference loop (detection-only, no beep) + +**Goal:** Replace the load-only smoke test with the full input stream + inference loop. Detection lines print to stdout but no beep yet — we want detection working in isolation before adding the second PortAudio stream. + +**Files:** +- Modify: `tests/02-wakeword-cs/Program.cs` (full rewrite) + +- [ ] **Step 1: Rewrite `tests/02-wakeword-cs/Program.cs`** + +```csharp +using System.Collections.Concurrent; +using System.Diagnostics; +using System.Runtime.InteropServices; +using PortAudioSharp; +using WakewordProbe; +using Stream = PortAudioSharp.Stream; + +const int SampleRate = WakewordModel.SampleRate; +const int Channels = 1; +const uint BlockFrames = WakewordModel.FrameSamples; // 1280 +const float Threshold = 0.5f; +const double CooldownSeconds = 1.0; + +// .NET-side mirror for any readers that go through Environment.GetEnvironmentVariable, +// AND libc setenv so PortAudio / onnxruntime (which use getenv()) see the values. +Environment.SetEnvironmentVariable("PA_ALSA_PLUGHW", "1"); +Libc.setenv("PA_ALSA_PLUGHW", "1", 1); +Environment.SetEnvironmentVariable("ORT_LOGGING_LEVEL", "3"); +Libc.setenv("ORT_LOGGING_LEVEL", "3", 1); + +PortAudio.Initialize(); +WakewordModel? model = null; +try +{ + int device = FindUsbDevice(); + Console.WriteLine($"Using device {device} ('{PortAudio.GetDeviceInfo(device).name}')"); + + Console.WriteLine("Loading WakewordModel..."); + var t0 = DateTime.UtcNow; + string modelsDir = Path.Combine(AppContext.BaseDirectory, "models"); + model = new WakewordModel( + Path.Combine(modelsDir, "melspectrogram.onnx"), + Path.Combine(modelsDir, "embedding_model.onnx"), + Path.Combine(modelsDir, "alexa.onnx")); + Console.WriteLine($"Loaded in {(DateTime.UtcNow - t0).TotalSeconds:0.00}s."); + + var queue = new BlockingCollection(boundedCapacity: 16); + using var cts = new CancellationTokenSource(); + Console.CancelKeyPress += (_, e) => { e.Cancel = true; cts.Cancel(); }; + + var inParams = new StreamParameters + { + device = device, + channelCount = Channels, + sampleFormat = SampleFormat.Int16, + suggestedLatency = PortAudio.GetDeviceInfo(device).defaultLowInputLatency, + hostApiSpecificStreamInfo = IntPtr.Zero, + }; + + Stream.Callback callback = (IntPtr input, IntPtr _, uint frameCount, + ref StreamCallbackTimeInfo _2, StreamCallbackFlags status, IntPtr _3) => + { + if (status.HasFlag(StreamCallbackFlags.InputOverflow)) + Console.Error.WriteLine("[status] input overflow"); + + if (frameCount != BlockFrames) + { + Console.Error.WriteLine($"[status] unexpected callback frameCount={frameCount} (want {BlockFrames})"); + return StreamCallbackResult.Continue; + } + + var frame = new short[BlockFrames]; + unsafe + { + short* src = (short*)input.ToPointer(); + fixed (short* dst = &frame[0]) + Buffer.MemoryCopy(src, dst, BlockFrames * sizeof(short), BlockFrames * sizeof(short)); + } + if (!queue.TryAdd(frame, 0)) + Console.Error.WriteLine("[status] consumer behind, dropping frame"); + return StreamCallbackResult.Continue; + }; + + using var stream = new Stream( + inParams, null, SampleRate, BlockFrames, StreamFlags.NoFlag, callback, IntPtr.Zero); + stream.Start(); + Console.WriteLine("Listening. Say 'alexa'. Ctrl-C to exit."); + + var sw = Stopwatch.StartNew(); + TimeSpan lastTrigger = TimeSpan.FromSeconds(-CooldownSeconds); + + try + { + while (!cts.IsCancellationRequested) + { + short[] frame = queue.Take(cts.Token); + float score = model.Predict(frame); + var now = sw.Elapsed; + if (score >= Threshold && (now - lastTrigger).TotalSeconds >= CooldownSeconds) + { + Console.WriteLine($"DETECTED alexa score={score:0.000} t={now.TotalSeconds:0.0}s"); + lastTrigger = now; + // Beep is added in Task 5. + } + } + } + catch (OperationCanceledException) { /* Ctrl-C path */ } + + stream.Stop(); + Console.WriteLine("\nbye"); +} +finally +{ + model?.Dispose(); + PortAudio.Terminate(); +} + +static int FindUsbDevice() +{ + int count = PortAudio.DeviceCount; + for (int i = 0; i < count; i++) + { + var info = PortAudio.GetDeviceInfo(i); + if (info.name.ToLowerInvariant().Contains("usb") && info.maxInputChannels >= 1) + return i; + } + Console.Error.WriteLine("USB audio device not found. Devices:"); + for (int i = 0; i < count; i++) + { + var info = PortAudio.GetDeviceInfo(i); + Console.Error.WriteLine( + $" [{i}] {info.name} in={info.maxInputChannels} out={info.maxOutputChannels}"); + } + Environment.Exit(1); + return -1; // unreachable +} + +static class Libc +{ + [DllImport("libc", EntryPoint = "setenv")] + public static extern int setenv(string name, string value, int overwrite); +} +``` + +Key reuses from Test 1's `Program.cs` (`tests/01-record-play-cs/Program.cs`): the `using Stream = PortAudioSharp.Stream;` disambiguation, the `Libc.setenv` P/Invoke, the `FindUsbDevice` body, the `unsafe` callback `Buffer.MemoryCopy` pattern, the dual `Environment.SetEnvironmentVariable` + `Libc.setenv` calls. + +- [ ] **Step 2: Build locally** + +```sh +dotnet build tests/02-wakeword-cs -c Release +``` + +Expected: success. + +- [ ] **Step 3: Deploy + run on the Pi** + +```sh +bin/probe-cs-2 +``` + +Expected on the Pi: + +``` +Using device ('USB ...') +Loading WakewordModel... +Loaded in 0.XXs. +Listening. Say 'alexa'. Ctrl-C to exit. +``` + +Then say "alexa" 3-4 times at normal volume from ~1 m. Expected: one `DETECTED alexa score=0.XXX t=YYs` line per spoken wakeword, fired within ~1 s of saying it. **No beep yet** — that's Task 5. + +What to check before moving on: + +- Detection actually fires. If 0/4 attempts trigger, the port has a bug — re-verify shapes and the `x/10 + 2` mel transform before continuing. +- No sustained `[status] input overflow` or `[status] consumer behind` lines. Occasional ones at startup are OK. +- No ORT GPU warnings (`/sys/class/drm/card0` etc.) — `ORT_LOGGING_LEVEL=3` should silence them. If they appear, the `Libc.setenv` for ORT_LOGGING_LEVEL isn't being honoured by the onnxruntime version pinned; consult Context7 for the correct env var name for that version. +- Ctrl-C from the workstation cleanly exits the probe (prints `bye`). If Ctrl-C just hangs the SSH session, check that `bin/probe-cs-2` uses `ssh -t`. + +- [ ] **Step 4: Commit** + +```sh +git add tests/02-wakeword-cs/Program.cs +git commit -m "Live wakeword loop: PortAudio input + queue + threshold detection" +``` + +--- + +## Task 5: Fire-and-forget beep + full hardware verification + `findings.md` update + +**Goal:** Add the detection beep without blocking input ingestion, then run the full hardware test against all three pass criteria from the spec, and capture findings. + +**Files:** +- Modify: `tests/02-wakeword-cs/Program.cs` (add beep helper + wire it into the detection branch) +- Modify: `findings.md` (append Test 2 outcome section) + +- [ ] **Step 1: Add the beep buffer + dispatch helper to `Program.cs`** + +At the top of `Program.cs`, after the `const double CooldownSeconds = 1.0;` line, add: + +```csharp +const double BeepHz = 880.0; +const double BeepSeconds = 0.2; +``` + +Before the line `Console.WriteLine("Listening. Say 'alexa'. Ctrl-C to exit.");`, add the beep buffer precomputation: + +```csharp + short[] beepBuffer = MakeBeep(BeepHz, BeepSeconds, SampleRate); +``` + +In the detection branch (inside the main loop, after `lastTrigger = now;`), add: + +```csharp + FireAndForgetBeep(device, beepBuffer); +``` + +At the bottom of the file (after `FindUsbDevice` but before `class Libc`), add: + +```csharp +static short[] MakeBeep(double freqHz, double durationS, int sampleRate) +{ + int n = (int)(sampleRate * durationS); + var buf = new short[n]; + for (int i = 0; i < n; i++) + { + double t = i / (double)sampleRate; + double v = 0.3 * Math.Sin(2.0 * Math.PI * freqHz * t); + buf[i] = (short)(v * short.MaxValue); + } + return buf; +} + +static void FireAndForgetBeep(int device, short[] beepBuffer) +{ + int offset = 0; + var done = new ManualResetEventSlim(false); + + var outParams = new StreamParameters + { + device = device, + channelCount = 1, + sampleFormat = SampleFormat.Int16, + suggestedLatency = PortAudio.GetDeviceInfo(device).defaultLowOutputLatency, + hostApiSpecificStreamInfo = IntPtr.Zero, + }; + + Stream.Callback playCb = (IntPtr _, IntPtr output, uint frameCount, + ref StreamCallbackTimeInfo _2, StreamCallbackFlags _3, IntPtr _4) => + { + int remaining = beepBuffer.Length - offset; + int take = (int)Math.Min(frameCount, (uint)remaining); + int silence = (int)frameCount - take; + unsafe + { + short* dst = (short*)output.ToPointer(); + if (take > 0) + { + fixed (short* src = &beepBuffer[offset]) + Buffer.MemoryCopy(src, dst, take * sizeof(short), take * sizeof(short)); + offset += take; + } + for (int i = take; i < frameCount; i++) dst[i] = 0; + } + if (offset >= beepBuffer.Length) + { + done.Set(); + return StreamCallbackResult.Complete; + } + return StreamCallbackResult.Continue; + }; + + var stream = new Stream( + null, outParams, WakewordModel.SampleRate, 1024, StreamFlags.NoFlag, playCb, IntPtr.Zero); + Task.Run(() => + { + done.Wait(); + // Brief drain pause matches Test 1's behaviour; the device buffer needs a moment after Complete. + Thread.Sleep(200); + stream.Stop(); + stream.Dispose(); + done.Dispose(); + }); + stream.Start(); +} +``` + +Note: `MakeBeep` and `FireAndForgetBeep` are `static` local-method-style helpers; with top-level statements they live at file scope alongside `FindUsbDevice` and `Libc`. They reference `Stream` (the `using` alias) and `PortAudio.GetDeviceInfo` — both already in scope. + +The `Task.Run` is captured by the closure but not awaited; the main loop returns immediately after `stream.Start()`. The stream + event handle are kept alive by the closure until the task disposes them. + +- [ ] **Step 2: Build locally** + +```sh +dotnet build tests/02-wakeword-cs -c Release +``` + +Expected: success. + +- [ ] **Step 3: Run on the Pi and verify pass criterion 1 (true-positive rate)** + +```sh +bin/probe-cs-2 +``` + +Wait for `Listening...`. From ~1 m, say "alexa" at normal volume, **10 times**, pausing ≥ 2 seconds between attempts. Count printed `DETECTED ...` lines AND audible beeps. + +**Pass criterion 1:** ≥ 8/10 detections, each within ~1 s of saying the word. + +If you get fewer: the issue is in the inference pipeline. Common causes: +- `x/10 + 2` mel transform wasn't applied (search `WakewordModel.cs` for `/10` — if missing, that's it). +- Off-by-one in the mel window slice (`-EmbeddingWindowMelFrames` vs `-EmbeddingWindowMelFrames-1`). +- Dtype mismatch — passing int16 directly instead of converting to float32. + +- [ ] **Step 4: Verify pass criterion 2 (false-positive rate)** + +While the probe is still running, read a newspaper or book aloud at normal volume from ~1 m for **3 continuous minutes** (anything that isn't "alexa"). Count any spurious `DETECTED ...` lines. + +**Pass criterion 2:** ≤ 1 false positive per minute (i.e., ≤ 3 over the 3-minute test). + +If you get more: the threshold (0.5) was right for the Python probe; if C# is over-triggering, the pipeline is producing systematically higher scores than Python — likely a normalization or transform bug (revisit Step 3 troubleshooting). + +- [ ] **Step 5: Verify pass criterion 3 (CPU budget)** + +In a **second terminal**, while the probe is still running: + +```sh +sshpass -p assistant ssh pi@192.168.50.115 htop +``` + +Find the `Probe` process row. Note the `CPU%` column over ~30 seconds while the probe is doing inference (speak occasionally to keep it busy). + +**Pass criterion 3:** CPU% stays under 50% of one core (the Pi 4 has 4 cores, so htop shows "100%" per core — pass = stays below 50 in that column). + +If higher: confirm `WakewordModel`'s `SessionOptions` set `IntraOpNumThreads = 1` and `InterOpNumThreads = 1` (Task 3 Step 1). If they're set and it's still hot, ONNX Runtime may be ignoring the limit — set the env var via `Libc.setenv("OMP_NUM_THREADS", "1", 1)` at the top of `Program.cs` alongside the others, redeploy, recheck. + +Ctrl-C the probe when done. Quit htop. + +- [ ] **Step 6: Commit the beep code** + +```sh +git add tests/02-wakeword-cs/Program.cs +git commit -m "Beep on detection: fire-and-forget OutputStream, input keeps flowing" +``` + +- [ ] **Step 7: Append Test 2 outcome section to `findings.md`** + +Open `findings.md` and append, after the existing `## C# probe outcome (2026-06-12 — Test 1 ported to C#)` section, a new section in the same shape: + +```markdown +## C# wakeword probe outcome (2026-06-12 — Test 2 ported to C#) + +The Python Test 2 (openwakeword "alexa" listener with beep on detect) was re-implemented in C# / .NET 9, with the openwakeword `Model.predict()` streaming pipeline ported to `Microsoft.ML.OnnxRuntime`. Probe lives at `tests/02-wakeword-cs/`; deploy script `bin/probe-cs-2`. All three pass criteria from the original spec held on hardware. + +### What we proved + +- **Microsoft.ML.OnnxRuntime ** runs on linux-arm64 from a self-contained `dotnet publish`. NuGet [bundles / does not bundle — confirm] `libonnxruntime.so` for the RID. +- The three openwakeword ONNX models (`melspectrogram.onnx`, `embedding_model.onnx`, `alexa.onnx`, vendored from openwakeword 0.6.0) load and chain correctly when fed real audio. +- **Two PortAudio streams on one USB device works**: the input stream stays open while a per-detection fire-and-forget output stream plays the beep. No errors, no input overflow during playback (fixing the bug `findings.md` § A flagged in the Python probe). +- True-positive rate: /10, false-positive rate: /3 min, CPU: % of one core. (Fill in observed values.) + +### Key gotchas + +- (Capture anything that took > 15 minutes to debug. Likely candidates: env-var name for ORT log level on this version; whether the mel transform `x/10 + 2` actually mattered in C#; whether `IntraOpNumThreads = 1` was honoured; whether shape assertions caught anything during development.) + +### Open questions for the main assistant + +- **Custom wakeword model.** Stock "alexa" works as a probe but the shipped assistant needs a custom "hey assistant" (or similar) model. openwakeword has a Colab notebook for synthetic training; budget ~1 hour to train + verify on the same pipeline this probe validates. +- **Speaker-is-mic self-trigger.** Not in probe scope, but the main assistant will need detector gating during TTS playback (findings.md § B), or hardware/software AEC. +- **Inference latency.** If `WakewordModel.Predict` ever exceeds 80 ms, the queue backs up; sample real measurements from this probe before sizing the queue / consumer thread in the main assistant. +``` + +Replace the bracketed placeholders (``, ``, ``, ``, the bundling note, the gotchas list) with the actual values you observed. **Do not commit with placeholders.** + +- [ ] **Step 8: Commit findings** + +```sh +git add findings.md +git commit -m "Findings: C# wakeword probe outcome" +``` + +--- + +## Notes on what is intentionally **not** in this plan + +- **No automated tests.** The spec is explicit; verification is the listening / counting test on real hardware. The silence smoke test in Task 3 is a pipeline sanity check, not a model-correctness test. +- **No numerical parity test vs Python.** If hardware passes, the port is good enough for the probe. If hardware fails, the diagnostic levers in the spec (Section "Diagnostic levers if it fails") cover the likely causes. +- **No detector gating during beep playback.** The 200 ms beep at 880 Hz is far enough from human speech to not self-trigger the "alexa" classifier; if it did, the spec's § B (gate detector during TTS) is the main-assistant solution, not the probe's. +- **No retry / restart on PortAudio init failure.** Probe — crash and surface. +- **No abstraction over PortAudio.** Same one-file Program.cs shape as Test 1. +- **No custom wakeword.** Stock alexa only. Custom is a separate task per `findings.md` § C. +- **No bin/probe-cs-2 → bin/probe-cs unification.** Decided in brainstorm — kept as separate scripts. + +If during implementation you feel the urge to add any of the above, stop and re-read the spec. The point of this probe is to be cheap and disposable.