Files
Assistant/docs/superpowers/plans/2026-06-12-test-2-wakeword-csharp.md
T

970 lines
44 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Test 2 wakeword C# probe — implementation plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Re-implement Test 2 (listen for the "alexa" wakeword via openwakeword's stock model → beep on detection) in C# on the Raspberry Pi, by porting `openwakeword/model.py`'s streaming inference pipeline to C# / `Microsoft.ML.OnnxRuntime`. Pass the same hardware criteria as the Python version.
**Architecture:** Single-folder throwaway probe under `tests/02-wakeword-cs/`. `WakewordModel.cs` chains three vendored ONNX models (mel-spectrogram → Google speech embedding → alexa keyword classifier) with the same streaming buffer geometry as the Python implementation. `Program.cs` drives a PortAudio input stream into a bounded `BlockingCollection`; the main thread pulls frames, runs `WakewordModel.Predict`, and dispatches a fire-and-forget output stream for the detection beep so input ingestion never stalls. `bin/probe-cs-2` builds a self-contained linux-arm64 binary and ships it to the Pi.
**Tech Stack:** .NET 9, C#, PortAudioSharp2 (NuGet), Microsoft.ML.OnnxRuntime (NuGet), bash, sshpass. Target host: Pi at `192.168.50.115`, user `pi`, password `assistant` (see `CLAUDE.md`). The Pi gets no .NET install — the binary is self-contained.
---
## File structure
| Path | Created in task | Purpose |
| --- | --- | --- |
| `tests/02-wakeword-cs/Probe.csproj` | Task 1 | net9.0 console project; linux-arm64 RID; PortAudioSharp2 + Microsoft.ML.OnnxRuntime refs; bundles `models/*.onnx` to output. |
| `tests/02-wakeword-cs/models/melspectrogram.onnx` | Task 1 | Vendored from the Pi's openwakeword install (~14 KB). |
| `tests/02-wakeword-cs/models/embedding_model.onnx` | Task 1 | Vendored from the Pi (~10 MB), Google's speech embedding model. |
| `tests/02-wakeword-cs/models/alexa.onnx` | Task 1 | Vendored from the Pi (~1.5 MB), stock alexa classifier. |
| `tests/02-wakeword-cs/Program.cs` | Task 1 (spike), rewritten in Task 4 & extended in Task 5 | Entry point: env vars → PortAudio init → device pick → input stream → main inference loop → optional beep. |
| `tests/02-wakeword-cs/WakewordModel.cs` | Task 3 | Loads the three sessions; `Predict(short[] frame1280) → float`. |
| `bin/probe-cs-2` | Task 2 | `dotnet publish``scp``ssh -t` flow. |
`tests/02-wakeword/` (Python), `tests/01-record-play-cs/` (Test 1 C#), and `bin/probe-cs` are untouched.
---
## Task 1: Scaffold project + vendor ONNX models + ORT sanity spike on the Pi
**Goal:** Stand up the .NET project with both NuGet refs, vendor the three openwakeword ONNX models into the repo, and prove on the Pi that `Microsoft.ML.OnnxRuntime` loads all three from a self-contained linux-arm64 publish. Capture each model's declared input / output shape (we'll use these as ground-truth assertions in Task 3).
**Files:**
- Create: `tests/02-wakeword-cs/Probe.csproj`
- Create: `tests/02-wakeword-cs/models/melspectrogram.onnx` (vendored)
- Create: `tests/02-wakeword-cs/models/embedding_model.onnx` (vendored)
- Create: `tests/02-wakeword-cs/models/alexa.onnx` (vendored)
- Create: `tests/02-wakeword-cs/Program.cs` (spike — replaced in Task 4)
- [ ] **Step 1: Pin the Microsoft.ML.OnnxRuntime version**
Use Context7 to fetch the current Microsoft.ML.OnnxRuntime NuGet docs and identify the latest stable version (or check `https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime`):
```sh
curl -s https://api.nuget.org/v3-flatcontainer/microsoft.ml.onnxruntime/index.json | jq -r '.versions[]' | grep -v -E '(rc|preview|alpha|beta)' | tail -1
```
Note the exact version (call it `<ORT_VER>`). Pin this in the csproj — do NOT use a floating `*`.
Same drill for PortAudioSharp2 (already pinned at 1.0.6 in Test 1; use that exact version unless a newer non-prerelease has shipped):
```sh
curl -s https://api.nuget.org/v3-flatcontainer/portaudiosharp2/index.json | jq -r '.versions[-1]'
```
- [ ] **Step 2: Vendor the three ONNX files from the Pi**
```sh
mkdir -p tests/02-wakeword-cs/models
sshpass -p assistant scp \
pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/melspectrogram.onnx \
pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/embedding_model.onnx \
pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/alexa.onnx \
tests/02-wakeword-cs/models/
ls -l tests/02-wakeword-cs/models/
sha256sum tests/02-wakeword-cs/models/*.onnx
```
Expected: three files, roughly 14 KB / 10 MB / 1.5 MB. Record the SHA-256 hashes in your commit message in Step 8 — that pins the exact model files used.
If `scp` fails because the Pi's openwakeword venv lives elsewhere, run:
```sh
sshpass -p assistant ssh pi@192.168.50.115 'find ~ -path "*openwakeword/resources/models/*.onnx" 2>/dev/null'
```
and substitute the correct directory.
- [ ] **Step 3: Write `tests/02-wakeword-cs/Probe.csproj`**
```xml
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<OutputType>Exe</OutputType>
<TargetFramework>net9.0</TargetFramework>
<RootNamespace>WakewordProbe</RootNamespace>
<AssemblyName>Probe</AssemblyName>
<Nullable>enable</Nullable>
<ImplicitUsings>enable</ImplicitUsings>
<RuntimeIdentifier>linux-arm64</RuntimeIdentifier>
<SelfContained>true</SelfContained>
<PublishSingleFile>false</PublishSingleFile>
<InvariantGlobalization>true</InvariantGlobalization>
<AllowUnsafeBlocks>true</AllowUnsafeBlocks>
</PropertyGroup>
<ItemGroup>
<PackageReference Include="PortAudioSharp2" Version="1.0.6" />
<PackageReference Include="Microsoft.ML.OnnxRuntime" Version="<ORT_VER>" />
</ItemGroup>
<ItemGroup>
<Content Include="models/*.onnx">
<CopyToOutputDirectory>PreserveNewest</CopyToOutputDirectory>
</Content>
</ItemGroup>
</Project>
```
Replace `<ORT_VER>` with the version from Step 1.
- [ ] **Step 4: Write the spike `tests/02-wakeword-cs/Program.cs`**
This temporary Program.cs loads each model and prints its declared shapes — exactly what Task 3 needs to assert against. It is replaced in Task 4.
```csharp
using Microsoft.ML.OnnxRuntime;
string modelsDir = Path.Combine(AppContext.BaseDirectory, "models");
string[] modelFiles = { "melspectrogram.onnx", "embedding_model.onnx", "alexa.onnx" };
var opts = new SessionOptions
{
IntraOpNumThreads = 1,
InterOpNumThreads = 1,
LogSeverityLevel = OrtLoggingLevel.ORT_LOGGING_LEVEL_ERROR,
};
foreach (var f in modelFiles)
{
string path = Path.Combine(modelsDir, f);
Console.WriteLine($"== {f} ==");
using var session = new InferenceSession(path, opts);
foreach (var kv in session.InputMetadata)
{
var dims = string.Join(",", kv.Value.Dimensions);
Console.WriteLine($" input '{kv.Key}': shape=[{dims}] dtype={kv.Value.ElementType.Name}");
}
foreach (var kv in session.OutputMetadata)
{
var dims = string.Join(",", kv.Value.Dimensions);
Console.WriteLine($" output '{kv.Key}': shape=[{dims}] dtype={kv.Value.ElementType.Name}");
}
}
Console.WriteLine("All three models loaded.");
```
- [ ] **Step 5: Build and publish locally**
```sh
dotnet publish tests/02-wakeword-cs -c Release -r linux-arm64 --self-contained \
-o /tmp/probe-cs-2-out
```
Expected: build succeeds. Confirm the publish dir contains the binary, the three model files, and the bundled ONNX Runtime native:
```sh
ls /tmp/probe-cs-2-out | grep -E 'libonnxruntime|libportaudio|Probe'
ls /tmp/probe-cs-2-out/models/
```
Expected: `Probe`, `libonnxruntime.so` (or `.so.<version>`), `libportaudio.so` (or similar), plus the three `.onnx` files under `models/`. If `libonnxruntime.so` is missing, the `Microsoft.ML.OnnxRuntime` package for your version may have split the native runtime — install the matching native package (`Microsoft.ML.OnnxRuntime.Native.<rid>` if it exists for `<ORT_VER>`) or downgrade `<ORT_VER>` to a version where the linux-arm64 native is included in the main package. Do NOT proceed to Step 6 if the native is missing.
- [ ] **Step 6: Push the spike to the Pi and run it**
```sh
sshpass -p assistant ssh pi@192.168.50.115 'rm -rf ~/probe-cs-2 && mkdir ~/probe-cs-2'
sshpass -p assistant scp -r /tmp/probe-cs-2-out/. pi@192.168.50.115:~/probe-cs-2/
sshpass -p assistant ssh pi@192.168.50.115 'cd ~/probe-cs-2 && chmod +x Probe && ./Probe'
```
Expected output (numbers and names matching the openwakeword Python source from `/tmp/oww-utils.py`):
```
== melspectrogram.onnx ==
input '<name>': shape=[-1,-1] dtype=Single (1D audio, variable length)
output '<name>': shape=[1,-1,1,32] dtype=Single (mel: 32 bins, variable n_frames)
== embedding_model.onnx ==
input 'input_1': shape=[-1,76,32,1] dtype=Single (76 mel frames × 32 bins × 1 channel, variable batch)
output '<name>': shape=[-1,1,1,96] dtype=Single (96-d embedding per batch)
== alexa.onnx ==
input '<name>': shape=[-1,16,96] dtype=Single (16 embeddings × 96 dims, variable batch)
output '<name>': shape=[-1,1] dtype=Single (binary classifier score)
All three models loaded.
```
Negative dims (`-1`) mean "dynamic / batch dimension" — that's fine, we'll pass `1` at inference time.
**Save the actual printed output to your scratchpad** — the input/output names (`'input_1'` etc.) and exact dimension lists become the assertion values in Task 3 Step 4. The dimension numbers (76, 32, 16, 96) should match the Python source; if they don't, stop and investigate before proceeding.
- [ ] **Step 7: Mark `bin/probe-cs-2` placeholder NOT created yet**
Task 2 creates the deploy script. For Task 1 we use the inline scp commands above. (Task 1 cannot depend on Task 2.)
- [ ] **Step 8: Commit**
```sh
git add tests/02-wakeword-cs/Probe.csproj tests/02-wakeword-cs/Program.cs tests/02-wakeword-cs/models/
git commit -m "Scaffold C# wakeword probe: vendored ONNX models + ORT sanity spike
Models sourced from openwakeword==0.6.0 on the Pi.
SHA-256:
melspectrogram.onnx <hash>
embedding_model.onnx <hash>
alexa.onnx <hash>"
```
---
## Task 2: Wrap the deploy flow as `bin/probe-cs-2`
**Goal:** Capture the build / scp / ssh sequence as a single script so subsequent tasks just run `bin/probe-cs-2`. Mirrors `bin/probe-cs` with new paths and adds `ssh -t` so workstation Ctrl-C propagates as SIGINT to the long-running probe.
**Files:**
- Create: `bin/probe-cs-2`
- [ ] **Step 1: Write `bin/probe-cs-2`**
```sh
#!/usr/bin/env bash
set -euo pipefail
PI_HOST="${PI_HOST:-192.168.50.115}"
PI_USER="${PI_USER:-pi}"
PI_PASS="${PI_PASS:-assistant}"
PUBLISH_DIR="${PUBLISH_DIR:-/tmp/probe-cs-2-out}"
repo_root="$(cd "$(dirname "$0")/.." && pwd)"
cd "$repo_root"
echo ">> dotnet publish (linux-arm64, self-contained)"
dotnet publish tests/02-wakeword-cs \
-c Release -r linux-arm64 --self-contained \
-o "$PUBLISH_DIR"
echo ">> scp to $PI_USER@$PI_HOST:~/probe-cs-2/ (wipe-and-replace)"
sshpass -p "$PI_PASS" ssh -o StrictHostKeyChecking=accept-new \
"$PI_USER@$PI_HOST" 'rm -rf ~/probe-cs-2 && mkdir ~/probe-cs-2'
sshpass -p "$PI_PASS" scp -r "$PUBLISH_DIR/." \
"$PI_USER@$PI_HOST:~/probe-cs-2/"
echo ">> ssh + run on Pi (Ctrl-C from this terminal stops the probe)"
sshpass -p "$PI_PASS" ssh -t -o StrictHostKeyChecking=accept-new \
"$PI_USER@$PI_HOST" 'cd ~/probe-cs-2 && chmod +x Probe && ./Probe'
```
The `ssh -t` (force pseudo-tty) makes Ctrl-C from the workstation propagate as SIGINT to the remote `./Probe` process — needed because Task 4's main loop runs until Ctrl-C. Test 1's `bin/probe-cs` didn't need it because the probe self-terminated after 5 s of recording + playback.
- [ ] **Step 2: Mark it executable**
```sh
chmod +x bin/probe-cs-2
```
- [ ] **Step 3: Smoke-test against Task 1's spike**
```sh
bin/probe-cs-2
```
Expected: same shape-dump output as Task 1 Step 6, ending with `All three models loaded.` Ctrl-C is not needed because the spike exits on its own. If it errors, fix the script — do not work around in later tasks.
- [ ] **Step 4: Commit**
```sh
git add bin/probe-cs-2
git commit -m "Add bin/probe-cs-2: build + scp + run the C# wakeword probe on the Pi"
```
---
## Task 3: `WakewordModel.cs` — port openwakeword's streaming pipeline
**Goal:** A `WakewordModel` class that wraps the three ONNX sessions and exposes a single `Predict(short[] frame1280) → float` method. Internal state matches `openwakeword/AudioFeatures._streaming_features()` and `Model.predict()` from the openwakeword 0.6.0 source on the Pi (already fetched to `/tmp/model.py` and `/tmp/oww-utils.py` during plan-writing — re-fetch if needed: `sshpass -p assistant scp pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/{model,utils}.py /tmp/`).
**Files:**
- Create: `tests/02-wakeword-cs/WakewordModel.cs`
- Modify: `tests/02-wakeword-cs/Program.cs` (replace spike with a load-and-print test)
### Pipeline geometry (derived from openwakeword/utils.py `AudioFeatures`)
Per call to `Predict(short[] frame1280)`:
1. **Raw-audio buffer.** Append the 1280 new samples to a ring of recent samples. Keep at least the last `1280 + 480 = 1760` samples (the +480 = `160 * 3` is the mel model's 30 ms warm-up context that openwakeword passes in via the `-n_samples-160*3:` slice in `_streaming_melspectrogram`).
2. **Mel stage.** Slice the last 1760 samples → run `melspectrogram.onnx` with input shape `(1, 1760)` dtype Single → output is `(1, n_frames, 1, 32)`; squeeze trailing dims to get `(n_frames, 32)` where `n_frames` ≈ 8 (`ceil(1760/160 - 3) = ceil(8) = 8`). Apply the openwakeword post-transform `x = x / 10 + 2` element-wise. Append the new mel frames to the mel-frame ring. Trim ring to the last 970 frames (= `10 * 97`, openwakeword's `melspectrogram_max_len`).
3. **Embedding stage.** Slice the last 76 mel frames from the ring → shape `(1, 76, 32, 1)`. Run `embedding_model.onnx` (input name `input_1`) → output shape `(1, 1, 1, 96)`; squeeze to a `float[96]` embedding. Append to the embedding ring. Trim ring to the last 120 entries (openwakeword's `feature_buffer_max_len`).
4. **Classifier stage.** Slice the last 16 embeddings from the ring → shape `(1, 16, 96)`. Run `alexa.onnx` → output shape `(1, 1)`; return `output[0, 0]` as the score.
5. **Initialization gate.** Return `0f` for the first 16 calls — until the embedding ring has 16 real entries, classifier output is meaningless. (openwakeword pre-populates with random-audio embeddings then zeros the first 5 scores; we do the simpler equivalent: skip outright until we have 16 real embeddings.)
- [ ] **Step 1: Document the geometry in `WakewordModel.cs`**
Create `tests/02-wakeword-cs/WakewordModel.cs` and start with constants and the class skeleton:
```csharp
using Microsoft.ML.OnnxRuntime;
using Microsoft.ML.OnnxRuntime.Tensors;
namespace WakewordProbe;
// Streaming wakeword inference port of openwakeword 0.6.0's predict pipeline.
// References (read alongside this file):
// openwakeword/utils.py:AudioFeatures._streaming_features (mel + embedding stages)
// openwakeword/utils.py:AudioFeatures._streaming_melspectrogram
// openwakeword/utils.py:AudioFeatures._get_embeddings
// openwakeword/model.py:Model.predict (classifier stage)
internal sealed class WakewordModel : IDisposable
{
// Audio I/O geometry — per-call input contract.
public const int FrameSamples = 1280; // 80 ms @ 16 kHz; each Predict() call.
public const int SampleRate = 16_000;
// Mel model: input is last (FrameSamples + MelContextSamples) raw samples
// -- the +480 is openwakeword's `-n_samples-160*3:` slice (3 hops of context).
public const int MelContextSamples = 480; // 160 * 3
public const int MelInputSamples = FrameSamples + MelContextSamples; // 1760
public const int MelBins = 32; // melspectrogram model output dim
public const int MelBufferMaxFrames = 970; // 10 * 97 (openwakeword `melspectrogram_max_len`)
// Embedding model: 76-mel-frame window in, 96-d embedding out.
public const int EmbeddingWindowMelFrames = 76;
public const int EmbeddingDim = 96;
public const int EmbeddingBufferMax = 120; // openwakeword `feature_buffer_max_len`
public const string EmbeddingInputName = "input_1"; // openwakeword convention; assert at startup
// Classifier: 16 embeddings in, scalar score out.
public const int ClassifierEmbeddings = 16;
// Skip the first N Predict() calls — buffer fill-up window.
public const int WarmupFrames = ClassifierEmbeddings; // 16 frames ≈ 1.28 s
private readonly InferenceSession _mel;
private readonly InferenceSession _emb;
private readonly InferenceSession _cls;
private readonly string _melInputName; // discovered at startup
private readonly string _clsInputName; // discovered at startup
private readonly short[] _rawRing = new short[MelInputSamples];
private int _rawRingFill = 0; // samples buffered (≤ MelInputSamples)
private readonly List<float[]> _melRing = new(MelBufferMaxFrames); // each entry is a length-32 mel frame
private readonly List<float[]> _embRing = new(EmbeddingBufferMax); // each entry is a length-96 embedding
private int _framesSeen = 0;
public WakewordModel(string melPath, string embeddingPath, string classifierPath)
{
var opts = new SessionOptions
{
IntraOpNumThreads = 1,
InterOpNumThreads = 1,
LogSeverityLevel = OrtLoggingLevel.ORT_LOGGING_LEVEL_ERROR,
};
_mel = new InferenceSession(melPath, opts);
_emb = new InferenceSession(embeddingPath, opts);
_cls = new InferenceSession(classifierPath, opts);
// Discover input names + assert shape geometry.
_melInputName = _mel.InputMetadata.Keys.Single();
AssertEmbeddingShape(_emb); // input_1: [batch, 76, 32, 1], dtype float
_clsInputName = _cls.InputMetadata.Keys.Single();
AssertClassifierShape(_cls); // [batch, 16, 96], dtype float
}
// ... methods follow in Step 2 ...
public void Dispose()
{
_mel.Dispose();
_emb.Dispose();
_cls.Dispose();
}
}
```
- [ ] **Step 2: Implement the shape-assertion helpers**
Append inside the `WakewordModel` class:
```csharp
private static void AssertEmbeddingShape(InferenceSession sess)
{
if (!sess.InputMetadata.TryGetValue(EmbeddingInputName, out var meta))
throw new InvalidOperationException(
$"embedding_model.onnx: expected input named '{EmbeddingInputName}', got [{string.Join(",", sess.InputMetadata.Keys)}]");
var d = meta.Dimensions;
// Expected: [batch, 76, 32, 1] — batch may be -1 (dynamic).
if (d.Length != 4 || d[1] != EmbeddingWindowMelFrames || d[2] != MelBins || d[3] != 1)
throw new InvalidOperationException(
$"embedding_model.onnx: expected input shape [batch,{EmbeddingWindowMelFrames},{MelBins},1], got [{string.Join(",", d)}]");
if (meta.ElementType != typeof(float))
throw new InvalidOperationException($"embedding_model.onnx: expected Single input, got {meta.ElementType.Name}");
}
private static void AssertClassifierShape(InferenceSession sess)
{
var inputName = sess.InputMetadata.Keys.Single();
var meta = sess.InputMetadata[inputName];
var d = meta.Dimensions;
// Expected: [batch, 16, 96] — batch may be -1.
if (d.Length != 3 || d[1] != ClassifierEmbeddings || d[2] != EmbeddingDim)
throw new InvalidOperationException(
$"alexa.onnx: expected input shape [batch,{ClassifierEmbeddings},{EmbeddingDim}], got [{string.Join(",", d)}]");
if (meta.ElementType != typeof(float))
throw new InvalidOperationException($"alexa.onnx: expected Single input, got {meta.ElementType.Name}");
}
```
- [ ] **Step 3: Implement `Predict`**
Append inside the `WakewordModel` class:
```csharp
public float Predict(short[] frame1280)
{
if (frame1280.Length != FrameSamples)
throw new ArgumentException($"Expected {FrameSamples} samples, got {frame1280.Length}");
// 1. Append 1280 new samples to the raw ring (shift older samples down if full).
if (_rawRingFill < MelInputSamples)
{
int copyToFront = Math.Min(MelInputSamples - _rawRingFill, FrameSamples);
Array.Copy(frame1280, 0, _rawRing, _rawRingFill, copyToFront);
_rawRingFill += copyToFront;
if (copyToFront < FrameSamples)
{
// Shouldn't happen on the very first call, but defensive.
int leftover = FrameSamples - copyToFront;
Array.Copy(_rawRing, leftover, _rawRing, 0, MelInputSamples - leftover);
Array.Copy(frame1280, copyToFront, _rawRing, MelInputSamples - leftover, leftover);
}
}
else
{
// Shift older samples left by FrameSamples, then append new at the tail.
Array.Copy(_rawRing, FrameSamples, _rawRing, 0, MelInputSamples - FrameSamples);
Array.Copy(frame1280, 0, _rawRing, MelInputSamples - FrameSamples, FrameSamples);
}
// Skip everything until we have the full mel-context window primed.
if (_rawRingFill < MelInputSamples)
{
_framesSeen++;
return 0f;
}
// 2. Mel stage: feed the full _rawRing as float32 (1, MelInputSamples) into mel model.
var melInputData = new float[MelInputSamples];
for (int i = 0; i < MelInputSamples; i++) melInputData[i] = _rawRing[i]; // int16 → float32, NO normalisation
var melInputTensor = new DenseTensor<float>(melInputData, new[] { 1, MelInputSamples });
using var melResults = _mel.Run(new[] {
NamedOnnxValue.CreateFromTensor(_melInputName, melInputTensor)
});
var melOut = melResults.First().AsTensor<float>(); // shape (1, n_frames, 1, 32)
// Apply openwakeword's `x / 10 + 2` transform and append each new frame to _melRing.
int nFrames = melOut.Dimensions[1];
for (int f = 0; f < nFrames; f++)
{
var bin = new float[MelBins];
for (int b = 0; b < MelBins; b++)
bin[b] = melOut[0, f, 0, b] / 10f + 2f;
_melRing.Add(bin);
}
if (_melRing.Count > MelBufferMaxFrames)
_melRing.RemoveRange(0, _melRing.Count - MelBufferMaxFrames);
// 3. Embedding stage: need ≥ 76 mel frames; slice the last 76 → (1, 76, 32, 1).
if (_melRing.Count < EmbeddingWindowMelFrames)
{
_framesSeen++;
return 0f;
}
var embInputData = new float[EmbeddingWindowMelFrames * MelBins];
int startMel = _melRing.Count - EmbeddingWindowMelFrames;
for (int f = 0; f < EmbeddingWindowMelFrames; f++)
Array.Copy(_melRing[startMel + f], 0, embInputData, f * MelBins, MelBins);
var embInputTensor = new DenseTensor<float>(embInputData, new[] { 1, EmbeddingWindowMelFrames, MelBins, 1 });
using var embResults = _emb.Run(new[] {
NamedOnnxValue.CreateFromTensor(EmbeddingInputName, embInputTensor)
});
var embOut = embResults.First().AsTensor<float>(); // shape (1, 1, 1, 96)
var newEmb = new float[EmbeddingDim];
for (int i = 0; i < EmbeddingDim; i++) newEmb[i] = embOut[0, 0, 0, i];
_embRing.Add(newEmb);
if (_embRing.Count > EmbeddingBufferMax)
_embRing.RemoveRange(0, _embRing.Count - EmbeddingBufferMax);
_framesSeen++;
// 4. Warm-up + classifier stage.
if (_embRing.Count < ClassifierEmbeddings || _framesSeen <= WarmupFrames)
return 0f;
var clsInputData = new float[ClassifierEmbeddings * EmbeddingDim];
int startEmb = _embRing.Count - ClassifierEmbeddings;
for (int i = 0; i < ClassifierEmbeddings; i++)
Array.Copy(_embRing[startEmb + i], 0, clsInputData, i * EmbeddingDim, EmbeddingDim);
var clsInputTensor = new DenseTensor<float>(clsInputData, new[] { 1, ClassifierEmbeddings, EmbeddingDim });
using var clsResults = _cls.Run(new[] {
NamedOnnxValue.CreateFromTensor(_clsInputName, clsInputTensor)
});
var clsOut = clsResults.First().AsTensor<float>(); // shape (1, 1)
return clsOut[0, 0];
}
```
- [ ] **Step 4: Replace `Program.cs` spike with a load-only smoke test**
Replace `tests/02-wakeword-cs/Program.cs` with:
```csharp
using WakewordProbe;
string modelsDir = Path.Combine(AppContext.BaseDirectory, "models");
Console.WriteLine("Loading WakewordModel...");
var t0 = DateTime.UtcNow;
using var model = new WakewordModel(
Path.Combine(modelsDir, "melspectrogram.onnx"),
Path.Combine(modelsDir, "embedding_model.onnx"),
Path.Combine(modelsDir, "alexa.onnx"));
Console.WriteLine($"Loaded in {(DateTime.UtcNow - t0).TotalSeconds:0.00}s.");
// Smoke test: feed 50 frames of silence (1280 zero samples each).
// Expect every score to be 0f (warmup gate + no signal).
var silentFrame = new short[WakewordModel.FrameSamples];
int nonZero = 0;
for (int i = 0; i < 50; i++)
{
float s = model.Predict(silentFrame);
if (s != 0f) nonZero++;
}
Console.WriteLine($"Silence test: {nonZero}/50 non-zero scores (expected: 0 — but small drift is OK).");
Console.WriteLine("Predict pipeline ran without throwing.");
```
The "non-zero on silence" count may be small but not literally 0 — the classifier sees ~zero input but its output is rarely exactly 0f. Anything under ~5 (and well below threshold 0.5) is fine. The point of the smoke test is: did `Predict` run end-to-end without throwing on shape mismatches? If it throws, the assertion bug is in the model load or pipeline geometry.
- [ ] **Step 5: Build locally**
```sh
dotnet build tests/02-wakeword-cs -c Release
```
Expected: success. Common compile errors:
- `NamedOnnxValue` / `DenseTensor` / `InferenceSession` not found → confirm `using Microsoft.ML.OnnxRuntime;` and `using Microsoft.ML.OnnxRuntime.Tensors;` are at top.
- `'InputMetadata' has no Single()` → add `using System.Linq;` (with `ImplicitUsings` enabled this should already be there, but the spike didn't need it).
- API surface drift in `Microsoft.ML.OnnxRuntime` `<ORT_VER>` (the package occasionally renames `NamedOnnxValue.CreateFromTensor`) → consult Context7 `/microsoft/onnxruntime` docs for the pinned version. Do NOT silently switch to a different API; document the version-specific shape if it diverges.
- [ ] **Step 6: Run on the Pi via `bin/probe-cs-2`**
```sh
bin/probe-cs-2
```
Expected output:
```
Loading WakewordModel...
Loaded in 0.XXs.
Silence test: <small>/50 non-zero scores (expected: 0 — but small drift is OK).
Predict pipeline ran without throwing.
```
If the run aborts with an `InvalidOperationException` from one of the shape assertions, the model in the repo disagrees with what this plan expects — STOP. Either the vendored ONNX file is from a newer openwakeword version with a different shape (re-verify the SHA-256 hashes against Step 8 of Task 1), or our shape derivation from `oww-utils.py` was wrong. Investigate by re-running the Task 1 spike and comparing dimensions before adjusting `WakewordModel.cs`.
- [ ] **Step 7: Commit**
```sh
git add tests/02-wakeword-cs/WakewordModel.cs tests/02-wakeword-cs/Program.cs
git commit -m "WakewordModel: port openwakeword streaming pipeline (mel→emb→cls)"
```
---
## Task 4: `Program.cs` — live input stream + main inference loop (detection-only, no beep)
**Goal:** Replace the load-only smoke test with the full input stream + inference loop. Detection lines print to stdout but no beep yet — we want detection working in isolation before adding the second PortAudio stream.
**Files:**
- Modify: `tests/02-wakeword-cs/Program.cs` (full rewrite)
- [ ] **Step 1: Rewrite `tests/02-wakeword-cs/Program.cs`**
```csharp
using System.Collections.Concurrent;
using System.Diagnostics;
using System.Runtime.InteropServices;
using PortAudioSharp;
using WakewordProbe;
using Stream = PortAudioSharp.Stream;
const int SampleRate = WakewordModel.SampleRate;
const int Channels = 1;
const uint BlockFrames = WakewordModel.FrameSamples; // 1280
const float Threshold = 0.5f;
const double CooldownSeconds = 1.0;
// .NET-side mirror for any readers that go through Environment.GetEnvironmentVariable,
// AND libc setenv so PortAudio / onnxruntime (which use getenv()) see the values.
Environment.SetEnvironmentVariable("PA_ALSA_PLUGHW", "1");
Libc.setenv("PA_ALSA_PLUGHW", "1", 1);
Environment.SetEnvironmentVariable("ORT_LOGGING_LEVEL", "3");
Libc.setenv("ORT_LOGGING_LEVEL", "3", 1);
PortAudio.Initialize();
WakewordModel? model = null;
try
{
int device = FindUsbDevice();
Console.WriteLine($"Using device {device} ('{PortAudio.GetDeviceInfo(device).name}')");
Console.WriteLine("Loading WakewordModel...");
var t0 = DateTime.UtcNow;
string modelsDir = Path.Combine(AppContext.BaseDirectory, "models");
model = new WakewordModel(
Path.Combine(modelsDir, "melspectrogram.onnx"),
Path.Combine(modelsDir, "embedding_model.onnx"),
Path.Combine(modelsDir, "alexa.onnx"));
Console.WriteLine($"Loaded in {(DateTime.UtcNow - t0).TotalSeconds:0.00}s.");
var queue = new BlockingCollection<short[]>(boundedCapacity: 16);
using var cts = new CancellationTokenSource();
Console.CancelKeyPress += (_, e) => { e.Cancel = true; cts.Cancel(); };
var inParams = new StreamParameters
{
device = device,
channelCount = Channels,
sampleFormat = SampleFormat.Int16,
suggestedLatency = PortAudio.GetDeviceInfo(device).defaultLowInputLatency,
hostApiSpecificStreamInfo = IntPtr.Zero,
};
Stream.Callback callback = (IntPtr input, IntPtr _, uint frameCount,
ref StreamCallbackTimeInfo _2, StreamCallbackFlags status, IntPtr _3) =>
{
if (status.HasFlag(StreamCallbackFlags.InputOverflow))
Console.Error.WriteLine("[status] input overflow");
if (frameCount != BlockFrames)
{
Console.Error.WriteLine($"[status] unexpected callback frameCount={frameCount} (want {BlockFrames})");
return StreamCallbackResult.Continue;
}
var frame = new short[BlockFrames];
unsafe
{
short* src = (short*)input.ToPointer();
fixed (short* dst = &frame[0])
Buffer.MemoryCopy(src, dst, BlockFrames * sizeof(short), BlockFrames * sizeof(short));
}
if (!queue.TryAdd(frame, 0))
Console.Error.WriteLine("[status] consumer behind, dropping frame");
return StreamCallbackResult.Continue;
};
using var stream = new Stream(
inParams, null, SampleRate, BlockFrames, StreamFlags.NoFlag, callback, IntPtr.Zero);
stream.Start();
Console.WriteLine("Listening. Say 'alexa'. Ctrl-C to exit.");
var sw = Stopwatch.StartNew();
TimeSpan lastTrigger = TimeSpan.FromSeconds(-CooldownSeconds);
try
{
while (!cts.IsCancellationRequested)
{
short[] frame = queue.Take(cts.Token);
float score = model.Predict(frame);
var now = sw.Elapsed;
if (score >= Threshold && (now - lastTrigger).TotalSeconds >= CooldownSeconds)
{
Console.WriteLine($"DETECTED alexa score={score:0.000} t={now.TotalSeconds:0.0}s");
lastTrigger = now;
// Beep is added in Task 5.
}
}
}
catch (OperationCanceledException) { /* Ctrl-C path */ }
stream.Stop();
Console.WriteLine("\nbye");
}
finally
{
model?.Dispose();
PortAudio.Terminate();
}
static int FindUsbDevice()
{
int count = PortAudio.DeviceCount;
for (int i = 0; i < count; i++)
{
var info = PortAudio.GetDeviceInfo(i);
if (info.name.ToLowerInvariant().Contains("usb") && info.maxInputChannels >= 1)
return i;
}
Console.Error.WriteLine("USB audio device not found. Devices:");
for (int i = 0; i < count; i++)
{
var info = PortAudio.GetDeviceInfo(i);
Console.Error.WriteLine(
$" [{i}] {info.name} in={info.maxInputChannels} out={info.maxOutputChannels}");
}
Environment.Exit(1);
return -1; // unreachable
}
static class Libc
{
[DllImport("libc", EntryPoint = "setenv")]
public static extern int setenv(string name, string value, int overwrite);
}
```
Key reuses from Test 1's `Program.cs` (`tests/01-record-play-cs/Program.cs`): the `using Stream = PortAudioSharp.Stream;` disambiguation, the `Libc.setenv` P/Invoke, the `FindUsbDevice` body, the `unsafe` callback `Buffer.MemoryCopy` pattern, the dual `Environment.SetEnvironmentVariable` + `Libc.setenv` calls.
- [ ] **Step 2: Build locally**
```sh
dotnet build tests/02-wakeword-cs -c Release
```
Expected: success.
- [ ] **Step 3: Deploy + run on the Pi**
```sh
bin/probe-cs-2
```
Expected on the Pi:
```
Using device <N> ('USB ...')
Loading WakewordModel...
Loaded in 0.XXs.
Listening. Say 'alexa'. Ctrl-C to exit.
```
Then say "alexa" 3-4 times at normal volume from ~1 m. Expected: one `DETECTED alexa score=0.XXX t=YYs` line per spoken wakeword, fired within ~1 s of saying it. **No beep yet** — that's Task 5.
What to check before moving on:
- Detection actually fires. If 0/4 attempts trigger, the port has a bug — re-verify shapes and the `x/10 + 2` mel transform before continuing.
- No sustained `[status] input overflow` or `[status] consumer behind` lines. Occasional ones at startup are OK.
- No ORT GPU warnings (`/sys/class/drm/card0` etc.) — `ORT_LOGGING_LEVEL=3` should silence them. If they appear, the `Libc.setenv` for ORT_LOGGING_LEVEL isn't being honoured by the onnxruntime version pinned; consult Context7 for the correct env var name for that version.
- Ctrl-C from the workstation cleanly exits the probe (prints `bye`). If Ctrl-C just hangs the SSH session, check that `bin/probe-cs-2` uses `ssh -t`.
- [ ] **Step 4: Commit**
```sh
git add tests/02-wakeword-cs/Program.cs
git commit -m "Live wakeword loop: PortAudio input + queue + threshold detection"
```
---
## Task 5: Fire-and-forget beep + full hardware verification + `findings.md` update
**Goal:** Add the detection beep without blocking input ingestion, then run the full hardware test against all three pass criteria from the spec, and capture findings.
**Files:**
- Modify: `tests/02-wakeword-cs/Program.cs` (add beep helper + wire it into the detection branch)
- Modify: `findings.md` (append Test 2 outcome section)
- [ ] **Step 1: Add the beep buffer + dispatch helper to `Program.cs`**
At the top of `Program.cs`, after the `const double CooldownSeconds = 1.0;` line, add:
```csharp
const double BeepHz = 880.0;
const double BeepSeconds = 0.2;
```
Before the line `Console.WriteLine("Listening. Say 'alexa'. Ctrl-C to exit.");`, add the beep buffer precomputation:
```csharp
short[] beepBuffer = MakeBeep(BeepHz, BeepSeconds, SampleRate);
```
In the detection branch (inside the main loop, after `lastTrigger = now;`), add:
```csharp
FireAndForgetBeep(device, beepBuffer);
```
At the bottom of the file (after `FindUsbDevice` but before `class Libc`), add:
```csharp
static short[] MakeBeep(double freqHz, double durationS, int sampleRate)
{
int n = (int)(sampleRate * durationS);
var buf = new short[n];
for (int i = 0; i < n; i++)
{
double t = i / (double)sampleRate;
double v = 0.3 * Math.Sin(2.0 * Math.PI * freqHz * t);
buf[i] = (short)(v * short.MaxValue);
}
return buf;
}
static void FireAndForgetBeep(int device, short[] beepBuffer)
{
int offset = 0;
var done = new ManualResetEventSlim(false);
var outParams = new StreamParameters
{
device = device,
channelCount = 1,
sampleFormat = SampleFormat.Int16,
suggestedLatency = PortAudio.GetDeviceInfo(device).defaultLowOutputLatency,
hostApiSpecificStreamInfo = IntPtr.Zero,
};
Stream.Callback playCb = (IntPtr _, IntPtr output, uint frameCount,
ref StreamCallbackTimeInfo _2, StreamCallbackFlags _3, IntPtr _4) =>
{
int remaining = beepBuffer.Length - offset;
int take = (int)Math.Min(frameCount, (uint)remaining);
int silence = (int)frameCount - take;
unsafe
{
short* dst = (short*)output.ToPointer();
if (take > 0)
{
fixed (short* src = &beepBuffer[offset])
Buffer.MemoryCopy(src, dst, take * sizeof(short), take * sizeof(short));
offset += take;
}
for (int i = take; i < frameCount; i++) dst[i] = 0;
}
if (offset >= beepBuffer.Length)
{
done.Set();
return StreamCallbackResult.Complete;
}
return StreamCallbackResult.Continue;
};
var stream = new Stream(
null, outParams, WakewordModel.SampleRate, 1024, StreamFlags.NoFlag, playCb, IntPtr.Zero);
Task.Run(() =>
{
done.Wait();
// Brief drain pause matches Test 1's behaviour; the device buffer needs a moment after Complete.
Thread.Sleep(200);
stream.Stop();
stream.Dispose();
done.Dispose();
});
stream.Start();
}
```
Note: `MakeBeep` and `FireAndForgetBeep` are `static` local-method-style helpers; with top-level statements they live at file scope alongside `FindUsbDevice` and `Libc`. They reference `Stream` (the `using` alias) and `PortAudio.GetDeviceInfo` — both already in scope.
The `Task.Run` is captured by the closure but not awaited; the main loop returns immediately after `stream.Start()`. The stream + event handle are kept alive by the closure until the task disposes them.
- [ ] **Step 2: Build locally**
```sh
dotnet build tests/02-wakeword-cs -c Release
```
Expected: success.
- [ ] **Step 3: Run on the Pi and verify pass criterion 1 (true-positive rate)**
```sh
bin/probe-cs-2
```
Wait for `Listening...`. From ~1 m, say "alexa" at normal volume, **10 times**, pausing ≥ 2 seconds between attempts. Count printed `DETECTED ...` lines AND audible beeps.
**Pass criterion 1:** ≥ 8/10 detections, each within ~1 s of saying the word.
If you get fewer: the issue is in the inference pipeline. Common causes:
- `x/10 + 2` mel transform wasn't applied (search `WakewordModel.cs` for `/10` — if missing, that's it).
- Off-by-one in the mel window slice (`-EmbeddingWindowMelFrames` vs `-EmbeddingWindowMelFrames-1`).
- Dtype mismatch — passing int16 directly instead of converting to float32.
- [ ] **Step 4: Verify pass criterion 2 (false-positive rate)**
While the probe is still running, read a newspaper or book aloud at normal volume from ~1 m for **3 continuous minutes** (anything that isn't "alexa"). Count any spurious `DETECTED ...` lines.
**Pass criterion 2:** ≤ 1 false positive per minute (i.e., ≤ 3 over the 3-minute test).
If you get more: the threshold (0.5) was right for the Python probe; if C# is over-triggering, the pipeline is producing systematically higher scores than Python — likely a normalization or transform bug (revisit Step 3 troubleshooting).
- [ ] **Step 5: Verify pass criterion 3 (CPU budget)**
In a **second terminal**, while the probe is still running:
```sh
sshpass -p assistant ssh pi@192.168.50.115 htop
```
Find the `Probe` process row. Note the `CPU%` column over ~30 seconds while the probe is doing inference (speak occasionally to keep it busy).
**Pass criterion 3:** CPU% stays under 50% of one core (the Pi 4 has 4 cores, so htop shows "100%" per core — pass = stays below 50 in that column).
If higher: confirm `WakewordModel`'s `SessionOptions` set `IntraOpNumThreads = 1` and `InterOpNumThreads = 1` (Task 3 Step 1). If they're set and it's still hot, ONNX Runtime may be ignoring the limit — set the env var via `Libc.setenv("OMP_NUM_THREADS", "1", 1)` at the top of `Program.cs` alongside the others, redeploy, recheck.
Ctrl-C the probe when done. Quit htop.
- [ ] **Step 6: Commit the beep code**
```sh
git add tests/02-wakeword-cs/Program.cs
git commit -m "Beep on detection: fire-and-forget OutputStream, input keeps flowing"
```
- [ ] **Step 7: Append Test 2 outcome section to `findings.md`**
Open `findings.md` and append, after the existing `## C# probe outcome (2026-06-12 — Test 1 ported to C#)` section, a new section in the same shape:
```markdown
## C# wakeword probe outcome (2026-06-12 — Test 2 ported to C#)
The Python Test 2 (openwakeword "alexa" listener with beep on detect) was re-implemented in C# / .NET 9, with the openwakeword `Model.predict()` streaming pipeline ported to `Microsoft.ML.OnnxRuntime`. Probe lives at `tests/02-wakeword-cs/`; deploy script `bin/probe-cs-2`. All three pass criteria from the original spec held on hardware.
### What we proved
- **Microsoft.ML.OnnxRuntime <ORT_VER>** runs on linux-arm64 from a self-contained `dotnet publish`. NuGet [bundles / does not bundle — confirm] `libonnxruntime.so` for the RID.
- The three openwakeword ONNX models (`melspectrogram.onnx`, `embedding_model.onnx`, `alexa.onnx`, vendored from openwakeword 0.6.0) load and chain correctly when fed real audio.
- **Two PortAudio streams on one USB device works**: the input stream stays open while a per-detection fire-and-forget output stream plays the beep. No errors, no input overflow during playback (fixing the bug `findings.md` § A flagged in the Python probe).
- True-positive rate: <X>/10, false-positive rate: <Y>/3 min, CPU: <Z>% of one core. (Fill in observed values.)
### Key gotchas
- (Capture anything that took > 15 minutes to debug. Likely candidates: env-var name for ORT log level on this version; whether the mel transform `x/10 + 2` actually mattered in C#; whether `IntraOpNumThreads = 1` was honoured; whether shape assertions caught anything during development.)
### Open questions for the main assistant
- **Custom wakeword model.** Stock "alexa" works as a probe but the shipped assistant needs a custom "hey assistant" (or similar) model. openwakeword has a Colab notebook for synthetic training; budget ~1 hour to train + verify on the same pipeline this probe validates.
- **Speaker-is-mic self-trigger.** Not in probe scope, but the main assistant will need detector gating during TTS playback (findings.md § B), or hardware/software AEC.
- **Inference latency.** If `WakewordModel.Predict` ever exceeds 80 ms, the queue backs up; sample real measurements from this probe before sizing the queue / consumer thread in the main assistant.
```
Replace the bracketed placeholders (`<ORT_VER>`, `<X>`, `<Y>`, `<Z>`, the bundling note, the gotchas list) with the actual values you observed. **Do not commit with placeholders.**
- [ ] **Step 8: Commit findings**
```sh
git add findings.md
git commit -m "Findings: C# wakeword probe outcome"
```
---
## Notes on what is intentionally **not** in this plan
- **No automated tests.** The spec is explicit; verification is the listening / counting test on real hardware. The silence smoke test in Task 3 is a pipeline sanity check, not a model-correctness test.
- **No numerical parity test vs Python.** If hardware passes, the port is good enough for the probe. If hardware fails, the diagnostic levers in the spec (Section "Diagnostic levers if it fails") cover the likely causes.
- **No detector gating during beep playback.** The 200 ms beep at 880 Hz is far enough from human speech to not self-trigger the "alexa" classifier; if it did, the spec's § B (gate detector during TTS) is the main-assistant solution, not the probe's.
- **No retry / restart on PortAudio init failure.** Probe — crash and surface.
- **No abstraction over PortAudio.** Same one-file Program.cs shape as Test 1.
- **No custom wakeword.** Stock alexa only. Custom is a separate task per `findings.md` § C.
- **No bin/probe-cs-2 → bin/probe-cs unification.** Decided in brainstorm — kept as separate scripts.
If during implementation you feel the urge to add any of the above, stop and re-read the spec. The point of this probe is to be cheap and disposable.