Plan: C# probe re-implementing Test 2 (wakeword) on the Pi
This commit is contained in:
@@ -0,0 +1,969 @@
|
|||||||
|
# Test 2 wakeword C# probe — implementation plan
|
||||||
|
|
||||||
|
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||||||
|
|
||||||
|
**Goal:** Re-implement Test 2 (listen for the "alexa" wakeword via openwakeword's stock model → beep on detection) in C# on the Raspberry Pi, by porting `openwakeword/model.py`'s streaming inference pipeline to C# / `Microsoft.ML.OnnxRuntime`. Pass the same hardware criteria as the Python version.
|
||||||
|
|
||||||
|
**Architecture:** Single-folder throwaway probe under `tests/02-wakeword-cs/`. `WakewordModel.cs` chains three vendored ONNX models (mel-spectrogram → Google speech embedding → alexa keyword classifier) with the same streaming buffer geometry as the Python implementation. `Program.cs` drives a PortAudio input stream into a bounded `BlockingCollection`; the main thread pulls frames, runs `WakewordModel.Predict`, and dispatches a fire-and-forget output stream for the detection beep so input ingestion never stalls. `bin/probe-cs-2` builds a self-contained linux-arm64 binary and ships it to the Pi.
|
||||||
|
|
||||||
|
**Tech Stack:** .NET 9, C#, PortAudioSharp2 (NuGet), Microsoft.ML.OnnxRuntime (NuGet), bash, sshpass. Target host: Pi at `192.168.50.115`, user `pi`, password `assistant` (see `CLAUDE.md`). The Pi gets no .NET install — the binary is self-contained.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## File structure
|
||||||
|
|
||||||
|
| Path | Created in task | Purpose |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `tests/02-wakeword-cs/Probe.csproj` | Task 1 | net9.0 console project; linux-arm64 RID; PortAudioSharp2 + Microsoft.ML.OnnxRuntime refs; bundles `models/*.onnx` to output. |
|
||||||
|
| `tests/02-wakeword-cs/models/melspectrogram.onnx` | Task 1 | Vendored from the Pi's openwakeword install (~14 KB). |
|
||||||
|
| `tests/02-wakeword-cs/models/embedding_model.onnx` | Task 1 | Vendored from the Pi (~10 MB), Google's speech embedding model. |
|
||||||
|
| `tests/02-wakeword-cs/models/alexa.onnx` | Task 1 | Vendored from the Pi (~1.5 MB), stock alexa classifier. |
|
||||||
|
| `tests/02-wakeword-cs/Program.cs` | Task 1 (spike), rewritten in Task 4 & extended in Task 5 | Entry point: env vars → PortAudio init → device pick → input stream → main inference loop → optional beep. |
|
||||||
|
| `tests/02-wakeword-cs/WakewordModel.cs` | Task 3 | Loads the three sessions; `Predict(short[] frame1280) → float`. |
|
||||||
|
| `bin/probe-cs-2` | Task 2 | `dotnet publish` → `scp` → `ssh -t` flow. |
|
||||||
|
|
||||||
|
`tests/02-wakeword/` (Python), `tests/01-record-play-cs/` (Test 1 C#), and `bin/probe-cs` are untouched.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 1: Scaffold project + vendor ONNX models + ORT sanity spike on the Pi
|
||||||
|
|
||||||
|
**Goal:** Stand up the .NET project with both NuGet refs, vendor the three openwakeword ONNX models into the repo, and prove on the Pi that `Microsoft.ML.OnnxRuntime` loads all three from a self-contained linux-arm64 publish. Capture each model's declared input / output shape (we'll use these as ground-truth assertions in Task 3).
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `tests/02-wakeword-cs/Probe.csproj`
|
||||||
|
- Create: `tests/02-wakeword-cs/models/melspectrogram.onnx` (vendored)
|
||||||
|
- Create: `tests/02-wakeword-cs/models/embedding_model.onnx` (vendored)
|
||||||
|
- Create: `tests/02-wakeword-cs/models/alexa.onnx` (vendored)
|
||||||
|
- Create: `tests/02-wakeword-cs/Program.cs` (spike — replaced in Task 4)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Pin the Microsoft.ML.OnnxRuntime version**
|
||||||
|
|
||||||
|
Use Context7 to fetch the current Microsoft.ML.OnnxRuntime NuGet docs and identify the latest stable version (or check `https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime`):
|
||||||
|
|
||||||
|
```sh
|
||||||
|
curl -s https://api.nuget.org/v3-flatcontainer/microsoft.ml.onnxruntime/index.json | jq -r '.versions[]' | grep -v -E '(rc|preview|alpha|beta)' | tail -1
|
||||||
|
```
|
||||||
|
|
||||||
|
Note the exact version (call it `<ORT_VER>`). Pin this in the csproj — do NOT use a floating `*`.
|
||||||
|
|
||||||
|
Same drill for PortAudioSharp2 (already pinned at 1.0.6 in Test 1; use that exact version unless a newer non-prerelease has shipped):
|
||||||
|
|
||||||
|
```sh
|
||||||
|
curl -s https://api.nuget.org/v3-flatcontainer/portaudiosharp2/index.json | jq -r '.versions[-1]'
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Vendor the three ONNX files from the Pi**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
mkdir -p tests/02-wakeword-cs/models
|
||||||
|
sshpass -p assistant scp \
|
||||||
|
pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/melspectrogram.onnx \
|
||||||
|
pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/embedding_model.onnx \
|
||||||
|
pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/resources/models/alexa.onnx \
|
||||||
|
tests/02-wakeword-cs/models/
|
||||||
|
ls -l tests/02-wakeword-cs/models/
|
||||||
|
sha256sum tests/02-wakeword-cs/models/*.onnx
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: three files, roughly 14 KB / 10 MB / 1.5 MB. Record the SHA-256 hashes in your commit message in Step 8 — that pins the exact model files used.
|
||||||
|
|
||||||
|
If `scp` fails because the Pi's openwakeword venv lives elsewhere, run:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sshpass -p assistant ssh pi@192.168.50.115 'find ~ -path "*openwakeword/resources/models/*.onnx" 2>/dev/null'
|
||||||
|
```
|
||||||
|
|
||||||
|
and substitute the correct directory.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Write `tests/02-wakeword-cs/Probe.csproj`**
|
||||||
|
|
||||||
|
```xml
|
||||||
|
<Project Sdk="Microsoft.NET.Sdk">
|
||||||
|
<PropertyGroup>
|
||||||
|
<OutputType>Exe</OutputType>
|
||||||
|
<TargetFramework>net9.0</TargetFramework>
|
||||||
|
<RootNamespace>WakewordProbe</RootNamespace>
|
||||||
|
<AssemblyName>Probe</AssemblyName>
|
||||||
|
<Nullable>enable</Nullable>
|
||||||
|
<ImplicitUsings>enable</ImplicitUsings>
|
||||||
|
<RuntimeIdentifier>linux-arm64</RuntimeIdentifier>
|
||||||
|
<SelfContained>true</SelfContained>
|
||||||
|
<PublishSingleFile>false</PublishSingleFile>
|
||||||
|
<InvariantGlobalization>true</InvariantGlobalization>
|
||||||
|
<AllowUnsafeBlocks>true</AllowUnsafeBlocks>
|
||||||
|
</PropertyGroup>
|
||||||
|
<ItemGroup>
|
||||||
|
<PackageReference Include="PortAudioSharp2" Version="1.0.6" />
|
||||||
|
<PackageReference Include="Microsoft.ML.OnnxRuntime" Version="<ORT_VER>" />
|
||||||
|
</ItemGroup>
|
||||||
|
<ItemGroup>
|
||||||
|
<Content Include="models/*.onnx">
|
||||||
|
<CopyToOutputDirectory>PreserveNewest</CopyToOutputDirectory>
|
||||||
|
</Content>
|
||||||
|
</ItemGroup>
|
||||||
|
</Project>
|
||||||
|
```
|
||||||
|
|
||||||
|
Replace `<ORT_VER>` with the version from Step 1.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Write the spike `tests/02-wakeword-cs/Program.cs`**
|
||||||
|
|
||||||
|
This temporary Program.cs loads each model and prints its declared shapes — exactly what Task 3 needs to assert against. It is replaced in Task 4.
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using Microsoft.ML.OnnxRuntime;
|
||||||
|
|
||||||
|
string modelsDir = Path.Combine(AppContext.BaseDirectory, "models");
|
||||||
|
string[] modelFiles = { "melspectrogram.onnx", "embedding_model.onnx", "alexa.onnx" };
|
||||||
|
|
||||||
|
var opts = new SessionOptions
|
||||||
|
{
|
||||||
|
IntraOpNumThreads = 1,
|
||||||
|
InterOpNumThreads = 1,
|
||||||
|
LogSeverityLevel = OrtLoggingLevel.ORT_LOGGING_LEVEL_ERROR,
|
||||||
|
};
|
||||||
|
|
||||||
|
foreach (var f in modelFiles)
|
||||||
|
{
|
||||||
|
string path = Path.Combine(modelsDir, f);
|
||||||
|
Console.WriteLine($"== {f} ==");
|
||||||
|
using var session = new InferenceSession(path, opts);
|
||||||
|
foreach (var kv in session.InputMetadata)
|
||||||
|
{
|
||||||
|
var dims = string.Join(",", kv.Value.Dimensions);
|
||||||
|
Console.WriteLine($" input '{kv.Key}': shape=[{dims}] dtype={kv.Value.ElementType.Name}");
|
||||||
|
}
|
||||||
|
foreach (var kv in session.OutputMetadata)
|
||||||
|
{
|
||||||
|
var dims = string.Join(",", kv.Value.Dimensions);
|
||||||
|
Console.WriteLine($" output '{kv.Key}': shape=[{dims}] dtype={kv.Value.ElementType.Name}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
Console.WriteLine("All three models loaded.");
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 5: Build and publish locally**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
dotnet publish tests/02-wakeword-cs -c Release -r linux-arm64 --self-contained \
|
||||||
|
-o /tmp/probe-cs-2-out
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: build succeeds. Confirm the publish dir contains the binary, the three model files, and the bundled ONNX Runtime native:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
ls /tmp/probe-cs-2-out | grep -E 'libonnxruntime|libportaudio|Probe'
|
||||||
|
ls /tmp/probe-cs-2-out/models/
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: `Probe`, `libonnxruntime.so` (or `.so.<version>`), `libportaudio.so` (or similar), plus the three `.onnx` files under `models/`. If `libonnxruntime.so` is missing, the `Microsoft.ML.OnnxRuntime` package for your version may have split the native runtime — install the matching native package (`Microsoft.ML.OnnxRuntime.Native.<rid>` if it exists for `<ORT_VER>`) or downgrade `<ORT_VER>` to a version where the linux-arm64 native is included in the main package. Do NOT proceed to Step 6 if the native is missing.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Push the spike to the Pi and run it**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sshpass -p assistant ssh pi@192.168.50.115 'rm -rf ~/probe-cs-2 && mkdir ~/probe-cs-2'
|
||||||
|
sshpass -p assistant scp -r /tmp/probe-cs-2-out/. pi@192.168.50.115:~/probe-cs-2/
|
||||||
|
sshpass -p assistant ssh pi@192.168.50.115 'cd ~/probe-cs-2 && chmod +x Probe && ./Probe'
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected output (numbers and names matching the openwakeword Python source from `/tmp/oww-utils.py`):
|
||||||
|
|
||||||
|
```
|
||||||
|
== melspectrogram.onnx ==
|
||||||
|
input '<name>': shape=[-1,-1] dtype=Single (1D audio, variable length)
|
||||||
|
output '<name>': shape=[1,-1,1,32] dtype=Single (mel: 32 bins, variable n_frames)
|
||||||
|
== embedding_model.onnx ==
|
||||||
|
input 'input_1': shape=[-1,76,32,1] dtype=Single (76 mel frames × 32 bins × 1 channel, variable batch)
|
||||||
|
output '<name>': shape=[-1,1,1,96] dtype=Single (96-d embedding per batch)
|
||||||
|
== alexa.onnx ==
|
||||||
|
input '<name>': shape=[-1,16,96] dtype=Single (16 embeddings × 96 dims, variable batch)
|
||||||
|
output '<name>': shape=[-1,1] dtype=Single (binary classifier score)
|
||||||
|
All three models loaded.
|
||||||
|
```
|
||||||
|
|
||||||
|
Negative dims (`-1`) mean "dynamic / batch dimension" — that's fine, we'll pass `1` at inference time.
|
||||||
|
|
||||||
|
**Save the actual printed output to your scratchpad** — the input/output names (`'input_1'` etc.) and exact dimension lists become the assertion values in Task 3 Step 4. The dimension numbers (76, 32, 16, 96) should match the Python source; if they don't, stop and investigate before proceeding.
|
||||||
|
|
||||||
|
- [ ] **Step 7: Mark `bin/probe-cs-2` placeholder NOT created yet**
|
||||||
|
|
||||||
|
Task 2 creates the deploy script. For Task 1 we use the inline scp commands above. (Task 1 cannot depend on Task 2.)
|
||||||
|
|
||||||
|
- [ ] **Step 8: Commit**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git add tests/02-wakeword-cs/Probe.csproj tests/02-wakeword-cs/Program.cs tests/02-wakeword-cs/models/
|
||||||
|
git commit -m "Scaffold C# wakeword probe: vendored ONNX models + ORT sanity spike
|
||||||
|
|
||||||
|
Models sourced from openwakeword==0.6.0 on the Pi.
|
||||||
|
SHA-256:
|
||||||
|
melspectrogram.onnx <hash>
|
||||||
|
embedding_model.onnx <hash>
|
||||||
|
alexa.onnx <hash>"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 2: Wrap the deploy flow as `bin/probe-cs-2`
|
||||||
|
|
||||||
|
**Goal:** Capture the build / scp / ssh sequence as a single script so subsequent tasks just run `bin/probe-cs-2`. Mirrors `bin/probe-cs` with new paths and adds `ssh -t` so workstation Ctrl-C propagates as SIGINT to the long-running probe.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `bin/probe-cs-2`
|
||||||
|
|
||||||
|
- [ ] **Step 1: Write `bin/probe-cs-2`**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
PI_HOST="${PI_HOST:-192.168.50.115}"
|
||||||
|
PI_USER="${PI_USER:-pi}"
|
||||||
|
PI_PASS="${PI_PASS:-assistant}"
|
||||||
|
PUBLISH_DIR="${PUBLISH_DIR:-/tmp/probe-cs-2-out}"
|
||||||
|
|
||||||
|
repo_root="$(cd "$(dirname "$0")/.." && pwd)"
|
||||||
|
cd "$repo_root"
|
||||||
|
|
||||||
|
echo ">> dotnet publish (linux-arm64, self-contained)"
|
||||||
|
dotnet publish tests/02-wakeword-cs \
|
||||||
|
-c Release -r linux-arm64 --self-contained \
|
||||||
|
-o "$PUBLISH_DIR"
|
||||||
|
|
||||||
|
echo ">> scp to $PI_USER@$PI_HOST:~/probe-cs-2/ (wipe-and-replace)"
|
||||||
|
sshpass -p "$PI_PASS" ssh -o StrictHostKeyChecking=accept-new \
|
||||||
|
"$PI_USER@$PI_HOST" 'rm -rf ~/probe-cs-2 && mkdir ~/probe-cs-2'
|
||||||
|
sshpass -p "$PI_PASS" scp -r "$PUBLISH_DIR/." \
|
||||||
|
"$PI_USER@$PI_HOST:~/probe-cs-2/"
|
||||||
|
|
||||||
|
echo ">> ssh + run on Pi (Ctrl-C from this terminal stops the probe)"
|
||||||
|
sshpass -p "$PI_PASS" ssh -t -o StrictHostKeyChecking=accept-new \
|
||||||
|
"$PI_USER@$PI_HOST" 'cd ~/probe-cs-2 && chmod +x Probe && ./Probe'
|
||||||
|
```
|
||||||
|
|
||||||
|
The `ssh -t` (force pseudo-tty) makes Ctrl-C from the workstation propagate as SIGINT to the remote `./Probe` process — needed because Task 4's main loop runs until Ctrl-C. Test 1's `bin/probe-cs` didn't need it because the probe self-terminated after 5 s of recording + playback.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Mark it executable**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
chmod +x bin/probe-cs-2
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 3: Smoke-test against Task 1's spike**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
bin/probe-cs-2
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: same shape-dump output as Task 1 Step 6, ending with `All three models loaded.` Ctrl-C is not needed because the spike exits on its own. If it errors, fix the script — do not work around in later tasks.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Commit**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git add bin/probe-cs-2
|
||||||
|
git commit -m "Add bin/probe-cs-2: build + scp + run the C# wakeword probe on the Pi"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 3: `WakewordModel.cs` — port openwakeword's streaming pipeline
|
||||||
|
|
||||||
|
**Goal:** A `WakewordModel` class that wraps the three ONNX sessions and exposes a single `Predict(short[] frame1280) → float` method. Internal state matches `openwakeword/AudioFeatures._streaming_features()` and `Model.predict()` from the openwakeword 0.6.0 source on the Pi (already fetched to `/tmp/model.py` and `/tmp/oww-utils.py` during plan-writing — re-fetch if needed: `sshpass -p assistant scp pi@192.168.50.115:assistant/.venv/lib/python3.13/site-packages/openwakeword/{model,utils}.py /tmp/`).
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `tests/02-wakeword-cs/WakewordModel.cs`
|
||||||
|
- Modify: `tests/02-wakeword-cs/Program.cs` (replace spike with a load-and-print test)
|
||||||
|
|
||||||
|
### Pipeline geometry (derived from openwakeword/utils.py `AudioFeatures`)
|
||||||
|
|
||||||
|
Per call to `Predict(short[] frame1280)`:
|
||||||
|
|
||||||
|
1. **Raw-audio buffer.** Append the 1280 new samples to a ring of recent samples. Keep at least the last `1280 + 480 = 1760` samples (the +480 = `160 * 3` is the mel model's 30 ms warm-up context that openwakeword passes in via the `-n_samples-160*3:` slice in `_streaming_melspectrogram`).
|
||||||
|
2. **Mel stage.** Slice the last 1760 samples → run `melspectrogram.onnx` with input shape `(1, 1760)` dtype Single → output is `(1, n_frames, 1, 32)`; squeeze trailing dims to get `(n_frames, 32)` where `n_frames` ≈ 8 (`ceil(1760/160 - 3) = ceil(8) = 8`). Apply the openwakeword post-transform `x = x / 10 + 2` element-wise. Append the new mel frames to the mel-frame ring. Trim ring to the last 970 frames (= `10 * 97`, openwakeword's `melspectrogram_max_len`).
|
||||||
|
3. **Embedding stage.** Slice the last 76 mel frames from the ring → shape `(1, 76, 32, 1)`. Run `embedding_model.onnx` (input name `input_1`) → output shape `(1, 1, 1, 96)`; squeeze to a `float[96]` embedding. Append to the embedding ring. Trim ring to the last 120 entries (openwakeword's `feature_buffer_max_len`).
|
||||||
|
4. **Classifier stage.** Slice the last 16 embeddings from the ring → shape `(1, 16, 96)`. Run `alexa.onnx` → output shape `(1, 1)`; return `output[0, 0]` as the score.
|
||||||
|
5. **Initialization gate.** Return `0f` for the first 16 calls — until the embedding ring has 16 real entries, classifier output is meaningless. (openwakeword pre-populates with random-audio embeddings then zeros the first 5 scores; we do the simpler equivalent: skip outright until we have 16 real embeddings.)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Document the geometry in `WakewordModel.cs`**
|
||||||
|
|
||||||
|
Create `tests/02-wakeword-cs/WakewordModel.cs` and start with constants and the class skeleton:
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using Microsoft.ML.OnnxRuntime;
|
||||||
|
using Microsoft.ML.OnnxRuntime.Tensors;
|
||||||
|
|
||||||
|
namespace WakewordProbe;
|
||||||
|
|
||||||
|
// Streaming wakeword inference port of openwakeword 0.6.0's predict pipeline.
|
||||||
|
// References (read alongside this file):
|
||||||
|
// openwakeword/utils.py:AudioFeatures._streaming_features (mel + embedding stages)
|
||||||
|
// openwakeword/utils.py:AudioFeatures._streaming_melspectrogram
|
||||||
|
// openwakeword/utils.py:AudioFeatures._get_embeddings
|
||||||
|
// openwakeword/model.py:Model.predict (classifier stage)
|
||||||
|
internal sealed class WakewordModel : IDisposable
|
||||||
|
{
|
||||||
|
// Audio I/O geometry — per-call input contract.
|
||||||
|
public const int FrameSamples = 1280; // 80 ms @ 16 kHz; each Predict() call.
|
||||||
|
public const int SampleRate = 16_000;
|
||||||
|
|
||||||
|
// Mel model: input is last (FrameSamples + MelContextSamples) raw samples
|
||||||
|
// -- the +480 is openwakeword's `-n_samples-160*3:` slice (3 hops of context).
|
||||||
|
public const int MelContextSamples = 480; // 160 * 3
|
||||||
|
public const int MelInputSamples = FrameSamples + MelContextSamples; // 1760
|
||||||
|
public const int MelBins = 32; // melspectrogram model output dim
|
||||||
|
public const int MelBufferMaxFrames = 970; // 10 * 97 (openwakeword `melspectrogram_max_len`)
|
||||||
|
|
||||||
|
// Embedding model: 76-mel-frame window in, 96-d embedding out.
|
||||||
|
public const int EmbeddingWindowMelFrames = 76;
|
||||||
|
public const int EmbeddingDim = 96;
|
||||||
|
public const int EmbeddingBufferMax = 120; // openwakeword `feature_buffer_max_len`
|
||||||
|
public const string EmbeddingInputName = "input_1"; // openwakeword convention; assert at startup
|
||||||
|
|
||||||
|
// Classifier: 16 embeddings in, scalar score out.
|
||||||
|
public const int ClassifierEmbeddings = 16;
|
||||||
|
|
||||||
|
// Skip the first N Predict() calls — buffer fill-up window.
|
||||||
|
public const int WarmupFrames = ClassifierEmbeddings; // 16 frames ≈ 1.28 s
|
||||||
|
|
||||||
|
private readonly InferenceSession _mel;
|
||||||
|
private readonly InferenceSession _emb;
|
||||||
|
private readonly InferenceSession _cls;
|
||||||
|
|
||||||
|
private readonly string _melInputName; // discovered at startup
|
||||||
|
private readonly string _clsInputName; // discovered at startup
|
||||||
|
|
||||||
|
private readonly short[] _rawRing = new short[MelInputSamples];
|
||||||
|
private int _rawRingFill = 0; // samples buffered (≤ MelInputSamples)
|
||||||
|
|
||||||
|
private readonly List<float[]> _melRing = new(MelBufferMaxFrames); // each entry is a length-32 mel frame
|
||||||
|
private readonly List<float[]> _embRing = new(EmbeddingBufferMax); // each entry is a length-96 embedding
|
||||||
|
|
||||||
|
private int _framesSeen = 0;
|
||||||
|
|
||||||
|
public WakewordModel(string melPath, string embeddingPath, string classifierPath)
|
||||||
|
{
|
||||||
|
var opts = new SessionOptions
|
||||||
|
{
|
||||||
|
IntraOpNumThreads = 1,
|
||||||
|
InterOpNumThreads = 1,
|
||||||
|
LogSeverityLevel = OrtLoggingLevel.ORT_LOGGING_LEVEL_ERROR,
|
||||||
|
};
|
||||||
|
|
||||||
|
_mel = new InferenceSession(melPath, opts);
|
||||||
|
_emb = new InferenceSession(embeddingPath, opts);
|
||||||
|
_cls = new InferenceSession(classifierPath, opts);
|
||||||
|
|
||||||
|
// Discover input names + assert shape geometry.
|
||||||
|
_melInputName = _mel.InputMetadata.Keys.Single();
|
||||||
|
|
||||||
|
AssertEmbeddingShape(_emb); // input_1: [batch, 76, 32, 1], dtype float
|
||||||
|
_clsInputName = _cls.InputMetadata.Keys.Single();
|
||||||
|
AssertClassifierShape(_cls); // [batch, 16, 96], dtype float
|
||||||
|
}
|
||||||
|
|
||||||
|
// ... methods follow in Step 2 ...
|
||||||
|
|
||||||
|
public void Dispose()
|
||||||
|
{
|
||||||
|
_mel.Dispose();
|
||||||
|
_emb.Dispose();
|
||||||
|
_cls.Dispose();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 2: Implement the shape-assertion helpers**
|
||||||
|
|
||||||
|
Append inside the `WakewordModel` class:
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
private static void AssertEmbeddingShape(InferenceSession sess)
|
||||||
|
{
|
||||||
|
if (!sess.InputMetadata.TryGetValue(EmbeddingInputName, out var meta))
|
||||||
|
throw new InvalidOperationException(
|
||||||
|
$"embedding_model.onnx: expected input named '{EmbeddingInputName}', got [{string.Join(",", sess.InputMetadata.Keys)}]");
|
||||||
|
var d = meta.Dimensions;
|
||||||
|
// Expected: [batch, 76, 32, 1] — batch may be -1 (dynamic).
|
||||||
|
if (d.Length != 4 || d[1] != EmbeddingWindowMelFrames || d[2] != MelBins || d[3] != 1)
|
||||||
|
throw new InvalidOperationException(
|
||||||
|
$"embedding_model.onnx: expected input shape [batch,{EmbeddingWindowMelFrames},{MelBins},1], got [{string.Join(",", d)}]");
|
||||||
|
if (meta.ElementType != typeof(float))
|
||||||
|
throw new InvalidOperationException($"embedding_model.onnx: expected Single input, got {meta.ElementType.Name}");
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void AssertClassifierShape(InferenceSession sess)
|
||||||
|
{
|
||||||
|
var inputName = sess.InputMetadata.Keys.Single();
|
||||||
|
var meta = sess.InputMetadata[inputName];
|
||||||
|
var d = meta.Dimensions;
|
||||||
|
// Expected: [batch, 16, 96] — batch may be -1.
|
||||||
|
if (d.Length != 3 || d[1] != ClassifierEmbeddings || d[2] != EmbeddingDim)
|
||||||
|
throw new InvalidOperationException(
|
||||||
|
$"alexa.onnx: expected input shape [batch,{ClassifierEmbeddings},{EmbeddingDim}], got [{string.Join(",", d)}]");
|
||||||
|
if (meta.ElementType != typeof(float))
|
||||||
|
throw new InvalidOperationException($"alexa.onnx: expected Single input, got {meta.ElementType.Name}");
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 3: Implement `Predict`**
|
||||||
|
|
||||||
|
Append inside the `WakewordModel` class:
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
public float Predict(short[] frame1280)
|
||||||
|
{
|
||||||
|
if (frame1280.Length != FrameSamples)
|
||||||
|
throw new ArgumentException($"Expected {FrameSamples} samples, got {frame1280.Length}");
|
||||||
|
|
||||||
|
// 1. Append 1280 new samples to the raw ring (shift older samples down if full).
|
||||||
|
if (_rawRingFill < MelInputSamples)
|
||||||
|
{
|
||||||
|
int copyToFront = Math.Min(MelInputSamples - _rawRingFill, FrameSamples);
|
||||||
|
Array.Copy(frame1280, 0, _rawRing, _rawRingFill, copyToFront);
|
||||||
|
_rawRingFill += copyToFront;
|
||||||
|
if (copyToFront < FrameSamples)
|
||||||
|
{
|
||||||
|
// Shouldn't happen on the very first call, but defensive.
|
||||||
|
int leftover = FrameSamples - copyToFront;
|
||||||
|
Array.Copy(_rawRing, leftover, _rawRing, 0, MelInputSamples - leftover);
|
||||||
|
Array.Copy(frame1280, copyToFront, _rawRing, MelInputSamples - leftover, leftover);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
else
|
||||||
|
{
|
||||||
|
// Shift older samples left by FrameSamples, then append new at the tail.
|
||||||
|
Array.Copy(_rawRing, FrameSamples, _rawRing, 0, MelInputSamples - FrameSamples);
|
||||||
|
Array.Copy(frame1280, 0, _rawRing, MelInputSamples - FrameSamples, FrameSamples);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Skip everything until we have the full mel-context window primed.
|
||||||
|
if (_rawRingFill < MelInputSamples)
|
||||||
|
{
|
||||||
|
_framesSeen++;
|
||||||
|
return 0f;
|
||||||
|
}
|
||||||
|
|
||||||
|
// 2. Mel stage: feed the full _rawRing as float32 (1, MelInputSamples) into mel model.
|
||||||
|
var melInputData = new float[MelInputSamples];
|
||||||
|
for (int i = 0; i < MelInputSamples; i++) melInputData[i] = _rawRing[i]; // int16 → float32, NO normalisation
|
||||||
|
var melInputTensor = new DenseTensor<float>(melInputData, new[] { 1, MelInputSamples });
|
||||||
|
using var melResults = _mel.Run(new[] {
|
||||||
|
NamedOnnxValue.CreateFromTensor(_melInputName, melInputTensor)
|
||||||
|
});
|
||||||
|
var melOut = melResults.First().AsTensor<float>(); // shape (1, n_frames, 1, 32)
|
||||||
|
|
||||||
|
// Apply openwakeword's `x / 10 + 2` transform and append each new frame to _melRing.
|
||||||
|
int nFrames = melOut.Dimensions[1];
|
||||||
|
for (int f = 0; f < nFrames; f++)
|
||||||
|
{
|
||||||
|
var bin = new float[MelBins];
|
||||||
|
for (int b = 0; b < MelBins; b++)
|
||||||
|
bin[b] = melOut[0, f, 0, b] / 10f + 2f;
|
||||||
|
_melRing.Add(bin);
|
||||||
|
}
|
||||||
|
if (_melRing.Count > MelBufferMaxFrames)
|
||||||
|
_melRing.RemoveRange(0, _melRing.Count - MelBufferMaxFrames);
|
||||||
|
|
||||||
|
// 3. Embedding stage: need ≥ 76 mel frames; slice the last 76 → (1, 76, 32, 1).
|
||||||
|
if (_melRing.Count < EmbeddingWindowMelFrames)
|
||||||
|
{
|
||||||
|
_framesSeen++;
|
||||||
|
return 0f;
|
||||||
|
}
|
||||||
|
var embInputData = new float[EmbeddingWindowMelFrames * MelBins];
|
||||||
|
int startMel = _melRing.Count - EmbeddingWindowMelFrames;
|
||||||
|
for (int f = 0; f < EmbeddingWindowMelFrames; f++)
|
||||||
|
Array.Copy(_melRing[startMel + f], 0, embInputData, f * MelBins, MelBins);
|
||||||
|
var embInputTensor = new DenseTensor<float>(embInputData, new[] { 1, EmbeddingWindowMelFrames, MelBins, 1 });
|
||||||
|
using var embResults = _emb.Run(new[] {
|
||||||
|
NamedOnnxValue.CreateFromTensor(EmbeddingInputName, embInputTensor)
|
||||||
|
});
|
||||||
|
var embOut = embResults.First().AsTensor<float>(); // shape (1, 1, 1, 96)
|
||||||
|
var newEmb = new float[EmbeddingDim];
|
||||||
|
for (int i = 0; i < EmbeddingDim; i++) newEmb[i] = embOut[0, 0, 0, i];
|
||||||
|
_embRing.Add(newEmb);
|
||||||
|
if (_embRing.Count > EmbeddingBufferMax)
|
||||||
|
_embRing.RemoveRange(0, _embRing.Count - EmbeddingBufferMax);
|
||||||
|
|
||||||
|
_framesSeen++;
|
||||||
|
|
||||||
|
// 4. Warm-up + classifier stage.
|
||||||
|
if (_embRing.Count < ClassifierEmbeddings || _framesSeen <= WarmupFrames)
|
||||||
|
return 0f;
|
||||||
|
|
||||||
|
var clsInputData = new float[ClassifierEmbeddings * EmbeddingDim];
|
||||||
|
int startEmb = _embRing.Count - ClassifierEmbeddings;
|
||||||
|
for (int i = 0; i < ClassifierEmbeddings; i++)
|
||||||
|
Array.Copy(_embRing[startEmb + i], 0, clsInputData, i * EmbeddingDim, EmbeddingDim);
|
||||||
|
var clsInputTensor = new DenseTensor<float>(clsInputData, new[] { 1, ClassifierEmbeddings, EmbeddingDim });
|
||||||
|
using var clsResults = _cls.Run(new[] {
|
||||||
|
NamedOnnxValue.CreateFromTensor(_clsInputName, clsInputTensor)
|
||||||
|
});
|
||||||
|
var clsOut = clsResults.First().AsTensor<float>(); // shape (1, 1)
|
||||||
|
return clsOut[0, 0];
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 4: Replace `Program.cs` spike with a load-only smoke test**
|
||||||
|
|
||||||
|
Replace `tests/02-wakeword-cs/Program.cs` with:
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using WakewordProbe;
|
||||||
|
|
||||||
|
string modelsDir = Path.Combine(AppContext.BaseDirectory, "models");
|
||||||
|
Console.WriteLine("Loading WakewordModel...");
|
||||||
|
var t0 = DateTime.UtcNow;
|
||||||
|
using var model = new WakewordModel(
|
||||||
|
Path.Combine(modelsDir, "melspectrogram.onnx"),
|
||||||
|
Path.Combine(modelsDir, "embedding_model.onnx"),
|
||||||
|
Path.Combine(modelsDir, "alexa.onnx"));
|
||||||
|
Console.WriteLine($"Loaded in {(DateTime.UtcNow - t0).TotalSeconds:0.00}s.");
|
||||||
|
|
||||||
|
// Smoke test: feed 50 frames of silence (1280 zero samples each).
|
||||||
|
// Expect every score to be 0f (warmup gate + no signal).
|
||||||
|
var silentFrame = new short[WakewordModel.FrameSamples];
|
||||||
|
int nonZero = 0;
|
||||||
|
for (int i = 0; i < 50; i++)
|
||||||
|
{
|
||||||
|
float s = model.Predict(silentFrame);
|
||||||
|
if (s != 0f) nonZero++;
|
||||||
|
}
|
||||||
|
Console.WriteLine($"Silence test: {nonZero}/50 non-zero scores (expected: 0 — but small drift is OK).");
|
||||||
|
Console.WriteLine("Predict pipeline ran without throwing.");
|
||||||
|
```
|
||||||
|
|
||||||
|
The "non-zero on silence" count may be small but not literally 0 — the classifier sees ~zero input but its output is rarely exactly 0f. Anything under ~5 (and well below threshold 0.5) is fine. The point of the smoke test is: did `Predict` run end-to-end without throwing on shape mismatches? If it throws, the assertion bug is in the model load or pipeline geometry.
|
||||||
|
|
||||||
|
- [ ] **Step 5: Build locally**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
dotnet build tests/02-wakeword-cs -c Release
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: success. Common compile errors:
|
||||||
|
|
||||||
|
- `NamedOnnxValue` / `DenseTensor` / `InferenceSession` not found → confirm `using Microsoft.ML.OnnxRuntime;` and `using Microsoft.ML.OnnxRuntime.Tensors;` are at top.
|
||||||
|
- `'InputMetadata' has no Single()` → add `using System.Linq;` (with `ImplicitUsings` enabled this should already be there, but the spike didn't need it).
|
||||||
|
- API surface drift in `Microsoft.ML.OnnxRuntime` `<ORT_VER>` (the package occasionally renames `NamedOnnxValue.CreateFromTensor`) → consult Context7 `/microsoft/onnxruntime` docs for the pinned version. Do NOT silently switch to a different API; document the version-specific shape if it diverges.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Run on the Pi via `bin/probe-cs-2`**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
bin/probe-cs-2
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected output:
|
||||||
|
|
||||||
|
```
|
||||||
|
Loading WakewordModel...
|
||||||
|
Loaded in 0.XXs.
|
||||||
|
Silence test: <small>/50 non-zero scores (expected: 0 — but small drift is OK).
|
||||||
|
Predict pipeline ran without throwing.
|
||||||
|
```
|
||||||
|
|
||||||
|
If the run aborts with an `InvalidOperationException` from one of the shape assertions, the model in the repo disagrees with what this plan expects — STOP. Either the vendored ONNX file is from a newer openwakeword version with a different shape (re-verify the SHA-256 hashes against Step 8 of Task 1), or our shape derivation from `oww-utils.py` was wrong. Investigate by re-running the Task 1 spike and comparing dimensions before adjusting `WakewordModel.cs`.
|
||||||
|
|
||||||
|
- [ ] **Step 7: Commit**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git add tests/02-wakeword-cs/WakewordModel.cs tests/02-wakeword-cs/Program.cs
|
||||||
|
git commit -m "WakewordModel: port openwakeword streaming pipeline (mel→emb→cls)"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 4: `Program.cs` — live input stream + main inference loop (detection-only, no beep)
|
||||||
|
|
||||||
|
**Goal:** Replace the load-only smoke test with the full input stream + inference loop. Detection lines print to stdout but no beep yet — we want detection working in isolation before adding the second PortAudio stream.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `tests/02-wakeword-cs/Program.cs` (full rewrite)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Rewrite `tests/02-wakeword-cs/Program.cs`**
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using System.Collections.Concurrent;
|
||||||
|
using System.Diagnostics;
|
||||||
|
using System.Runtime.InteropServices;
|
||||||
|
using PortAudioSharp;
|
||||||
|
using WakewordProbe;
|
||||||
|
using Stream = PortAudioSharp.Stream;
|
||||||
|
|
||||||
|
const int SampleRate = WakewordModel.SampleRate;
|
||||||
|
const int Channels = 1;
|
||||||
|
const uint BlockFrames = WakewordModel.FrameSamples; // 1280
|
||||||
|
const float Threshold = 0.5f;
|
||||||
|
const double CooldownSeconds = 1.0;
|
||||||
|
|
||||||
|
// .NET-side mirror for any readers that go through Environment.GetEnvironmentVariable,
|
||||||
|
// AND libc setenv so PortAudio / onnxruntime (which use getenv()) see the values.
|
||||||
|
Environment.SetEnvironmentVariable("PA_ALSA_PLUGHW", "1");
|
||||||
|
Libc.setenv("PA_ALSA_PLUGHW", "1", 1);
|
||||||
|
Environment.SetEnvironmentVariable("ORT_LOGGING_LEVEL", "3");
|
||||||
|
Libc.setenv("ORT_LOGGING_LEVEL", "3", 1);
|
||||||
|
|
||||||
|
PortAudio.Initialize();
|
||||||
|
WakewordModel? model = null;
|
||||||
|
try
|
||||||
|
{
|
||||||
|
int device = FindUsbDevice();
|
||||||
|
Console.WriteLine($"Using device {device} ('{PortAudio.GetDeviceInfo(device).name}')");
|
||||||
|
|
||||||
|
Console.WriteLine("Loading WakewordModel...");
|
||||||
|
var t0 = DateTime.UtcNow;
|
||||||
|
string modelsDir = Path.Combine(AppContext.BaseDirectory, "models");
|
||||||
|
model = new WakewordModel(
|
||||||
|
Path.Combine(modelsDir, "melspectrogram.onnx"),
|
||||||
|
Path.Combine(modelsDir, "embedding_model.onnx"),
|
||||||
|
Path.Combine(modelsDir, "alexa.onnx"));
|
||||||
|
Console.WriteLine($"Loaded in {(DateTime.UtcNow - t0).TotalSeconds:0.00}s.");
|
||||||
|
|
||||||
|
var queue = new BlockingCollection<short[]>(boundedCapacity: 16);
|
||||||
|
using var cts = new CancellationTokenSource();
|
||||||
|
Console.CancelKeyPress += (_, e) => { e.Cancel = true; cts.Cancel(); };
|
||||||
|
|
||||||
|
var inParams = new StreamParameters
|
||||||
|
{
|
||||||
|
device = device,
|
||||||
|
channelCount = Channels,
|
||||||
|
sampleFormat = SampleFormat.Int16,
|
||||||
|
suggestedLatency = PortAudio.GetDeviceInfo(device).defaultLowInputLatency,
|
||||||
|
hostApiSpecificStreamInfo = IntPtr.Zero,
|
||||||
|
};
|
||||||
|
|
||||||
|
Stream.Callback callback = (IntPtr input, IntPtr _, uint frameCount,
|
||||||
|
ref StreamCallbackTimeInfo _2, StreamCallbackFlags status, IntPtr _3) =>
|
||||||
|
{
|
||||||
|
if (status.HasFlag(StreamCallbackFlags.InputOverflow))
|
||||||
|
Console.Error.WriteLine("[status] input overflow");
|
||||||
|
|
||||||
|
if (frameCount != BlockFrames)
|
||||||
|
{
|
||||||
|
Console.Error.WriteLine($"[status] unexpected callback frameCount={frameCount} (want {BlockFrames})");
|
||||||
|
return StreamCallbackResult.Continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
var frame = new short[BlockFrames];
|
||||||
|
unsafe
|
||||||
|
{
|
||||||
|
short* src = (short*)input.ToPointer();
|
||||||
|
fixed (short* dst = &frame[0])
|
||||||
|
Buffer.MemoryCopy(src, dst, BlockFrames * sizeof(short), BlockFrames * sizeof(short));
|
||||||
|
}
|
||||||
|
if (!queue.TryAdd(frame, 0))
|
||||||
|
Console.Error.WriteLine("[status] consumer behind, dropping frame");
|
||||||
|
return StreamCallbackResult.Continue;
|
||||||
|
};
|
||||||
|
|
||||||
|
using var stream = new Stream(
|
||||||
|
inParams, null, SampleRate, BlockFrames, StreamFlags.NoFlag, callback, IntPtr.Zero);
|
||||||
|
stream.Start();
|
||||||
|
Console.WriteLine("Listening. Say 'alexa'. Ctrl-C to exit.");
|
||||||
|
|
||||||
|
var sw = Stopwatch.StartNew();
|
||||||
|
TimeSpan lastTrigger = TimeSpan.FromSeconds(-CooldownSeconds);
|
||||||
|
|
||||||
|
try
|
||||||
|
{
|
||||||
|
while (!cts.IsCancellationRequested)
|
||||||
|
{
|
||||||
|
short[] frame = queue.Take(cts.Token);
|
||||||
|
float score = model.Predict(frame);
|
||||||
|
var now = sw.Elapsed;
|
||||||
|
if (score >= Threshold && (now - lastTrigger).TotalSeconds >= CooldownSeconds)
|
||||||
|
{
|
||||||
|
Console.WriteLine($"DETECTED alexa score={score:0.000} t={now.TotalSeconds:0.0}s");
|
||||||
|
lastTrigger = now;
|
||||||
|
// Beep is added in Task 5.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
catch (OperationCanceledException) { /* Ctrl-C path */ }
|
||||||
|
|
||||||
|
stream.Stop();
|
||||||
|
Console.WriteLine("\nbye");
|
||||||
|
}
|
||||||
|
finally
|
||||||
|
{
|
||||||
|
model?.Dispose();
|
||||||
|
PortAudio.Terminate();
|
||||||
|
}
|
||||||
|
|
||||||
|
static int FindUsbDevice()
|
||||||
|
{
|
||||||
|
int count = PortAudio.DeviceCount;
|
||||||
|
for (int i = 0; i < count; i++)
|
||||||
|
{
|
||||||
|
var info = PortAudio.GetDeviceInfo(i);
|
||||||
|
if (info.name.ToLowerInvariant().Contains("usb") && info.maxInputChannels >= 1)
|
||||||
|
return i;
|
||||||
|
}
|
||||||
|
Console.Error.WriteLine("USB audio device not found. Devices:");
|
||||||
|
for (int i = 0; i < count; i++)
|
||||||
|
{
|
||||||
|
var info = PortAudio.GetDeviceInfo(i);
|
||||||
|
Console.Error.WriteLine(
|
||||||
|
$" [{i}] {info.name} in={info.maxInputChannels} out={info.maxOutputChannels}");
|
||||||
|
}
|
||||||
|
Environment.Exit(1);
|
||||||
|
return -1; // unreachable
|
||||||
|
}
|
||||||
|
|
||||||
|
static class Libc
|
||||||
|
{
|
||||||
|
[DllImport("libc", EntryPoint = "setenv")]
|
||||||
|
public static extern int setenv(string name, string value, int overwrite);
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Key reuses from Test 1's `Program.cs` (`tests/01-record-play-cs/Program.cs`): the `using Stream = PortAudioSharp.Stream;` disambiguation, the `Libc.setenv` P/Invoke, the `FindUsbDevice` body, the `unsafe` callback `Buffer.MemoryCopy` pattern, the dual `Environment.SetEnvironmentVariable` + `Libc.setenv` calls.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Build locally**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
dotnet build tests/02-wakeword-cs -c Release
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: success.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Deploy + run on the Pi**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
bin/probe-cs-2
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected on the Pi:
|
||||||
|
|
||||||
|
```
|
||||||
|
Using device <N> ('USB ...')
|
||||||
|
Loading WakewordModel...
|
||||||
|
Loaded in 0.XXs.
|
||||||
|
Listening. Say 'alexa'. Ctrl-C to exit.
|
||||||
|
```
|
||||||
|
|
||||||
|
Then say "alexa" 3-4 times at normal volume from ~1 m. Expected: one `DETECTED alexa score=0.XXX t=YYs` line per spoken wakeword, fired within ~1 s of saying it. **No beep yet** — that's Task 5.
|
||||||
|
|
||||||
|
What to check before moving on:
|
||||||
|
|
||||||
|
- Detection actually fires. If 0/4 attempts trigger, the port has a bug — re-verify shapes and the `x/10 + 2` mel transform before continuing.
|
||||||
|
- No sustained `[status] input overflow` or `[status] consumer behind` lines. Occasional ones at startup are OK.
|
||||||
|
- No ORT GPU warnings (`/sys/class/drm/card0` etc.) — `ORT_LOGGING_LEVEL=3` should silence them. If they appear, the `Libc.setenv` for ORT_LOGGING_LEVEL isn't being honoured by the onnxruntime version pinned; consult Context7 for the correct env var name for that version.
|
||||||
|
- Ctrl-C from the workstation cleanly exits the probe (prints `bye`). If Ctrl-C just hangs the SSH session, check that `bin/probe-cs-2` uses `ssh -t`.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Commit**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git add tests/02-wakeword-cs/Program.cs
|
||||||
|
git commit -m "Live wakeword loop: PortAudio input + queue + threshold detection"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 5: Fire-and-forget beep + full hardware verification + `findings.md` update
|
||||||
|
|
||||||
|
**Goal:** Add the detection beep without blocking input ingestion, then run the full hardware test against all three pass criteria from the spec, and capture findings.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `tests/02-wakeword-cs/Program.cs` (add beep helper + wire it into the detection branch)
|
||||||
|
- Modify: `findings.md` (append Test 2 outcome section)
|
||||||
|
|
||||||
|
- [ ] **Step 1: Add the beep buffer + dispatch helper to `Program.cs`**
|
||||||
|
|
||||||
|
At the top of `Program.cs`, after the `const double CooldownSeconds = 1.0;` line, add:
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
const double BeepHz = 880.0;
|
||||||
|
const double BeepSeconds = 0.2;
|
||||||
|
```
|
||||||
|
|
||||||
|
Before the line `Console.WriteLine("Listening. Say 'alexa'. Ctrl-C to exit.");`, add the beep buffer precomputation:
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
short[] beepBuffer = MakeBeep(BeepHz, BeepSeconds, SampleRate);
|
||||||
|
```
|
||||||
|
|
||||||
|
In the detection branch (inside the main loop, after `lastTrigger = now;`), add:
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
FireAndForgetBeep(device, beepBuffer);
|
||||||
|
```
|
||||||
|
|
||||||
|
At the bottom of the file (after `FindUsbDevice` but before `class Libc`), add:
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
static short[] MakeBeep(double freqHz, double durationS, int sampleRate)
|
||||||
|
{
|
||||||
|
int n = (int)(sampleRate * durationS);
|
||||||
|
var buf = new short[n];
|
||||||
|
for (int i = 0; i < n; i++)
|
||||||
|
{
|
||||||
|
double t = i / (double)sampleRate;
|
||||||
|
double v = 0.3 * Math.Sin(2.0 * Math.PI * freqHz * t);
|
||||||
|
buf[i] = (short)(v * short.MaxValue);
|
||||||
|
}
|
||||||
|
return buf;
|
||||||
|
}
|
||||||
|
|
||||||
|
static void FireAndForgetBeep(int device, short[] beepBuffer)
|
||||||
|
{
|
||||||
|
int offset = 0;
|
||||||
|
var done = new ManualResetEventSlim(false);
|
||||||
|
|
||||||
|
var outParams = new StreamParameters
|
||||||
|
{
|
||||||
|
device = device,
|
||||||
|
channelCount = 1,
|
||||||
|
sampleFormat = SampleFormat.Int16,
|
||||||
|
suggestedLatency = PortAudio.GetDeviceInfo(device).defaultLowOutputLatency,
|
||||||
|
hostApiSpecificStreamInfo = IntPtr.Zero,
|
||||||
|
};
|
||||||
|
|
||||||
|
Stream.Callback playCb = (IntPtr _, IntPtr output, uint frameCount,
|
||||||
|
ref StreamCallbackTimeInfo _2, StreamCallbackFlags _3, IntPtr _4) =>
|
||||||
|
{
|
||||||
|
int remaining = beepBuffer.Length - offset;
|
||||||
|
int take = (int)Math.Min(frameCount, (uint)remaining);
|
||||||
|
int silence = (int)frameCount - take;
|
||||||
|
unsafe
|
||||||
|
{
|
||||||
|
short* dst = (short*)output.ToPointer();
|
||||||
|
if (take > 0)
|
||||||
|
{
|
||||||
|
fixed (short* src = &beepBuffer[offset])
|
||||||
|
Buffer.MemoryCopy(src, dst, take * sizeof(short), take * sizeof(short));
|
||||||
|
offset += take;
|
||||||
|
}
|
||||||
|
for (int i = take; i < frameCount; i++) dst[i] = 0;
|
||||||
|
}
|
||||||
|
if (offset >= beepBuffer.Length)
|
||||||
|
{
|
||||||
|
done.Set();
|
||||||
|
return StreamCallbackResult.Complete;
|
||||||
|
}
|
||||||
|
return StreamCallbackResult.Continue;
|
||||||
|
};
|
||||||
|
|
||||||
|
var stream = new Stream(
|
||||||
|
null, outParams, WakewordModel.SampleRate, 1024, StreamFlags.NoFlag, playCb, IntPtr.Zero);
|
||||||
|
Task.Run(() =>
|
||||||
|
{
|
||||||
|
done.Wait();
|
||||||
|
// Brief drain pause matches Test 1's behaviour; the device buffer needs a moment after Complete.
|
||||||
|
Thread.Sleep(200);
|
||||||
|
stream.Stop();
|
||||||
|
stream.Dispose();
|
||||||
|
done.Dispose();
|
||||||
|
});
|
||||||
|
stream.Start();
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Note: `MakeBeep` and `FireAndForgetBeep` are `static` local-method-style helpers; with top-level statements they live at file scope alongside `FindUsbDevice` and `Libc`. They reference `Stream` (the `using` alias) and `PortAudio.GetDeviceInfo` — both already in scope.
|
||||||
|
|
||||||
|
The `Task.Run` is captured by the closure but not awaited; the main loop returns immediately after `stream.Start()`. The stream + event handle are kept alive by the closure until the task disposes them.
|
||||||
|
|
||||||
|
- [ ] **Step 2: Build locally**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
dotnet build tests/02-wakeword-cs -c Release
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: success.
|
||||||
|
|
||||||
|
- [ ] **Step 3: Run on the Pi and verify pass criterion 1 (true-positive rate)**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
bin/probe-cs-2
|
||||||
|
```
|
||||||
|
|
||||||
|
Wait for `Listening...`. From ~1 m, say "alexa" at normal volume, **10 times**, pausing ≥ 2 seconds between attempts. Count printed `DETECTED ...` lines AND audible beeps.
|
||||||
|
|
||||||
|
**Pass criterion 1:** ≥ 8/10 detections, each within ~1 s of saying the word.
|
||||||
|
|
||||||
|
If you get fewer: the issue is in the inference pipeline. Common causes:
|
||||||
|
- `x/10 + 2` mel transform wasn't applied (search `WakewordModel.cs` for `/10` — if missing, that's it).
|
||||||
|
- Off-by-one in the mel window slice (`-EmbeddingWindowMelFrames` vs `-EmbeddingWindowMelFrames-1`).
|
||||||
|
- Dtype mismatch — passing int16 directly instead of converting to float32.
|
||||||
|
|
||||||
|
- [ ] **Step 4: Verify pass criterion 2 (false-positive rate)**
|
||||||
|
|
||||||
|
While the probe is still running, read a newspaper or book aloud at normal volume from ~1 m for **3 continuous minutes** (anything that isn't "alexa"). Count any spurious `DETECTED ...` lines.
|
||||||
|
|
||||||
|
**Pass criterion 2:** ≤ 1 false positive per minute (i.e., ≤ 3 over the 3-minute test).
|
||||||
|
|
||||||
|
If you get more: the threshold (0.5) was right for the Python probe; if C# is over-triggering, the pipeline is producing systematically higher scores than Python — likely a normalization or transform bug (revisit Step 3 troubleshooting).
|
||||||
|
|
||||||
|
- [ ] **Step 5: Verify pass criterion 3 (CPU budget)**
|
||||||
|
|
||||||
|
In a **second terminal**, while the probe is still running:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sshpass -p assistant ssh pi@192.168.50.115 htop
|
||||||
|
```
|
||||||
|
|
||||||
|
Find the `Probe` process row. Note the `CPU%` column over ~30 seconds while the probe is doing inference (speak occasionally to keep it busy).
|
||||||
|
|
||||||
|
**Pass criterion 3:** CPU% stays under 50% of one core (the Pi 4 has 4 cores, so htop shows "100%" per core — pass = stays below 50 in that column).
|
||||||
|
|
||||||
|
If higher: confirm `WakewordModel`'s `SessionOptions` set `IntraOpNumThreads = 1` and `InterOpNumThreads = 1` (Task 3 Step 1). If they're set and it's still hot, ONNX Runtime may be ignoring the limit — set the env var via `Libc.setenv("OMP_NUM_THREADS", "1", 1)` at the top of `Program.cs` alongside the others, redeploy, recheck.
|
||||||
|
|
||||||
|
Ctrl-C the probe when done. Quit htop.
|
||||||
|
|
||||||
|
- [ ] **Step 6: Commit the beep code**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git add tests/02-wakeword-cs/Program.cs
|
||||||
|
git commit -m "Beep on detection: fire-and-forget OutputStream, input keeps flowing"
|
||||||
|
```
|
||||||
|
|
||||||
|
- [ ] **Step 7: Append Test 2 outcome section to `findings.md`**
|
||||||
|
|
||||||
|
Open `findings.md` and append, after the existing `## C# probe outcome (2026-06-12 — Test 1 ported to C#)` section, a new section in the same shape:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
## C# wakeword probe outcome (2026-06-12 — Test 2 ported to C#)
|
||||||
|
|
||||||
|
The Python Test 2 (openwakeword "alexa" listener with beep on detect) was re-implemented in C# / .NET 9, with the openwakeword `Model.predict()` streaming pipeline ported to `Microsoft.ML.OnnxRuntime`. Probe lives at `tests/02-wakeword-cs/`; deploy script `bin/probe-cs-2`. All three pass criteria from the original spec held on hardware.
|
||||||
|
|
||||||
|
### What we proved
|
||||||
|
|
||||||
|
- **Microsoft.ML.OnnxRuntime <ORT_VER>** runs on linux-arm64 from a self-contained `dotnet publish`. NuGet [bundles / does not bundle — confirm] `libonnxruntime.so` for the RID.
|
||||||
|
- The three openwakeword ONNX models (`melspectrogram.onnx`, `embedding_model.onnx`, `alexa.onnx`, vendored from openwakeword 0.6.0) load and chain correctly when fed real audio.
|
||||||
|
- **Two PortAudio streams on one USB device works**: the input stream stays open while a per-detection fire-and-forget output stream plays the beep. No errors, no input overflow during playback (fixing the bug `findings.md` § A flagged in the Python probe).
|
||||||
|
- True-positive rate: <X>/10, false-positive rate: <Y>/3 min, CPU: <Z>% of one core. (Fill in observed values.)
|
||||||
|
|
||||||
|
### Key gotchas
|
||||||
|
|
||||||
|
- (Capture anything that took > 15 minutes to debug. Likely candidates: env-var name for ORT log level on this version; whether the mel transform `x/10 + 2` actually mattered in C#; whether `IntraOpNumThreads = 1` was honoured; whether shape assertions caught anything during development.)
|
||||||
|
|
||||||
|
### Open questions for the main assistant
|
||||||
|
|
||||||
|
- **Custom wakeword model.** Stock "alexa" works as a probe but the shipped assistant needs a custom "hey assistant" (or similar) model. openwakeword has a Colab notebook for synthetic training; budget ~1 hour to train + verify on the same pipeline this probe validates.
|
||||||
|
- **Speaker-is-mic self-trigger.** Not in probe scope, but the main assistant will need detector gating during TTS playback (findings.md § B), or hardware/software AEC.
|
||||||
|
- **Inference latency.** If `WakewordModel.Predict` ever exceeds 80 ms, the queue backs up; sample real measurements from this probe before sizing the queue / consumer thread in the main assistant.
|
||||||
|
```
|
||||||
|
|
||||||
|
Replace the bracketed placeholders (`<ORT_VER>`, `<X>`, `<Y>`, `<Z>`, the bundling note, the gotchas list) with the actual values you observed. **Do not commit with placeholders.**
|
||||||
|
|
||||||
|
- [ ] **Step 8: Commit findings**
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git add findings.md
|
||||||
|
git commit -m "Findings: C# wakeword probe outcome"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Notes on what is intentionally **not** in this plan
|
||||||
|
|
||||||
|
- **No automated tests.** The spec is explicit; verification is the listening / counting test on real hardware. The silence smoke test in Task 3 is a pipeline sanity check, not a model-correctness test.
|
||||||
|
- **No numerical parity test vs Python.** If hardware passes, the port is good enough for the probe. If hardware fails, the diagnostic levers in the spec (Section "Diagnostic levers if it fails") cover the likely causes.
|
||||||
|
- **No detector gating during beep playback.** The 200 ms beep at 880 Hz is far enough from human speech to not self-trigger the "alexa" classifier; if it did, the spec's § B (gate detector during TTS) is the main-assistant solution, not the probe's.
|
||||||
|
- **No retry / restart on PortAudio init failure.** Probe — crash and surface.
|
||||||
|
- **No abstraction over PortAudio.** Same one-file Program.cs shape as Test 1.
|
||||||
|
- **No custom wakeword.** Stock alexa only. Custom is a separate task per `findings.md` § C.
|
||||||
|
- **No bin/probe-cs-2 → bin/probe-cs unification.** Decided in brainstorm — kept as separate scripts.
|
||||||
|
|
||||||
|
If during implementation you feel the urge to add any of the above, stop and re-read the spec. The point of this probe is to be cheap and disposable.
|
||||||
Reference in New Issue
Block a user