Language is a realtime speech translation app for noisy, multi-speaker rooms.
The product goal is to listen to overlapping speech from multiple people, detect each source language, translate it to English, synthesize the English audio in the same speaker voice at the same perceived volume, and suppress the original source voice so the user hears the translated mix instead.
The repo is still mock-first, but the scaffolding, evidence gates, and audio fixtures are now shaped around that full audio loop.
apps/field_app_flutter- Flutter field console.services/gateway- FastAPI gateway, mock scene orchestration, session API, and provider-facing adapter boundaries.crates/- Rust realtime session primitives, prioritization policy, and protobuf bindings.proto/session.proto- canonical shared contract.fixtures/audio_eval- deterministic audio-evaluation fixtures.scripts/- local, Docker, audio, release-gate, and packaging commands.docs/- development, release, testing, architecture, and research notes.
Bootstrap local SDKs and the gateway virtual environment:
bash scripts/bootstrap_dev.shRun the local smoke baseline:
make smoke-local-demoWindows without Make or Bash:
pwsh -NoProfile -File scripts/smoke_local_demo.ps1Run the gateway:
make gateway-runRun the Flutter app:
make flutter-runUse the categorized runner first. It keeps the normal path small by writing command output to log
files by default; pass --verbose only when you want live logs in the terminal.
python3 scripts/run_test_category.py quick
python3 scripts/run_test_category.py list
python3 scripts/run_test_category.py release-status
python3 scripts/run_test_category.py release-progress
python3 scripts/run_test_category.py release-artifacts
python3 scripts/run_test_category.py smoke-local
python3 scripts/run_test_category.py allWindows:
pwsh -NoProfile -File scripts/dev_container.ps1 test-category quick
pwsh -NoProfile -File scripts/dev_container.ps1 test-category list
pwsh -NoProfile -File scripts/dev_container.ps1 test-category release-status
pwsh -NoProfile -File scripts/dev_container.ps1 test-category release-progress
pwsh -NoProfile -File scripts/dev_container.ps1 test-category release-artifacts
pwsh -NoProfile -File scripts/dev_container.ps1 test-category smoke-local
pwsh -NoProfile -File scripts/dev_container.ps1 test-category allMake shortcuts:
make test-quick
make test-contracts
make test-audio-fixtures
make test-allCommon categories:
| Category | Use it for | Notes |
|---|---|---|
quick |
Fast local contract sanity checks | No Docker, model downloads, or hardware access. |
contracts |
Audio/report contract self-tests | Still local and hardware-free. |
core |
Repository contract, Rust, gateway, and app checks | Uses make check or scripts/check_local.ps1. |
smoke-local |
Local gateway smoke baseline | Verifies health, session, and SSE demo endpoints. |
audio-fixtures |
Disposable Docker audio fixtures | Capture, translation, playback, fallback TTS. |
hardware |
Host audio discovery and listener-ear planning | Run deliberately when testing devices. |
route-triage |
Host headphone route preflight and probe handoff | Prints the probe command; does not run it. |
guided-capture |
Strict host-guided listener-ear capture | Requires explicit device IDs, concrete labels, and a confirmed preflight report. |
physical-audio-handoff |
Current host devices, route, manual kit, and checklist | Best first command before a hardware session. |
evidence-kit |
Manual listener-ear recording kit/dropbox | Creates/checks the folder for the three release WAVs. |
stage-recordings-check |
Check recorder WAV exports | Validates the three take files without copying them. |
stage-recordings |
Stage recorder WAV exports | Copies/validates the three takes into the dropbox using exact release filenames. |
recording-status |
Listener-ear WAV readiness | Use after adding the three manual recordings. |
reference-playback-dry-run |
Validate manual reference playback routing | Requires source/headphone output env vars; plays no audio. |
recording-session-dry-run |
Rehearse a manual recording session | Prepares the kit, validates playback routing, and prints status without audio. |
reference-playback |
Play manual references for an external recorder | Explicit hardware action; not release evidence. |
release-evidence |
One-command listener-ear evidence handoff | Prepare/import/check the kit, then print release status. |
release-evidence-score |
Score complete listener-ear evidence | Requires real WAVs and concrete hardware labels. |
release-status |
Concise release blocker summary | Low-token next-action handoff; exits zero by default. |
release-progress |
Milestone percentages | Evidence-linked estimate for push summaries. |
release-artifacts |
Local source/gateway release handoff | Builds clean artifacts, checksums them, then smokes the packaged gateway wheel and write-token auth. |
release |
Strict release-gate status | Expected to fail until physical evidence is present. |
all |
Automated non-interactive suites | Excludes hardware, guided capture, release, optional model downloads, and artifact-dependent voice checks. |
Detailed test categories and the old command matrix live in
docs/testing/test-categories.md. Disposable environment details live in
docs/development/disposable-test-environments.md.
The contracts category creates .venv-audio-contract/ for fixture contract dependencies
(numpy, scipy, soundfile). The first run may need PyPI access or a warm pip cache; later runs
reuse the pinned local venv. Set LANGUAGE_AUDIO_CONTRACT_PYTHON only when you need to pin the base
interpreter.
On Windows, core can use a portable Flutter SDK at C:\tmp\flutter\bin\flutter.bat, or set
LANGUAGE_FLUTTER to another Flutter executable.
For host audio on Windows, Docker is the wrong boundary because Bluetooth, WASAPI, USB, and built-in devices need direct access to the host audio stack.
Start with a non-recording preflight:
pwsh -NoProfile -File scripts/headphone_isolation_local.ps1 -Action self-test
pwsh -NoProfile -File scripts/headphone_isolation_local.ps1 -Action list-devices
pwsh -NoProfile -File scripts/headphone_isolation_local.ps1 -Action preflight --sample-rate-hz 48000 --input-channels 1 --output-channels 2The wrapper auto-selects a supported Python >=3.11,<3.14; set LANGUAGE_PYTHON only when you need
to point it at a specific interpreter.
A laptop built-in mic and speakers are useful for route triage. Release evidence still needs real
listener-ear recordings: open-ear source, isolated source, and translated headphone playback from the
same listener-ear position. See docs/development/headphone-isolation-release-runbook.md.
To refresh preflight and print a deliberate non-release route probe command:
python scripts/run_test_category.py route-triageBefore a physical recording session, refresh route triage, prepare/check the kit, and write the current checklist:
python scripts/run_test_category.py physical-audio-handoffTo verify and then play the manual references through explicit outputs for an external recorder:
$env:LANGUAGE_SOURCE_OUTPUT_DEVICE = "REPLACE_WITH_SOURCE_OUTPUT_ID"
$env:LANGUAGE_HEADPHONE_OUTPUT_DEVICE = "REPLACE_WITH_HEADPHONE_OUTPUT_ID"
python scripts/run_test_category.py reference-playback-dry-run
python scripts/run_test_category.py recording-session-dry-run
python scripts/run_test_category.py reference-playbackIf preflight reports a capture-ready route and you have a real listener-ear mic in place, set the device IDs and labels, then run:
python scripts/run_test_category.py guided-capture --dry-run
python scripts/run_test_category.py guided-capturePrepare/check the manual recording kit, import any complete dropbox WAVs, and print release status:
python scripts/run_test_category.py release-evidenceIf recorder files have different names, stage them into the dropbox first:
$env:LANGUAGE_SOURCE_OPEN_EAR_RECORDING = "C:\Path\source-open.wav"
$env:LANGUAGE_SOURCE_ISOLATED_EAR_RECORDING = "C:\Path\source-isolated.wav"
$env:LANGUAGE_TRANSLATED_HEADPHONE_RECORDING = "C:\Path\translated-headphone.wav"
python scripts/run_test_category.py stage-recordings-check
python scripts/run_test_category.py stage-recordings
python scripts/run_test_category.py release-evidenceAfter the three WAVs are in the dropbox, set concrete labels and score the evidence:
$env:LANGUAGE_HEADPHONE_DEVICE_LABEL = "Sony WH-1000XM6"
$env:LANGUAGE_ISOLATION_FIXTURE_LABEL = "left earcup sealed over phone recorder"
$env:LANGUAGE_MEASUREMENT_MICROPHONE_LABEL = "phone WAV recorder at listener-ear point"
python scripts/run_test_category.py release-evidence-scoreBuild local source and gateway package artifacts from a clean tree, then smoke the packaged gateway:
python scripts/run_test_category.py release-artifactsThe focused commands remain available when you only want one part:
python scripts/run_test_category.py evidence-kit
python scripts/run_test_category.py recording-statusThe release gate is intentionally stricter than fixture tests:
python3 scripts/release_audio_gate.pyFor a shorter status and next-action handoff:
python3 scripts/run_test_category.py release-statusThat command also writes artifacts/release/physical-audio-checklist.md for the current hardware
handoff.
Use python3 scripts/release_audio_status.py --full-commands when you need the full hardware
command list in the terminal.
Use this for the milestone percentages reported after pushes:
python3 scripts/run_test_category.py release-progressIt writes:
artifacts/release/audio-gate-report.jsonartifacts/release/audio-gate-report.md
Use the Markdown report for the operator handoff and the JSON report as the authoritative pass/fail artifact. The current release path is still blocked until playback/source-suppression evidence is collected from real listener-ear recordings or true room-cancellation evidence.
This is mainly about keeping Codex/OpenAI development threads small. For long agent or CI runs, prefer the quiet default:
python3 scripts/run_test_category.py quick
python3 scripts/run_test_category.py all --continue-on-failureThe runner stores full logs under artifacts/test-categories/ and prints only summaries plus bounded
failure tails. Optional context-compression tools such as Headroom can be evaluated later, but they
should sit outside the release evidence path unless explicitly adopted and pinned.
When a development thread is close to the token limit, use release-status and release-progress
for compact handoff instead of pasting full reports or logs.
See docs/development/token-budget.md.
Use the sibling Robin checkout before choosing real audio models or providers:
python3 scripts/prepare_robin_research_pack.py
python3 scripts/prepare_robin_research_pack.py --check
python3 scripts/check_audio_corpus_catalog.pyResearch notes and decisions live under docs/research/.
The repo currently includes:
- typed Rust session and speaker primitives with prioritization tests
- FastAPI gateway with health, session, speaker, mock-scene, live-ingest, persistence, and SSE routes
- Flutter operator shell for speaker lanes, translated captions, audio metadata, and lock controls
- protobuf-backed contract generation across gateway, Flutter, and Rust
- disposable Docker profiles for core and audio-eval work
- deterministic and small real-speech audio fixtures for overlap, translation, playback, fallback TTS, diarization, target-speaker extraction, and same-voice candidate validation
- release-gate reporting that rejects fixture-only, synthetic, self-attested, or incomplete physical evidence
- Keep changes scoped to the subsystem you are touching.
- Add or update disposable test scaffolding when introducing new runtimes, SDKs, models, or providers.
- Run the smallest relevant category before handoff.
- Keep research-backed audio decisions in
docs/research/decisions/. - Keep the README short; put detailed runbooks in
docs/.