Scoring Japanese Pronunciation Entirely On-Device
DTW alignment, a Rust engine behind Dart FFI, signed lesson packs — and the calibration questions synthetic data cannot answer yet.
Executive summary
Kaiwa Coach scores spoken Japanese on the phone itself: no audio leaves the device. Getting there meant building a Rust speech engine (VAD, YIN pitch tracking, mora segmentation, DTW alignment), shipping licensed content as cryptographically signed packs, and fighting build-system issues that made a working engine look broken. This case study covers the architecture, the specific bugs that shaped it, and an honest account of what still needs real-speaker validation.
The Challenge
Cloud speech recognition is the default way to score pronunciation: upload audio, get a verdict per word. For a Japanese shadowing app that approach fails on three counts — cost per minute, privacy expectations for voice data, and offline usefulness. The app had to score mora timing, pitch and intelligibility locally, on mid-range phones, while shipping content it could prove had not been tampered with.
Key problems
- 1Mora-level feedback is required, but ASR transcripts do not carry pitch or timing detail
- 2Voice data is sensitive — users expect pronunciation practice to stay on their device
- 3Learners practice offline (commutes, flights), so scoring cannot depend on a network round trip
- 4Lesson content needed clear licensing and tamper-evident packaging
Constraints
- !On-device CPU budget on mid-range Android and iOS devices
- !Flutter app calling a native engine through Dart FFI
- !Only CC0 corpora could ship (Common Voice ja)
- !iOS static linking quirks with Rust LTO builds
An On-Device Speech Engine in Rust
The engine pipeline is deliberately traditional signal processing rather than a neural scorer: voice activity detection segments the utterance, mora segmentation aligns it to the expected reading, YIN tracks pitch, and DTW aligns the learner contour against a native reference contour. Every stage is testable in isolation, which mattered when results looked wrong.
VAD and mora segmentation
Voice activity detection cuts the utterance; mora boundaries are estimated from the expected kana reading rather than from energy alone. An early version split the VAD span evenly across morae — which turned out to make long-vowel verdicts a speaking-rate artifact.
- Mora segmentation derived from the lesson reading, not audio energy
- Content-mismatch gate: refuse to score when the recording does not match the exercise
- Median mora duration measured at ~130 ms — the commonly assumed 200 ms was wrong for real speech
YIN pitch tracking and DTW alignment
Pitch contours are tracked with YIN and aligned with dynamic time warping. Reference contours must be regridded to the engine 16 ms hop, because DTW step penalties assume near-diagonal paths — mixing frame rates silently biases every score.
- Reference contours regridded to the engine 16 ms hop
- Pipeline vs engine pitch verified: 100% voicing agreement, 0.018 cents median error
- Per-mora outputs: timing, pitch, intelligibility
Signed, versioned lesson packs
107 lessons ship as versioned packs signed with Ed25519 and verified by SHA-256, with rollback support. Content was mined from Common Voice ja 26.0: 584,920 clips filtered down to 7,048 N5-friendly candidates, then to 87 pronunciation lessons plus 20 shadowing lessons.
- Signed pack manifest with checksum verification and rollback
- Contours-only packs cut app assets from 9.6 MB to 988 KB
- Reward and entitlement logic gated server-side
Engine behind Dart FFI
The Flutter app calls the Rust engine through FFI. This worked on Android immediately and failed completely on iOS — Rust LTO was stripping every extern symbol from the static library, so 0 of 9 expected symbols were exported.
- LTO disabled for the FFI build plus -u linker flags on iOS
- 74 engine tests run against the same code the app links
- No JNI/bridge drift: one engine, both platforms
Key technical decisions
DTW over a neural pronunciation model
A traditional alignment pipeline is explainable per mora, runs within the CPU budget, and can be tested without a labeled dataset. A neural scorer would need training data we did not have and would hide failure modes instead of exposing them.
Contours-only packs first, audio optional later
Shipping reference contours brought app assets from 9.6 MB to 988 KB, making the content pipeline provable before optimizing audio delivery.
Sign packs at build time, verify at runtime
A content pipeline that can silently fail to verify (a hardcoded checksum placeholder did exactly that once) is worse than no signing. Verification is a hard gate with tests.
Implementation
Project timeline
Phase 0 — research spike
- Corpus mining from Common Voice ja with licensing review
- Prototype VAD + YIN + DTW in Rust
- Flutter shell with encrypted local storage
Engine hardening
- Per-mora scoring v3 with content-mismatch gate
- iOS symbol export fixed (LTO off, -u flags)
- Pitch power gate recalibrated against real speech
Content pipeline
- Ed25519 + SHA-256 pack signing with rollback
- 107 lessons packed; pipeline tests to guard regressions
- Strict checksum verification with fail-closed behavior
App integration
- Pitch Studio F0 contour UI
- Spaced-repetition error profiles per lesson
- Account deletion, storage and entitlement fixes
Implementation challenges
Long-vowel detection sat at 20% recall while geminates were not detected at all, against an >80% gate.
Solution: The cause was structural: VAD spans were split evenly across morae, so longer vowels looked like slower speech. Segmentation now derives from the expected reading and real timing distributions.
The engine worked on Android and exported 0 of 9 symbols on iOS.
Solution: Rust LTO dropped extern "C" symbols from the iOS static library. Keep LTO off for the FFI build and add -u linker flags to force symbol retention.
A pitch power gate of 5.0 rejected every frame of real speech — real values sat around 1.7.
Solution: Recalibrated the gate against recorded human speech instead of synthetic tones, then added regression tests using real clips.
The pack verifier accepted anything because a placeholder checksum was never replaced.
Solution: The verification became fail-closed with the real SHA-256 manifest, and a pipeline test now fails if the placeholder string ever returns.
Results & Impact
Key achievements
Business impact
Lessons Learned
A benchmark is only as honest as its denominator
Synthetic calibration reported 100/100 reference frames and 10/10 long vowels — with tiny denominators and no geminate coverage. The numbers were technically true and practically misleading. Real-speaker validation is now a release gate.
FFI failures are usually build-system failures
Two days were lost to "the engine is broken" when the real cause was linker settings stripping symbols on one platform. Validate cross-platform symbol export in CI before writing engine logic.
Calibrate every gate against real data
A pitch power threshold tuned on synthetic tones rejected 100% of real speech. Any gate that reads the physical world needs recorded human input in its tests.
What we’d do differently
- Run a real-speaker calibration study before app-store submission
- Add geminate and long-vowel test cases with meaningful sample sizes
- Move recordings out of the OS temp directory with a retention policy
- Bundle CJK fonts to eliminate tofu on offline devices
Technologies used
The project
See the full project
More details, screenshots, and information about Kaiwa Coach.
View project