Voice AI in production. On-premise, low-latency, and measured.
PhD in privacy-preserving speech (Inria / University of Lille). Co-creator of the VoicePrivacy Challenge, the community benchmark for voice anonymization. Co-founder of Nijta, where we deploy speech systems inside regulated environments, mostly public safety, transport police and healthcare.
That work has been shaped by constraints rather than by scale: multilingual voice agents running entirely on Android, CPU-only on-premise deployment cleared through a government security review, and no data leaving the building. Most of it lives behind customer NDAs. This account is where I rebuild it in the open, on public models and public data.
- Latency lab. Where the wait goes in a cascaded voice agent, split into four stages and reported at the tail instead of the mean. Seven of ten turns sit on a 0.58s endpointing floor set by a VAD parameter that appears in neither of the two latency knobs people reach for. The same agent costs 1.2 or 2.4 LLM calls per turn depending only on how the caller speaks. Turns that produced no audio at all are counted rather than dropped, and claims that turned out to be wrong are recorded rather than edited away.
- Multilingual turn-taking evaluation. Full-duplex turn-taking is evaluated almost entirely in English, against metrics nobody has grounded in human preference. Both gaps are named by the benchmark authors themselves, and turn-taking is one of the most cross-linguistically variable things people do.
- 20+ peer-reviewed publications, 1,530 citations, h-index 15, i10-index 21 (Sep 2026) · Google Scholar
- Reviewer for NeurIPS, ICLR and Interspeech
- Organizing Committee, ISCA SPSC Symposium
- Previously: Microsoft Research (code-switched Hindi-English ASR), Inria


