2026-06-09 · HuggingFace

Can Voice Agents Handle Bilingual Customers? Benchmarking Frontier ASR on Code-Switched Speech

modelsresearch

read at source ↗ huggingface.co

Can Voice Agents Handle Bilingual Customers? Benchmarking Frontier ASR on Code-Switched Speech

Source: HuggingFace Date: 2026-06-09 URL: https://huggingface.co/blog/ServiceNow-AI/code-switching

Summary

ServiceNow AI Research published a benchmark evaluating seven frontier ASR systems on code-switched speech — utterances where speakers alternate between two languages mid-sentence. The 918-utterance synthetic dataset spans four language pairs (Spanish-English, French-English, Canadian French-English, German-English) in HR and IT service domains. Top performers were ElevenLabs Scribe V2, Google Gemini 3 Flash, and AssemblyAI Universal 3-Pro. The study found that switching frequency predicts whether errors occur, while code-mixing density predicts error severity; errors concentrated on embedded English segments rather than at the switching points themselves.

Implications

  • Voice agents have a measurable bilingual gap. Enterprise deployments that assumed monolingual accuracy benchmarks carry over to bilingual customers are over-counting reliability. The benchmark gives a concrete failure-rate surface to reason from before production deployment.
  • Model selection for multilingual pipelines is non-trivial. Performance spread across the seven tested systems was significant, and the optimal choice varies by language pair. Teams choosing ASR for voice agents based on general WER leaderboards will likely pick wrong for code-switching workloads.
  • Feeds: voice AI capabilities, agentic coding tooling (agent pipeline reliability), voices/power dynamics (multilingual user representation in evaluation).

← all signals