
Apple researchers have developed a new training method for speech recognition systems that handle code-switching—when speakers mix Mandarin and English in the same sentence.
The iterative pseudo-labeling approach uses unlabeled audio data to improve accuracy without requiring massive amounts of manually transcribed bilingual speech.
This addresses a long-standing challenge in ASR for multilingual speakers.
What happened
Apple researchers applied an iterative pseudo-labeling training approach to Mandarin-English code-switching ASR (automatic speech recognition) for the first time. The method uses three phases: pseudo-label generation from unlabeled data, two-stage bilingual model training, and iterative improvements to enhance performance on speech that mixes both languages within a single utterance.
Why it matters
Code-switching—alternating between Mandarin and English in the same spoken sentence—has been difficult for speech recognition systems because training data is scarce. The pseudo-labeling approach leverages large unlabeled datasets to create semi-supervised training material, potentially making speech recognition more accurate for bilingual speakers without requiring as much manually transcribed code-switched speech.
What to watch
The paper demonstrates the approach's effectiveness but does not yet state a deployment timeline or commercial availability. The research appears in Apple's machine learning publication, suggesting the company is exploring this technique for future ASR products.
Ask the AI about this article →
Code-switching presents a genuine technical challenge for automatic speech recognition: bilingual speakers naturally alternate between languages mid-sentence, but the volume of publicly available transcribed code-switched speech is far smaller than monolingual datasets. Traditional supervised learning approaches require large amounts of labeled training data, making them impractical for this use case. Apple's research addresses this data scarcity by adopting pseudo-labeling, a semi-supervised technique that leverages unlabeled audio. Rather than requiring human transcription of every code-switched utterance, the system first generates candidate transcriptions automatically, then refines them iteratively. The three-phase structure—initial pseudo-label generation, bilingual model training, and iterative improvement—is designed to progressively reduce transcription errors introduced by the initial automatic labeling, making the synthetic training data more reliable with each cycle.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Apple Music will add “Made With AI” labels to songs where “a material portion” was created using AI, starting…

A software engineer created Schmaudio, a platform that generates interactive audio stories where listeners mak…

Researchers tested 11 widely used open-source speech recognition models and found that several of the highest-…

Adobe is releasing three AI audio tools — Generate Music (royalty-free music for videos), Generate Speech (scr…

Adobe announced general availability of audio generation capabilities in Firefly, its creative AI suite
Scammers are increasingly using AI-powered voice-cloning and deepfake technology to imitate trusted contacts…
