AIToday
Audio & SpeechApple Machine LearningPublished: Aug 21, 2026, 10:00 JST2 min read

Apple improves code-switching speech recognition with iterative pseudo-labeling

Apple improves code-switching speech recognition with iterative pseudo-labeling

Key takeaway

  • Apple researchers have developed a new training method for speech recognition systems that handle code-switching—when speakers mix Mandarin and English in the same sentence.

  • The iterative pseudo-labeling approach uses unlabeled audio data to improve accuracy without requiring massive amounts of manually transcribed bilingual speech.

  • This addresses a long-standing challenge in ASR for multilingual speakers.

3 Key Points

  1. What happened

    Apple researchers applied an iterative pseudo-labeling training approach to Mandarin-English code-switching ASR (automatic speech recognition) for the first time. The method uses three phases: pseudo-label generation from unlabeled data, two-stage bilingual model training, and iterative improvements to enhance performance on speech that mixes both languages within a single utterance.

  2. Why it matters

    Code-switching—alternating between Mandarin and English in the same spoken sentence—has been difficult for speech recognition systems because training data is scarce. The pseudo-labeling approach leverages large unlabeled datasets to create semi-supervised training material, potentially making speech recognition more accurate for bilingual speakers without requiring as much manually transcribed code-switched speech.

  3. What to watch

    The paper demonstrates the approach's effectiveness but does not yet state a deployment timeline or commercial availability. The research appears in Apple's machine learning publication, suggesting the company is exploring this technique for future ASR products.

Ask the AI about this article →

Context & Analysis

Code-switching presents a genuine technical challenge for automatic speech recognition: bilingual speakers naturally alternate between languages mid-sentence, but the volume of publicly available transcribed code-switched speech is far smaller than monolingual datasets. Traditional supervised learning approaches require large amounts of labeled training data, making them impractical for this use case. Apple's research addresses this data scarcity by adopting pseudo-labeling, a semi-supervised technique that leverages unlabeled audio. Rather than requiring human transcription of every code-switched utterance, the system first generates candidate transcriptions automatically, then refines them iteratively. The three-phase structure—initial pseudo-label generation, bilingual model training, and iterative improvement—is designed to progressively reduce transcription errors introduced by the initial automatic labeling, making the synthetic training data more reliable with each cycle.

FAQ

What is code-switching in speech recognition?
Code-switching is alternating between two languages within the same spoken utterance—for example, a speaker saying a sentence that contains both Mandarin and English words. ASR systems struggle with this because training data for code-switched speech is limited.
How does the pseudo-labeling approach work?
The method generates pseudo-labels (automatically assigned transcriptions) from a large unlabeled audio corpus, creating a semi-supervised dataset. This is then used in two-stage bilingual model training followed by iterative refinements to improve code-switching speech recognition performance.
Apple Machine LearningRead Original Article

Get the latest Audio & Speech news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSchwab Opens India Tech Hub, Plans 2,000 Workers by 2027