AIToday
Audio & SpeecharXiv cs.CLPublished: Apr 1, 2026, 13:00 JST1 min read

New Thiomi Dataset provides over 601,000 text annotations and 385,000 audio recordings across ten African languages to advance low-resource language AI models.

New Thiomi Dataset provides over 601,000 text annotations and 385,000 audio recordings across ten African languages to advance low-resource language AI models.

3 Key Points

  1. Dataset covers ten African languages including Swahili, Kikuyu, Wolof, Somali, and Fulani across four language families, collected through a community platform with over 100 contributors

  2. Achieves 86-100% text approval rates through multi-tier quality assurance pipeline for six primary languages

  3. Establishes baselines for ASR, machine translation, and text-to-speech models, with best ASR system reaching 3.24% word error rate on Swahili—significantly improving prior academic performance from 8% WER

Ask the AI about this article →

Get the latest Audio & Speech news every morning

For example, today's edition would include:

  • Mitsubishi Electric develops task-general sound separation AITop Companies AI · 13h ago
  • Musician Detectives Hunt AI Music GriftersThe Verge AI · 2d ago
  • Beatport bans AI-generated music from DJ marketplaceTHE DECODER · 3d ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMeta's semi-formal reasoning technique improves LLM code review accuracy to 93% by requiring AI to explicitly trace execution paths before answering