AIToday
Large Language ModelsarXiv cs.CLPublished: Mar 25, 2026, 13:13 JST1 min read

New evaluation framework reveals that only 35% of LLM responses on sexual and reproductive health in Nepali meet quality standards, highlighting gaps in low-resource language support for sensitive health topics.

New evaluation framework reveals that only 35% of LLM responses on sexual and reproductive health in Nepali meet quality standards, highlighting gaps in low-resource language support for sensitive health topics.

3 Key Points

  1. Researchers introduced LEAF (LLM Evaluation Framework) to assess AI responses on sexual and reproductive health (SRH) queries in Nepali, a low-resource language

  2. The framework evaluates responses across four dimensions: accuracy, language quality, usability gaps (relevance, adequacy, cultural appropriateness), and safety gaps (safety, sensitivity, confidentiality)

  3. Study analyzed 14,000 SRH queries from over 9,000 Nepali users with manual annotations by SRH experts, finding only 35.1% of LLM responses met quality standards

  4. Current LLM evaluation methods focus primarily on accuracy for objective queries in high-resource languages, overlooking usability and safety needs for culturally sensitive health topics

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DataAgent launches with $10M to auto-fix Kubernetes faultsSiliconANGLE AI · 49m ago
  • SK Hynix custom HBM boosts inference up to 5.15xDIGITIMES Asia · 49m ago
  • Nvidia Earnings: Boring by Design, Avoiding a Consolidated WorldStratechery (Ben Thompson) · 49m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNew Progressive Quantization method addresses a fundamental flaw in vector tokenization used by multimodal AI models