AIToday
Large Language ModelsHacker NewsPublished: Sep 12, 2026, 22:00 JST2 min read

NüshuRescue hits 48.69% accuracy on endangered script

NüshuRescue hits 48.69% accuracy on endangered script

3 Key Points

  1. What happened

    Researchers Ivory Yang, Weicheng Ma and Soroush Vosoughi unveiled NüshuRescue, an AI framework that trained GPT-4-Turbo with only 35 short examples from NCGold and reached 48.69% translation accuracy on 50 withheld sentences.

  2. Why it matters

    The framework also produced NCGold, described as the first publicly available 500-sentence Nüshu-Chinese parallel corpus, and generated NCSilver, 98 newly translated modern Chinese sentences, so reconstruction work that normally needs extensive human input now has a public starting dataset.

  3. What to watch

    The test is whether this minimal-data approach scales to other low-resource languages, since the results cover one script and small evaluation sets. Watch adoption of the released datasets and code at the project's GitHub repository.

WHO IT HITSResearchers and archivists working on endangered-language preservation gain a public, minimal-data pipeline for building translation corpora. The authors position it as scalable to other low-resource languages, though the reported results cover only Nüshu.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The paper, presented at COLING 2025 in Abu Dhabi as part of the Proceedings of the 31st International Conference on Computational Linguistics, tackles a problem the authors describe as typically labor-intensive and costly: languages with very little written material. Nüshu is their example — a script historically used by Yao women in China — and the scale of the challenge is visible in the numbers: a 500-sentence parallel corpus, 35 training examples, and 50 sentences held back for testing.

The team's approach rests on two supporting pieces. NCGold, the 500-sentence Nüshu-Chinese parallel corpus, is described as the first publicly available dataset of its kind, and NCSilver, a set of 98 newly translated modern Chinese sentences of varying lengths, extends the target corpus. Alongside the GPT-4-Turbo work the authors built FastText-based and Seq2Seq models, so the release covers more than a single method. All datasets and code were made public.

The reported 48.69% accuracy is a starting point rather than a finished translation tool, and the evaluation rests on one small held-out set, so how far this carries to other endangered languages is not yet settled. The near-term test is whether other researchers pick up the released corpus and code, since the paper's claim of minimizing human input depends on that reuse.

FAQ
What is Nüshu?
Nüshu is a rare script historically used by Yao women in China for self-expression within a patriarchal society.
How much data did the AI need?
GPT-4-Turbo had no prior exposure to Nüshu and used only 35 short examples from NCGold to reach 48.69% translation accuracy on 50 withheld sentences.
Are the datasets and code available?
Yes. NCGold, a 500-sentence Nüshu-Chinese parallel corpus, is described as the first publicly available dataset of its kind, and all datasets and code were made publicly available at a GitHub repository.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Meta's Muse AI agent prices in at $20 and $100 a monthYahoo Finance AI · 1h ago
  • OpenAI agents hit RubyGems with 2,000+ malicious packagesTHE DECODER · 1h ago
  • AI hangover hits firms; cure isn't more GenAI useFortune AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAnthropic finds Claude misuse everywhere