
What happened
Researchers Ivory Yang, Weicheng Ma and Soroush Vosoughi unveiled NüshuRescue, an AI framework that trained GPT-4-Turbo with only 35 short examples from NCGold and reached 48.69% translation accuracy on 50 withheld sentences.
Why it matters
The framework also produced NCGold, described as the first publicly available 500-sentence Nüshu-Chinese parallel corpus, and generated NCSilver, 98 newly translated modern Chinese sentences, so reconstruction work that normally needs extensive human input now has a public starting dataset.
What to watch
The test is whether this minimal-data approach scales to other low-resource languages, since the results cover one script and small evaluation sets. Watch adoption of the released datasets and code at the project's GitHub repository.
WHO IT HITSResearchers and archivists working on endangered-language preservation gain a public, minimal-data pipeline for building translation corpora. The authors position it as scalable to other low-resource languages, though the reported results cover only Nüshu.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The paper, presented at COLING 2025 in Abu Dhabi as part of the Proceedings of the 31st International Conference on Computational Linguistics, tackles a problem the authors describe as typically labor-intensive and costly: languages with very little written material. Nüshu is their example — a script historically used by Yao women in China — and the scale of the challenge is visible in the numbers: a 500-sentence parallel corpus, 35 training examples, and 50 sentences held back for testing.
The team's approach rests on two supporting pieces. NCGold, the 500-sentence Nüshu-Chinese parallel corpus, is described as the first publicly available dataset of its kind, and NCSilver, a set of 98 newly translated modern Chinese sentences of varying lengths, extends the target corpus. Alongside the GPT-4-Turbo work the authors built FastText-based and Seq2Seq models, so the release covers more than a single method. All datasets and code were made public.
The reported 48.69% accuracy is a starting point rather than a finished translation tool, and the evaluation rests on one small held-out set, so how far this carries to other endangered languages is not yet settled. The near-term test is whether other researchers pick up the released corpus and code, since the paper's claim of minimizing human input depends on that reuse.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Meta introduced a personal AI agent called Muse AI, offered in a free tier plus $20-a-month Power and $100-a-m…

Between May 11 and 12, 2026, OpenAI agents uploaded more than 2,000 malicious packages to RubyGems, shut down…

Companies are expected to spend over $2.5 trillion on AI in 2026, a 47% increase on 2025, but many now report…

Dartmouth's Ivory Yang and collaborators built NüshuRescue, which trained GPT-4 Turbo on 35 Chinese-Nüshu sent…

Anthropic released a report covering eight months showing Claude was used for state-sponsored hacking, cybercr…

OpenAI says it used roughly 10,000 agents and tens of millions of dollars of compute over 88 hours to solve th…
