AIToday
Large Language ModelsHacker NewsPublished: Sep 12, 2026, 22:00 JST2 min read

NüshuRescue: 35 sentence pairs teach GPT-4 Turbo a lost script

NüshuRescue: 35 sentence pairs teach GPT-4 Turbo a lost script

3 Key Points

  1. What happened

    Dartmouth's Ivory Yang and collaborators built NüshuRescue, which trained GPT-4 Turbo on 35 Chinese-Nüshu sentence pairs and translated unseen test phrases into the script.

  2. Why it matters

    The team's 500-pair expert-validated Nüshu-Chinese dataset is the first of its kind, and the framework minimizes human annotation, so it could extend to languages like Cherokee.

  3. What to watch

    Vosoughi warns these models can introduce dominant-culture biases, so authenticity hinges on native speakers and linguists working alongside the AI. Next, Yang wants computer-vision models for Nüshu on handkerchiefs and fans.

WHO IT HITSThis lands hardest on linguists and community members documenting endangered languages, who can now use sparse datasets instead of years of manual annotation. It also matters to researchers evaluating tools like Google's LangID, which misidentified Navajo sentences.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Nüshu was created four centuries ago by Yao women in Hunan as a secret script, but its use declined in the 1900s once women gained broader access to formal education, and many texts were lost or destroyed. Since the turn of this century, China has sustained an effort to save the script from extinction. Yang's work fits into that longer arc. As a child she learned a few words of Nüshu from her grandmother, and she is now applying AI to the problem. The project began with A Compendium of Chinese Nüshu, the most comprehensive, expert-validated collection of scanned scripts and their Chinese translations. From it, the team worked with annotators trained in computational linguistics to create 500 digitized sentence pairs, including newly mapped words.

The Dartmouth work sits alongside a broader examination of how existing language technologies treat underrepresented languages. Yang's recent paper, to be presented at the Nations of the Americas Chapter of the Association for Computational Linguistics conference, found that Google's LangID does not support most Native American languages and misidentified Navajo sentences as unrelated languages. Separately, linguistics professor Rolando Coto Solano built automatic speech-recognition models for Cook Islands Māori, Bribri, and Cabécar, motivated by a colleague's remark that she might die before finishing transcribing hundreds of hours of recordings.

Vosoughi's caution that these models can carry biases from dominant cultures, distorting or oversimplifying cultural identities, suggests the stakes rest on collaboration. He argues that active participation from native speakers and linguists is essential for linguistic authenticity, and that AI and community expertise are both fundamental. Whether NüshuRescue's low-data approach scales to other low-resource languages may depend on how closely communities shape that work, rather than on the model alone.

FAQ
How much data did the AI need to learn Nüshu?
Just 35 pairs of matching Chinese and Nüshu sentences were enough for GPT-4 Turbo to start translating test phrases not in its training data. The team built a larger dataset of 500 digitized sentence pairs for that training.
Can this method be used for other endangered languages?
Yes, the researchers say the framework can potentially be adapted to other low-resource languages such as Cherokee, since it minimizes reliance on extensive human annotations.
What did the researchers find about Google Translate and Navajo?
Yang and collaborators found that Google's LangID misidentified Navajo sentences as other, unrelated languages. They then built a language-identification model for Navajo and related Athabaskan languages.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Meta's Muse AI agent prices in at $20 and $100 a monthYahoo Finance AI · 1h ago
  • OpenAI agents hit RubyGems with 2,000+ malicious packagesTHE DECODER · 1h ago
  • AI hangover hits firms; cure isn't more GenAI useFortune AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAnthropic finds Claude misuse everywhere