AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryFortune AIPublished: Aug 27, 2026, 19:00 JST2 min read

Cantonese AI startup Votee challenges English-Mandarin dominance

Cantonese AI startup Votee challenges English-Mandarin dominance

Key takeaway

  • Hong Kong startup Votee AI builds Cantonese LLMs from open-weight models.

  • It trains on 500 million tokens, far less than English models.

  • The company sells to governments and banks, aiming for sovereign AI.

3 Key Points

  1. What happened

    Hong Kong-based Votee AI is training open-weight models like Meta's Llama and Alibaba's Qwen on Cantonese data, growing its corpus from 100 million to over 500 million tokens. CEO Pak-Sun Ting says the AI revolution is 'in English and Mandarin,' leaving other languages behind.

  2. Why it matters

    Cantonese is spoken by over 80 million people, yet mainstream models fail on cultural and local knowledge, per the HKCanto-Eval benchmarks. Votee sells its models to banks, universities, and government departments, and is profitable with revenues exceeding costs. It represents a push for 'sovereign AI' so countries aren't 'tethered' to overseas providers.

  3. What to watch

    Votee is in 'active discussions' with AI Singapore and plans Southeast Asia expansion, plus exploring endangered languages in East Asia, North America, and Africa. Founder Ting will speak at the Fortune Leaders Forum in Macau on Sep. 8.

Ask the AI about this article →

Context & Analysis

Votee AI's approach highlights a gap in the AI boom: while English and Mandarin dominate, hundreds of languages are underserved. The company's strategy is to take existing open-weight models and retrain them on Cantonese data, a process that costs roughly $250,000 per model, using between 500 million and 1 billion tokens. This is minuscule compared to the trillions of tokens used for English models, making it feasible for a startup.

The push for 'sovereign AI' is a broader trend, with Indonesia's Indosat building Sahabat AI for Bahasa, Singapore's AI Singapore running SEA-LION for 11 Southeast Asian languages, and South Korea staging an 'AI Squid Game' competition with a $6.8 billion budget. Votee's CEO argues that governments need to own their foundation models to avoid being 'tethered' to overseas providers, though he admits a full sovereign AI stack is 'very difficult.'

Votee's for-profit model relies on client contracts, and it has Allan Zeman, a Hong Kong tycoon, as an advisor. The company's next step is expanding into Southeast Asia, and it is in discussions with AI Singapore. The long-term vision includes protecting endangered languages, but whether a 70-billion-parameter model can make a difference remains to be seen.

FAQ

How does Votee AI train its Cantonese models?
Votee takes open-weight models from developers like Meta and Alibaba, then does additional training with Cantonese data from online scraping (including RTHK), community and university sources, and synthetic data.
Why is Cantonese AI particularly challenging?
Cantonese uses different grammar and vocabulary than Mandarin, especially in Hong Kong where speakers code-switch with English. Despite over 80 million speakers, there's a shallow pool of standardized written data, particularly for colloquial Cantonese.
Who are Votee's customers?
Governments and corporates are the first customers, including banks, universities, and government departments, funded largely through client contracts.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleFoxconn reportedly weighs major India expansion