A Hong Kong startup called Votee AI has built a functional Cantonese large language model. It cost roughly $250,000 to train, using between 500 million and 1 billion tokens.
By comparison, English-language models use trillions of tokens and cost hundreds of millions of dollars. GPT-4's training cost was estimated at over $100 million. GPT-5 pushed that to an estimated $1 billion or more. Votee built a commercially deployed model for a fraction of one percent of that.
The implications are significant. A 70-billion-parameter model can be trained for $250,000 to serve 85 million Cantonese speakers. That changes the economics of AI for low-resource languages entirely. The barriers to entry are not technological. They are attentional. Big Tech is spending billions on models that serve a fraction of the world's population, while ignoring markets that could be served for pocket change.
The Cantonese Gap
Cantonese is spoken by more than 80 million people globally—roughly equal to the number of Korean speakers, and more than the populations speaking Italian or Thai. Yet it is a textbook "low-resource language" for AI training. It has a much smaller pool of standardized digital text than Mandarin or English, and its grammar, vocabulary, and syntax are distinct from Mandarin. In Hong Kong, speakers frequently switch between English and Cantonese words within the same sentence.
Despite its scale, Cantonese hasn't produced anywhere near the volume of standardized written text that Mandarin or English have. Leading AI models can handle everyday Cantonese at a reasonable level, but they routinely fail on cultural and local knowledge.
As Pak-Sun Ting, CEO of Votee AI, put it: "The whole AI revolution is in English and Mandarin." "There's only a very small fraction that represents other languages." When AI fails to cover Cantonese properly, he said, "AI is essentially useless."
The Cost Structure That Changes Everything
Votee's approach is methodical and inexpensive. The company takes open-weight models from developers like Meta and Alibaba and retrains them on Cantonese data. It scrapes data from public sources, including Radio Television Hong Kong, works with universities, and uses synthetic data to expand its corpus. Together, these efforts grew the Cantonese dataset from 100 million tokens to more than 500 million.
The resulting models have around 70 billion parameters—significantly smaller than frontier systems, but capable enough to understand and reason in Cantonese. Votee sells the result to banks, universities, and government departments.
The company also developed HKCanto-Eval, a set of benchmarks to measure Cantonese AI performance, in collaboration with Kyushu University and the Education University of Hong Kong. This is not a research project. It is a commercial product serving real customers in high-stakes environments.
Ting notes that for banks, corporations, and government agencies, "you need systems that are 85%, 90%, even 95% accurate." Below that threshold, models can misunderstand inputs or generate completely irrelevant outputs—a more serious problem than the common "hallucinations" in English systems.
The Big Tech Blind Spot
The global language gap is enormous. Ting points out that Southeast Asia alone has roughly 2,000 different languages. Africa has about 1,000. "How do you bridge those gaps? That's why we see this not just as an economic opportunity, but as a significant social mission," he said.
Big Tech's neglect is not a technical limitation. It is an economic choice. The companies leading the field—OpenAI, Anthropic, DeepSeek, Moonshot AI, Z.ai, and others—are almost all headquartered in the United States or mainland China. They optimize for the largest possible markets. English and Mandarin represent billions of potential users.
Cantonese, with 80 million speakers, does not.
But the cost equation suggests this logic may be flawed. If a functional Cantonese model costs $250,000, then serving any language with a few million speakers becomes economically viable. The fixed cost of training a model for a low-resource language is no longer prohibitive. The addressable market just expanded by orders of magnitude.
The Chinese Market That Doesn't Speak Mandarin
The Cantonese opportunity is just one piece of a much larger puzzle.
China's linguistic diversity is enormous. Estimates suggest that over 200 million Chinese citizens do not speak Mandarin as their first language.
The major dialect groups include:
- Wu (Shanghainese): Approximately 80 to 90 million speakers
- Min (including Hokkien/Taiwanese): Approximately 80 million speakers
- Yue (Cantonese): Approximately 85 to 100 million speakers
- Hakka: Tens of millions of speakers
- Xiang, Gan, and other dialects: Tens of millions more
For each of these dialect groups, the current state of AI is poor. A 2026 industry analysis found that leading general-purpose models struggle significantly with dialect recognition, particularly for tone-complex dialects like Minnan and Hakka.
Chinese AI companies are beginning to respond. Z.ai has open-sourced GLM-ASR-Nano-2512, a 1.5-billion-parameter speech recognition model specifically optimized for dialects including Cantonese, Sichuanese, and Minnan. In multiple benchmarks, it outperforms OpenAI's Whisper V3 on Mandarin and Cantonese recognition tasks.
But the scale of the opportunity remains largely unaddressed. A startup called Hai Xia Ji Yin has built a 70-billion-parameter MoE dialect model, demonstrating that the technology exists to serve these populations. The question is whether the market will catch up.
The Sovereignty Play
There is another dimension to this story that Western readers often miss: data sovereignty.
Ting frames Votee's work as part of a broader "sovereign AI" push. In Asia, governments and enterprises are increasingly seeking control over their own data and infrastructure. "You definitely want to have your own large language model within the company," Ting said. "That way, you don't have to source data from outside, and you can also control the cost of using your own model."
For Cantonese speakers, this is not just about economics. It is about cultural preservation. Votee's Cantonese LLM "documents Hong Kong's culture—from cha chaan teng orders to boardroom Cantonese and everything in between." The language is "forever digitised."
The same logic applies to every other dialect and low-resource language.
If the only AI models available are trained on English and Mandarin data, then the world's linguistic diversity is being systematically narrowed. The $250,000 model is not just a technical achievement. It is a preservation effort.
What This Means
The Votee case suggests that Big Tech's approach to AI language coverage is fundamentally misaligned with the economics of the problem.
The companies spending billions on frontier models are optimizing for scale. But the marginal cost of covering a new language is far lower than anyone assumed. If a 70-billion-parameter model can be trained for $250,000, then the barrier to entry for any language with a few million speakers is negligible.
The global low-resource language AI market is already substantial.
Industry estimates put the global low-resource speech recognition market at approximately $4.7 billion in 2026, up from $3.7 billion in 2025.
Ting's company is now in talks to expand into Southeast Asia through AI Singapore, and plans to replicate its foundry model in Africa and other regions. The model is replicable. The economics are proven. The only missing ingredient is attention.
The next time someone tells you that AI requires billion-dollar training runs, ask them about Votee. Ask them about the $250,000 model that serves 85 million Cantonese speakers. Ask them why Big Tech is spending billions on languages that are already covered, while ignoring billions of speakers who are still waiting.
English data is a bubble. Low-resource languages are the blue ocean. And the barrier to entry is lower than anyone imagined.
Sources: Cryptonomist (August 27, 2026); Yahoo Tech (August 27, 2026); The AI Innovator (May 1, 2026); SCMP (July 20, 2026); EET China (August 18, 2025); Tencent Cloud (July 24, 2026); Low-Resource Language AI Development White Paper (2026 Q2); Global Low-Resource Speech Recognition Industry Deep Trend Report (August 2026).
Disclaimer
The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.
Limitations
This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.
Forecasts from third-party analysts can change with market conditions.
Cost and pricing examples are point-in-time estimates; actual rates vary.
Country and company comparisons rely on public reporting, not operational data.
This sector moves fast; timelines and deal terms may be updated later.
Company deals and regulatory rulings may evolve; verify current status.
AI infrastructure is changing quickly; claims can become outdated soon.
Sources
- Cryptonomist (August 27, 2026)
- Yahoo Tech (August 27, 2026)
- The AI Innovator (May 1, 2026)
- SCMP (July 20, 2026)
- EET China (August 18, 2025)
- Tencent Cloud (July 24, 2026)
- Low-Resource Language AI Development White Paper (2026 Q2)
- Global Low-Resource Speech Recognition Industry Deep Trend Report (August 2026).
The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.
Limitations: This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.; Forecasts from third-party analysts can change with market conditions.; Cost and pricing examples are point-in-time estimates; actual rates vary.; Country and company comparisons rely on public reporting, not operational data.; This sector moves fast; timelines and deal terms may be updated later.; Company deals and regulatory rulings may evolve; verify current status.; AI infrastructure is changing quickly; claims can become outdated soon.