Alibaba Cloud founder Wang Haifeng has challenged the prevailing industry approach to AI development, arguing that next-generation foundation models need scientific data rather than continued scaling of model size and computing resources. For three years, the field has pursued bigger models with larger parameter counts and longer context windows, but Wang's intervention suggests the real constraint for future progress may lie in data quality and composition, not raw computational power.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Alibaba Cloud founder Wang Haifeng has argued that the next generation of large language models (LLMs—AI systems that understand and generate text) require scientific data as a foundation, breaking from the industry's recent focus on model size and computing resources.
Why it matters
For the past three years, the field has pursued bigger models with more parameters and extended context windows, but Wang's position suggests the bottleneck for progress may have shifted—away from raw scale toward the quality and nature of training data. This could redirect how companies approach model development and where they invest resources.
What to watch
The debate marks a turning point in how the industry measures progress: whether continued advancement depends primarily on computing power, or whether sourcing and curating scientific-quality data will become the decisive factor in building more capable systems.
Over the past three years, large language model development has followed a consistent trajectory centered on expansion. Model sizes have grown, parameter counts have climbed steadily, reasoning capabilities have improved, AI agents have matured, and context windows—the span of text a model can consider at once—have repeatedly broken previous records. This trajectory has shaped both the industry's technical direction and its public conversation, with the central debate focused on model capability, the scale of computing resources required, and how close the field stands to artificial general intelligence.
Against this backdrop, Alibaba Cloud founder Wang Haifeng has introduced a divergent perspective. He contends that the next generation of foundation models will require scientific data as their core foundation, rather than continued emphasis on scale. This argument represents a significant reorientation of the industry's priorities. Where much of the field has bet on the benefits of larger models trained on ever-greater volumes of data, Wang's position suggests that the composition, quality, and scientific grounding of that data may be more consequential than raw quantity or compute investment. His intervention opens a fundamental question about what truly constrains progress in AI development and where resources should be concentrated in the years ahead.
The AI industry's development path over the past three years has been characterized by a clear escalation strategy: bigger models, higher parameter counts, more capable reasoning, and record-breaking context windows. This direction was undergirded by a shared assumption that scale—whether in model parameters or compute—was the primary lever for progress. The central industry conversation reflected this belief, with discussions revolving around model capability relative to computing investment and the distance to AGI.
Wang Haifeng's intervention marks a significant counterpoint to this consensus. By asserting that scientific data—not scale—should be the foundation of next-generation models, he points to a potential shift in the field's bottleneck. Rather than suggesting that larger models or more computing power will unlock the next level of performance, his position implies that the quality, relevance, and scientific rigor of training data may be the limiting factor the industry has underestimated. This reframing does not negate the value of scale but repositions where the real constraint lies.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion




Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime