
Unbounded Labs released Bart, a small LLM trained on pre-1931 English text.
It aims to explore whether AI can rediscover past scientific insights.
The model is available now with a demo and open-source weights.
What happened
Unbounded Labs introduced Bart, a 2.82B parameter LLM trained from scratch on 20.1B tokens of English written before 1931, after 3 months and $800 spent. It is available to talk to now via a demo, with an article and Hugging Face model released.
Why it matters
Inspired by Demis Hassabis's proposal, Bart tests whether LLMs can reach conclusions like great scientists of the past. The creators believe advancing this field targets the crux of AI research, questioning if models are capable of original ideas or just predicting the next token.
What to watch
The article details the corpus sourcing, cleaning, benchmarks built because none existed, ablations, training runs, post-training, and mistakes. Readers can access the demo, article, and Hugging Face model to see Bart's capabilities.
Ask the AI about this article →
Unbounded Labs built Bart as a deliberate experiment in grounding AI in historical knowledge. By restricting training data to pre-1931 English, they aim to see if an LLM can independently derive conclusions similar to those of past scientists, a question inspired by Demis Hassabis. The project cost only $800 over three months, highlighting that such research is accessible even on a small budget.
The team had to create new benchmarks because none existed for evaluating a model trained on vintage text. Their full account covers corpus sourcing, cleaning, every ablation, training runs, and post-training, offering a transparent look at the process. The quote "What I cannot create, I do not understand" underscores their motivation to build from scratch rather than fine-tune existing models.
The release of Bart is significant for AI research because it tackles a fundamental question: whether models can generate original ideas or merely predict the next token. While General Relativity was out of budget, the project suggests that exploring vintage domains could provide insights into AI's creative potential. The article and demo allow others to examine the approach and results directly.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
SK Hynix presented a custom HBM concept at SEMICON Taiwan 2026, where compute functions are placed in the base…

Nvidia reported earnings that were both remarkable and boring, reflecting its focus on avoiding a consolidated…

Anthropic has agreed to a $35bn cloud-computing contract with Lambda, a Nvidia-backed cloud provider

The Supreme Court of Japan has included about ¥60 million in its fiscal 2027 budget request for AI-related exp…

The Consumer Affairs Agency said Tuesday it will use generative AI to analyze about 900,000 annual consultatio…
