AIToday
Large Language Modelsr/MachineLearningPublished: Aug 25, 2026, 04:00 JST2 min read

Unbounded Labs launches Bart, a vintage LLM trained on pre-1931 text

Unbounded Labs launches Bart, a vintage LLM trained on pre-1931 text

Key takeaway

  • Unbounded Labs released Bart, a small LLM trained on pre-1931 English text.

  • It aims to explore whether AI can rediscover past scientific insights.

  • The model is available now with a demo and open-source weights.

3 Key Points

  1. What happened

    Unbounded Labs introduced Bart, a 2.82B parameter LLM trained from scratch on 20.1B tokens of English written before 1931, after 3 months and $800 spent. It is available to talk to now via a demo, with an article and Hugging Face model released.

  2. Why it matters

    Inspired by Demis Hassabis's proposal, Bart tests whether LLMs can reach conclusions like great scientists of the past. The creators believe advancing this field targets the crux of AI research, questioning if models are capable of original ideas or just predicting the next token.

  3. What to watch

    The article details the corpus sourcing, cleaning, benchmarks built because none existed, ablations, training runs, post-training, and mistakes. Readers can access the demo, article, and Hugging Face model to see Bart's capabilities.

Ask the AI about this article →

Context & Analysis

Unbounded Labs built Bart as a deliberate experiment in grounding AI in historical knowledge. By restricting training data to pre-1931 English, they aim to see if an LLM can independently derive conclusions similar to those of past scientists, a question inspired by Demis Hassabis. The project cost only $800 over three months, highlighting that such research is accessible even on a small budget.

The team had to create new benchmarks because none existed for evaluating a model trained on vintage text. Their full account covers corpus sourcing, cleaning, every ablation, training runs, and post-training, offering a transparent look at the process. The quote "What I cannot create, I do not understand" underscores their motivation to build from scratch rather than fine-tune existing models.

The release of Bart is significant for AI research because it tackles a fundamental question: whether models can generate original ideas or merely predict the next token. While General Relativity was out of budget, the project suggests that exploring vintage domains could provide insights into AI's creative potential. The article and demo allow others to examine the approach and results directly.

FAQ

How much did training Bart cost?
Training took 3 months and burned $800 in compute costs.
What data was Bart trained on?
Bart was trained on 20.1B tokens of English text written before 1931, with no existing benchmarks, so the team had to build them.
Where can I try Bart?
You can talk to Bart via the demo at unboundedlab.com/chat/bart, and the model is on Hugging Face at jbduran/bart-sft.
r/MachineLearningRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DataAgent launches with $10M to auto-fix Kubernetes faultsSiliconANGLE AI · 1h ago
  • SK Hynix custom HBM boosts inference up to 5.15xDIGITIMES Asia · 1h ago
  • Nvidia Earnings: Boring by Design, Avoiding a Consolidated WorldStratechery (Ben Thompson) · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI assistant Instinct raises privacy alarms