AIToday
Large Language ModelsHugging Face BlogPublished: Apr 30, 2026, 01:00 JST1 min read

IBM releases Granite 4.1 LLM family trained on ~15T tokens with five-phase pre-training pipeline and long-context extension to 512K tokens

IBM releases Granite 4.1 LLM family trained on ~15T tokens with five-phase pre-training pipeline and long-context extension to 512K tokens

3 Key Points

  1. Granite 4.1 comprises three dense transformer models (3B, 8B, and 30B parameters) trained from scratch on approximately 15 trillion tokens using a five-stage pre-training strategy that progressively shifts from broad web data to curated, domain-specific content, with context window extended to 512K tokens in the final phase.

  2. The 8B instruct model matches or surpasses the previous Granite 4.0-H-Small (32B-A9B MoE, a mixture-of-experts architecture) despite using fewer parameters and a simpler dense architecture, refined through supervised fine-tuning on ~4.1M curated samples and reinforcement learning via on-policy GRPO with DAPO loss.

  3. All Granite 4.1 models are released under the Apache 2.0 license and available via Hugging Face Collection and GitHub Repository.

Ask the AI about this article →

Hugging Face BlogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Walmart settles opioid claims for $50MTop Companies AI · 3h ago
  • Tim Cook's legacy hinges on Apple's AI betTop Companies AI · 3h ago
  • CrowdStrike Falcon Guardian Targets AI SecurityTop Companies AI · 3h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNvidia releases Nemotron 3 Nano Omni, a 30-billion-parameter open multimodal model trained on data from competing AI labs including Qwen, OpenAI, and DeepSeek.