
Granite 4.1 comprises three dense transformer models (3B, 8B, and 30B parameters) trained from scratch on approximately 15 trillion tokens using a five-stage pre-training strategy that progressively shifts from broad web data to curated, domain-specific content, with context window extended to 512K tokens in the final phase.
The 8B instruct model matches or surpasses the previous Granite 4.0-H-Small (32B-A9B MoE, a mixture-of-experts architecture) despite using fewer parameters and a simpler dense architecture, refined through supervised fine-tuning on ~4.1M curated samples and reinforcement learning via on-policy GRPO with DAPO loss.
All Granite 4.1 models are released under the Apache 2.0 license and available via Hugging Face Collection and GitHub Repository.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Walmart settled opioid dispensing claims for $50 million

Tim Cook's legacy as Apple CEO is now tied to the company's push into artificial intelligence, according to a…

AT&T, Dell Technologies, and AMD have announced OTel 2.0, the largest and best-performing open-source model bu…

John Deere introduced JD, an AI assistant designed to help farmers manage and interpret their farm data, as re…

CrowdStrike is introducing Falcon Guardian, its flagship solution for the AI Detection and Response (AIDR) cat…

AT&T's legal department built an in-house center of expertise called Legal Edge, described as an AI-first lega…
