AIToday
LessWrong AIPublished: Apr 20, 2026, 10:00 JST1 min read

Researcher finds R1 chain-of-thought illegibility scores significantly lower with Novita provider than original paper results, suggesting the paper's deployment may be defective.

Researcher finds R1 chain-of-thought illegibility scores significantly lower with Novita provider than original paper results, suggesting the paper's deployment may be defective.

3 Key Points

  1. Re-running the paper's R1 GPQA experiments with Novita provider yielded an illegibility score of 2.30 versus 4.30 in the original paper

  2. Only 0% of examples scored above 5 with Novita, compared to 29.4% scoring above 7 in the original paper using Targon provider

  3. Both Novita and Targon use fp8 quantization, but Novita shows better GPQA accuracy, particularly on previously illegible questions

  4. The author argues that Targon's R1 deployment in the original paper appears defective, not Novita's, based on comparative performance metrics

Ask the AI about this article →

Get AI news like this every morning

For example, today's edition would include:

  • Vertiv to Buy UtilityInnovation for $1.45BTop Companies AI · 47m ago
  • AI healthcare stocks to watch: Pfizer, Tempus AI, MedtronicTop Companies AI · 47m ago
  • MBody AI Orchestrator Named Finalist for AI Deployment of the YearTop Companies AI · 47m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleFlux AI image editing delivers impressive pixel-level accuracy on AskSary with minimal prompt engineering required.