
Re-running the paper's R1 GPQA experiments with Novita provider yielded an illegibility score of 2.30 versus 4.30 in the original paper
Only 0% of examples scored above 5 with Novita, compared to 29.4% scoring above 7 in the original paper using Targon provider
Both Novita and Targon use fp8 quantization, but Novita shows better GPQA accuracy, particularly on previously illegible questions
The author argues that Targon's R1 deployment in the original paper appears defective, not Novita's, based on comparative performance metrics
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.