AIToday
Large Language ModelsarXiv cs.AIPublished: Apr 27, 2026, 13:00 JST1 min read

Study finds 11 frontier LLMs fail to faithfully sample from probability distributions, with batch generation achieving only 7% median pass rate and independent requests collapsing almost entirely.

3 Key Points

  1. Researchers benchmarked 11 frontier LLMs across 15 distributions using two protocols: Batch Generation (N=1000 samples in one response) and Independent Requests (N=1000 stateless calls). Batch generation achieved a 7% median pass rate; 10 of 11 models passed none of the distributions in independent requests.

  2. Sampling fidelity degrades monotonically with distributional complexity and worsens as the sampling horizon N increases. The sharp protocol asymmetry suggests current LLMs lack a functional internal sampler (a mechanism to generate random samples matching specified probability distributions).

  3. Downstream failures introduce systematic biases: models fail to enforce uniform answer-position constraints in Multiple Choice Question generation and systematically violate demographic targets in attribute-constrained text-to-image prompt synthesis, indicating a need for external tools when statistical guarantees are required.

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Interactive Brokers bets on AI-assisted investingTop Companies AI · 2h ago
  • Google courts Hollywood with AI licensing dealsTop Companies AI · 2h ago
  • AI adoption in K-12 outruns school readiness, IBM survey findsTop Companies AI · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMicrosoft and Stellantis launch five-year partnership to build 100+ AI features for software-driven cars