Researchers benchmarked 11 frontier LLMs across 15 distributions using two protocols: Batch Generation (N=1000 samples in one response) and Independent Requests (N=1000 stateless calls). Batch generation achieved a 7% median pass rate; 10 of 11 models passed none of the distributions in independent requests.
Sampling fidelity degrades monotonically with distributional complexity and worsens as the sampling horizon N increases. The sharp protocol asymmetry suggests current LLMs lack a functional internal sampler (a mechanism to generate random samples matching specified probability distributions).
Downstream failures introduce systematic biases: models fail to enforce uniform answer-position constraints in Multiple Choice Question generation and systematically violate demographic targets in attribute-constrained text-to-image prompt synthesis, indicating a need for external tools when statistical guarantees are required.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Interactive Brokers has begun connecting its platform with AI tools including ChatGPT, Claude, and Grok, and o…

Google has reportedly approached major studios such as Disney, Warner Bros

Google DeepMind introduced Gemini 3.8 Flash and Gemini 3.8 Flash Cyber

IBM released a survey showing AI is already used weekly in 76% of middle school and 73% of high school classro…

Qualcomm Technologies and HUMAIN launched Horizon Ultra, a next-generation AI PC, at LEAP 2026

Dentsu, Dentsu Digital, and SB Intuitions announced that they have built a dataset of over 110,000 ad copy eva…
