AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Sep 3, 2026, 13:00 JST1 min read

Safety org proposal: pay 1,000 to read AI transcripts

Safety org proposal: pay 1,000 to read AI transcripts

Key takeaway

  • A new proposal suggests paying 1,000 people to read AI transcripts.

  • It would cost $5M monthly and catch about 15 incidents per month.

  • The goal is to catch subtle misalignment without top talent.

3 Key Points

  1. What happened

    A proposal outlines an AI safety organization that would pay about 1,000 people to read transcripts from AI models like Claude, focusing on flagged code, RL, and eval traces.

  2. Why it matters

    This approach could catch warning shots, reward hacking, and unusual behavior without absorbing top talent, and it appears such an organization does not currently exist.

  3. What to watch

    The estimated cost is $5M per month, which could process all tokens in a frontier RL run and catch about 15 serious incidents per month under bearish estimates.

Ask the AI about this article →

Context & Analysis

The proposal responds to a perceived gap: no existing organization pays people specifically to read through AI transcripts at scale. By focusing on a high-recall, low-precision monitor to flag suspicious traces, a large workforce could review the data without needing to hire away scarce AI safety experts.

The author argues that even with advanced models, humans are still needed for certain slices of detection, such as spotting obvious misalignment or subtle reward hacking. The suggested scale—1,000 people and $5M per month—would allow processing all tokens in a frontier reinforcement learning run, with an estimated yield of around 15 serious incidents per month.

The cost estimate suggests this could be a practical complement to existing safety measures, though the proposal does not detail implementation specifics or address potential privacy or logistical concerns beyond anonymization.

FAQ

What kind of transcripts would the people read?
They would read anonymized transcripts from AI models like Claude, specifically code, reinforcement learning, and evaluation traces that a very high recall, low precision monitor flags as suspicious.
Why is human reading still considered necessary?
Humans are expected to remain necessary for cases where monitors don't catch issues, especially for obvious egregious misalignment and potentially scary inner alignment or scheming behavior.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Meta's Muse Spark 1.3 matches GPT-5.6-Sol, claims #3 modelLatent Space · 5h ago
  • CABiNet vs YOLO26-sem on UAVid: Accuracy, Compute, and GPU Latencyr/MachineLearning · 5h ago
  • C++ PCN library nears backprop accuracyr/MachineLearning · 5h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMeta ends tokenmaxxing, starts testing agentic AI Hatch