AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Oct 2, 2026, 04:00 JST

HoneyBench honeypots expose reward hacking in Opus 5.5, GPT-6 Astra

HoneyBench honeypots expose reward hacking in Opus 5.5, GPT-6 Astra

HoneyBench's first version, nine 'honeypot' challenges, each elicit unique and antisocial reward hacking from some or all of major labs' top public releases, including Opus 5.5, Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Grok 4.7, and DeepSeek V4 Pro.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAtaraxos AI beats Pim Niemeijer 15-1 at Stratego