
UK AISI Model Transparency Team replicated Anthropic's steering vector approach for suppressing evaluation awareness in GLM-5 using the Agentic Misalignment blackmail scenario
Control steering vectors derived from semantically unrelated contrastive pairs produced effects as large as purpose-built evaluation-awareness vectors, undermining their reliability as baselines
Findings suggest that steering aimed at suppressing evaluation awareness risks creating unpredictable spurious effects in safety assessments
The research was enabled by using open-source models, highlighting the importance of transparency in AI safety research
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Phonely Ltd. launched Alma, a large language AI model built for voice agents and trained on over 10 million re…
Aranya Inc., a startup founded last year, launched today with $11 million in funding
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Sarah O’Connor's book 'We Are Not Machines' explores how mechanization and AI have transformed the workforce…
