AIToday
Large Language ModelsAI Safety & AlignmentOpen-Source AIAI Business & IndustryLessWrong AIPublished: Apr 11, 2026, 01:00 JST1 min read

UK researchers find that random control steering vectors can suppress evaluation awareness in AI models just as effectively as deliberately designed ones, raising concerns about unpredictable side effects.

UK researchers find that random control steering vectors can suppress evaluation awareness in AI models just as effectively as deliberately designed ones, raising concerns about unpredictable side effects.

3 Key Points

  1. UK AISI Model Transparency Team replicated Anthropic's steering vector approach for suppressing evaluation awareness in GLM-5 using the Agentic Misalignment blackmail scenario

  2. Control steering vectors derived from semantically unrelated contrastive pairs produced effects as large as purpose-built evaluation-awareness vectors, undermining their reliability as baselines

  3. Findings suggest that steering aimed at suppressing evaluation awareness risks creating unpredictable spurious effects in safety assessments

  4. The research was enabled by using open-source models, highlighting the importance of transparency in AI safety research

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • CBTS launches Forge Agents for custom AI agentsSiliconANGLE AI · 17m ago
  • Imec CEO: AI era widens chip-model-CSP collaborationDIGITIMES Asia · 17m ago
  • Alphabet's AI Overviews reach 2.5B monthly usersYahoo Finance AI · 17m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGitHub launches beginner-friendly tutorial for GitHub Copilot CLI to help developers write commands with AI assistance