
A LessWrong essay, grown from conversations with Danaja Rutar, Paul Colognese and Eric Michaud, proposes "self-inoculation" — a virtuous form of gradient hacking — and demonstrates a possible circuit using a toy model.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Recursive CEO Richard Socher said in an X post that "There is NO realistic scenario where AI wipes out all of…

At the Yale CEO Caucus, 93% of attendees disagreed with Trump's Truth Social post calling AI warnings a hoax…

Anthropic CEO Dario Amodei urged the US and China to cooperate on slowing AI in his Saturday blog post "We Mus…

Last week, artificial intelligence rose to prominence as a crucial global issue after prominent officials and…

The heads of leading U.S

A LessWrong author posted hurried guidance for "technical profiles" — analytic, sometimes shy people, possibly…
