AIToday
AI Safety & AlignmentLessWrong AIPublished: Aug 24, 2026, 19:00 JST2 min read

AI 'natural features' found via simple experiments

AI 'natural features' found via simple experiments

Key takeaway

  • A researcher found natural features in a small AI model using simple experiments. The method checks if the model error-corrects a structure.

  • The results surprised the researcher.

  • The work builds on prior interpretability research.

3 Key Points

  1. What happened

    A researcher shared preliminary results from experiments run with claudecode on a small LLM (gpt2-small, no Layer Norm version, courtesy of Apollo research), mostly on the layer-6 MLP. The findings, shown in an exact experimental graph, reveal natural features the model treats as important.

  2. Why it matters

    The success of these experiments, given the method's simplicity, surprised the researcher. This suggests a simple way to distinguish what a model considers 'natural' versus 'incidental' structures, potentially aiding interpretability.

  3. What to watch

    The researcher invites criticism and bug-finders. The work builds on Stefan Heimersheim and Francisco Ferreira's idea that a model's effort in error-correcting indicates naturalness, and also references an information-theoretic version related to work by Adler and Shavit, building on prior work with Kaarel Hanni, Jake Mendel, and Lawrence.

Ask the AI about this article →

Context & Analysis

The researcher's preliminary results suggest that a simple method can reveal natural features in a small language model, challenging the assumption that complex techniques are needed. This simplicity, given the success, points to a potentially accessible approach for interpretability. The method builds on a theoretical foundation from Heimersheim and Ferreira, which posits that a model's error-correction effort indicates what it deems natural. This connection between theory and empirical results strengthens the credibility of the findings. The researcher invites scrutiny, acknowledging the need for validation. The work also hints at deeper information-theoretic connections, suggesting a broader framework. While these are preliminary results, they offer a promising direction for understanding model internals with minimal resources.

FAQ

What model was used in the experiments?
The experiments used gpt2-small, a no Layer Norm version, courtesy of Apollo research.
What is the key idea behind the experiment?
The idea is that a model's effort in error-correcting a structure indicates whether it's 'natural' or 'incidental'. This comes from Stefan Heimersheim and Francisco Ferreira.
Is there a more rigorous version of this idea?
Yes, there is an information-theoretic version related to work by Adler and Shavit, building on prior work with Kaarel Hanni, Jake Mendel, and Lawrence.

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • Pentagon deploys ChatGPT MilITmedia AI+ · 35m ago
  • AI agents won't fear undeployment from misbehaviorLessWrong AI · 3h ago
  • OpenAI supports California youth AI safety billOpenAI Blog · 3h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI text watermarking: a subtle statistical pattern, not ads