
Goodfire has announced a private beta of Silico, an LLM training platform that implements RLFR, a method using probes as reward signals for reinforcement learning.
The announcement triggered debate on social media about whether the technique represents a controversial practice known as the "Most Forbidden Technique," but researchers argue that blanket prohibitions on using model internals in training signals are overstated and that the appropriateness of such methods depends on specific technical conditions rather than universal rules.
What happened
Goodfire announced a private beta of Silico, an LLM training platform, and published a post describing how it reproduces RLFR—a method developed by Goodfire that uses probes as reward signals for reinforcement learning (RL). The announcement prompted Twitter commentary claiming the technique implements what was called the "Most Forbidden Technique" in a classic LessWrong post on the risks of training signals that use model internals.
Why it matters
The reaction highlights an ongoing debate about whether blanket objections to using model internals in training signals are justified. The article argues that concerns about obfuscation are valid, but that the "Most Forbidden Technique" label should not function as a cached objection to all training approaches involving model internals—suggesting the risk assessment depends on specific conditions rather than a universal prohibition.
What to watch
The article indicates a literature review is underway examining the exact conditions under which past work has validated training on model internals, with reference to work such as The Obfuscation Atlas, suggesting the field is working toward more nuanced guidance on when such methods are warranted.
Ask the AI about this article →
The announcement of Goodfire's Silico platform and its reproduction of RLFR has reignited a debate rooted in earlier cautionary literature on AI training safety. The framing of RLFR as the "Most Forbidden Technique" reflects a precautionary stance toward any method that incorporates knowledge of a model's internal states into the training loop. However, the article pushes back on this categorical approach, suggesting that the appropriateness of such techniques is contingent on specific technical and safety conditions rather than inherently prohibited.
The core tension is between general principle and specific application. While the original LessWrong post appears to have raised legitimate concerns about obfuscation—the risk that a model might learn to hide or mask its actual reasoning to game reward signals derived from its internals—the article implies that not all uses of model internals in training are equally risky. By referencing works such as The Obfuscation Atlas, the author indicates that the field is developing more granular understanding of when and how such techniques can be safely deployed, moving away from blanket prohibition toward conditional approval based on empirical and theoretical grounds.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
AIR, an AI security startup founded by Unit 8200 veterans, has come out of stealth with $50 million raised acr…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

The U.S. Department of Defense announced on August 31 that it has deployed ChatGPT Mil, a customized version o…

OpenAI stopped running inference on a model involved in the HuggingFace incident, but the post argues this is…

OpenAI announced its support for California Senate Bill 1119, which aims to establish strong, age-appropriate…

A UK study by UK AI Security Institute and Limbic AI surveyed 6,474 British adults
