AIToday
Large Language ModelsAI Stocks & MarketsAI Business & IndustryQiita 機械学習Published: Oct 3, 2026, 22:00 JST

ALPHA FORGE 3D map: Gemini, Claude and 4 ML models split the work

ALPHA FORGE 3D map: Gemini, Claude and 4 ML models split the work

3 Key Points

  1. What happened

    A developer behind the Japan-equity platform ALPHA FORGE published a Three.js 3D simulation of its pipeline, where Gemini scores news sentiment, four ML models (XGBoost, Random Forest, LSTM, Transformer) run in parallel, and Claude sets three price levels.

  2. Why it matters

    The design isolates raw news text from the main reasoning model, so Gemini passes only numbers and enum values onward — a barrier the developer frames as defence against indirect prompt injection, with Claude deciding whether a retrained model is promoted.

  3. What to watch

    Promotion hinges on all three gates — beating the incumbent in holdout tests, no worse calibration error, and non-degrading recent paper results — after which Claude issues the proposal and a human approves the swap. The full code and a 60-second video sit behind a 100 yen note article.

WHO IT HITSIndependent developers and small teams building ML-plus-LLM pipelines get a concrete reference architecture for separating untrusted text, numerical scoring and final approval — though the article is a personal verification log, not an audited system.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The article is written from the perspective of someone running ALPHA FORGE as an individual project, and the stated motivation is that static architecture diagrams fail to show how a system behaves over time. That is why the author turned to Three.js and a real-time 3D simulation: the point is not a prettier diagram but making the sequence — data flowing in, parallel scoring, review, and nightly retraining — visible as a process rather than a snapshot.

The architecture described is layered by trust. Gemini is deliberately placed as an isolated front-line node that never hands raw news text to the downstream reasoning model; the four ML models take structured price data, technical indicators and Gemini's sentiment figures and produce probability scores; Claude then aggregates those numbers into three values — a simulation starting price, a scenario invalidation level and an expected range ceiling — with server-side validation forcing those values into the correct order. Alongside this sit two reliability mechanisms: an append-only prediction ledger in SQLite that freezes the feature snapshot at the moment of prediction and settles outcomes after 1–3 business days and 20–60 business days, and isotonic regression that corrects model confidence toward measured win rates.

How well the design holds up is likely to hinge on the promotion gate rather than the prediction step, since the author's own framing is that nightly continual learning and model replacement are the genuinely hard part. Whether the three conditions — a clear holdout win over the champion, no worsening calibration error, and non-degrading recent paper results — prove sufficient to keep a drifting model out is the open question, and the article presents itself as a personal technical log rather than investment advice or a guarantee of future performance.

FAQ
Why is Gemini kept separate from Claude in this system?
Feeding raw news headlines straight into the main reasoning model (Claude) would create a risk of indirect prompt injection, so Gemini sits as an isolated outpost node. It passes only objective figures such as sentiment scores and impact levels, plus enum values, to the layers below.
What stops an AI model from being overconfident in its own predictions?
ALPHA FORGE applies Isotonic Regression as a post-hoc calibration, using past ledger results to correct output probabilities toward the true observed win rate. The example given is a model that is 90% confident while its actual win rate is 60%.
How does a new model replace the current one?
A nightly batch retrains models and detects drift via the Population Stability Index, while a challenger runs in shadow inference. Only if it clearly beats the champion in holdout validation, does not worsen calibration error and shows non-degrading recent paper results does Claude propose promotion, which a human must approve.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleFocus Flow AI (v7.0) ships unfinished to train open-source focus AI