
What happened
A developer behind the Japan-equity platform ALPHA FORGE published a Three.js 3D simulation of its pipeline, where Gemini scores news sentiment, four ML models (XGBoost, Random Forest, LSTM, Transformer) run in parallel, and Claude sets three price levels.
Why it matters
The design isolates raw news text from the main reasoning model, so Gemini passes only numbers and enum values onward — a barrier the developer frames as defence against indirect prompt injection, with Claude deciding whether a retrained model is promoted.
What to watch
Promotion hinges on all three gates — beating the incumbent in holdout tests, no worse calibration error, and non-degrading recent paper results — after which Claude issues the proposal and a human approves the swap. The full code and a 60-second video sit behind a 100 yen note article.
WHO IT HITSIndependent developers and small teams building ML-plus-LLM pipelines get a concrete reference architecture for separating untrusted text, numerical scoring and final approval — though the article is a personal verification log, not an audited system.
Summaries like this, in your inbox every morning.
The article is written from the perspective of someone running ALPHA FORGE as an individual project, and the stated motivation is that static architecture diagrams fail to show how a system behaves over time. That is why the author turned to Three.js and a real-time 3D simulation: the point is not a prettier diagram but making the sequence — data flowing in, parallel scoring, review, and nightly retraining — visible as a process rather than a snapshot.
The architecture described is layered by trust. Gemini is deliberately placed as an isolated front-line node that never hands raw news text to the downstream reasoning model; the four ML models take structured price data, technical indicators and Gemini's sentiment figures and produce probability scores; Claude then aggregates those numbers into three values — a simulation starting price, a scenario invalidation level and an expected range ceiling — with server-side validation forcing those values into the correct order. Alongside this sit two reliability mechanisms: an append-only prediction ledger in SQLite that freezes the feature snapshot at the moment of prediction and settles outcomes after 1–3 business days and 20–60 business days, and isotonic regression that corrects model confidence toward measured win rates.
How well the design holds up is likely to hinge on the promotion gate rather than the prediction step, since the author's own framing is that nightly continual learning and model replacement are the genuinely hard part. Whether the three conditions — a clear holdout win over the champion, no worsening calibration error, and non-degrading recent paper results — prove sufficient to keep a drifting model out is the open question, and the article presents itself as a personal technical log rather than investment advice or a guarantee of future performance.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
On The 8:30, Allianz chief economic adviser and Wharton School professor Mohamed El-Erian said AI's 'labor enh…

NVIDIA posted $89.02 billion in Data Center revenue for the second quarter of fiscal 2027, reported August 26…

Deepmind Institute researchers proposed "Artificial Symbiotic Intelligence," calling for coordination of netwo…

Moongazer RL, released as foxn2000/moongazer-rl, ran a fully asynchronous RL loop — rollout generation and mod…

Of 475 published Opus 5.5 videos, about 60% are motion graphics; Canvas was the most common rendering route at…

Developer ttokunaga reported September 2026 usage of about 41.574 billion Codex tokens and roughly 2.098 billi…
