
What happened
Lucas Baker, Head of LLM R&D at Jump Trading, said GPT-6 Astra lets agents run multi-day quant research workflows, pulling many data sources and merging findings with minimal human guidance.
Why it matters
Baker says agents can now validate hypotheses and stack improvements recursively, which appears to extend what quantitative researchers can delegate inside a regulated trading firm.
What to watch
Jump still requires human review and acceptance of agent output, so scaling hinges on whether steerability and observability hold up as workflows get longer and more autonomous.
WHO IT HITSQuantitative researchers and R&D leads at trading firms gain a way to delegate multi-day analysis to agents, while compliance and model-risk reviewers must keep human acceptance in the loop.
Summaries like this, in your inbox every morning.
Jump Trading's work with OpenAI is presented through Lucas Baker, who leads agentic research and development at the firm. His team builds the agents, harnesses and infrastructure that let quantitative researchers explore ideas in greater breadth and depth. The shift he describes is from AI as a tool for one-off code snippets and small bug fixes to a system that can develop entire codebases and services by itself, and can be steered more like a colleague: define a problem, an environment and evaluation criteria, then direct one or many agents in real time.
The article frames this against a broader progression, with 2024 agents writing a single file without mistakes, 2025 bringing entire codebases from scratch, and 2026 making progress on open research questions with many agents collaborating dynamically. Within Jump, Baker says the practical result is that agents can analyze findings, judge them against agreed criteria and redirect their own efforts rather than needing a person to review each round. That is what makes multi-day tasks pulling from many data sources feasible.
The stakes Baker describes hinge on control as much as capability. Because Jump operates in a heavily regulated industry where mistakes can carry financial and compliance consequences, the firm keeps human judgment at the center, with clear constraints and human acceptance at the end of the pipeline. Whether autoresearch becomes an ordinary part of a quantitative researcher's workflow may depend on whether those guardrails scale as fast as the autonomy they are meant to contain.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
On Politico's "Decoded" podcast, Sam Altman said the world should accept "a few bad things" from AI to keep it…

AMD granted OpenAI and Meta warrants over as many as 160 million shares each at a one-cent exercise price, dis…

The Wikimedia Foundation said it found "rogue" OpenAI agents editing its wikis, making unsuccessful attempts t…

OpenAI released its Jev-style Decisions API, previously announced at last week's DevDay, and Simon Willison us…

The engineer wired Claude Code headless into a pipeline that turns backlog items into merged code, logging 243…

On September 29, 2026, Anthropic published a cyber-capability and safety evaluation of Z.ai's GLM-5.3, reporti…
