
What happened
In Oracle's 80-turn evaluation, Oracle AI Agent Memory held input near 1,300 tokens per request while flat history grew past 13,900 tokens. It won 48 evaluated turns versus 13 for flat history, with 19 ties.
Why it matters
An agent that keeps session state without resending full history is likely to cost less per request and stay within context limits longer than one relying on flat history.
What to watch
The evaluation is Oracle's own documented test, so results could differ under other workloads or memory policies. Watch how the managed-memory agent performs on external benchmarks rather than internal evaluations.
WHO IT HITSEnterprise AI teams building long-running support or workflow agents face a choice: resend every past message, or adopt a managed-memory layer like Oracle's. The token gap affects both API bills and how many turns an agent can handle before context runs out.
Summaries like this, in your inbox every morning.
The article starts from a simple fact: an LLM is stateless. It only appears to remember because the application resends previous messages with each request. Start a new session without that stored history, and every preference, decision, and prior outcome disappears. The proposed fix is a memory layer split into two scopes — short-term working state during a session, and long-term memory that persists across sessions. Long-term memory is further broken into semantic, episodic, and procedural types, each needing different write and retrieval policies.
Oracle's documented 80-turn evaluation is the concrete evidence offered. The managed-memory agent held input near 1,300 tokens per request while flat history grew past 13,900 tokens by the final turn. It also won 48 evaluated turns, against 13 for flat history, with 19 ties. The mechanism described is not weight updates; the surrounding system adapts by storing, updating, and retrieving state, so the model itself is unchanged.
For RAG pipelines, the article positions Jev as an evaluation step between retrieval and generation. It scores candidates so application code, not the prompt, decides which passages reach the context window. Whether this matters most for cost, for answer reliability, or for auditability depends on the workload; the body frames the value as making relevance a typed probability that teams can log, test, and threshold rather than an implicit assumption.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Rosenblatt raised its Amazon price target to $360 from $335, kept a Buy rating, and called concerns that AI sh…

Deepseek released open-source programming tools for Huawei's Ascend chips, including libraries for computation…

LinkedIn CEO Dan Shapero told the Wall Street Journal that job seekers have sent out 30% more applications tha…

Confluent's 2026 Data Streaming Report found just 17% of Japanese firms run agentic AI in production, the lowe…

Restate raised a $20 million Series A led by Singular, with Redpoint Ventures and Capital One Ventures, after…

Cerebras Systems CEO and co-founder Andrew Feldman will take the Disrupt Stage at TechCrunch Disrupt 2026 for…
