
What happened
Testing a fake LLM call on CPython 3.14.8, the developer saw the first await succeed and the second raise RuntimeError: "cannot reuse already awaited coroutine", because lru_cache saves a coroutine object, not its result.
Why it matters
The cache can hit and still be unusable, so teams caching async LLM calls may find parallel duplicate requests are not actually collapsed, while sequential-only code is fine if it stores the value after the await.
What to watch
His fix shares an asyncio Task per key, and the test shows two parallel "same" calls run once (calls: 1) and a failed call retries once (calls: 2), but this is for single-event-loop batches with no cancellation — resident services need limits and TTL.
WHO IT HITSPython developers and data/ML engineers who wrap async LLM or API calls in caches, especially those running batched parallel jobs, are the ones who would hit this second-call failure.
Summaries like this, in your inbox every morning.
The article walks through a debugging session whose starting point is simple: the author added a cache and only the second call broke. The reason sits in how Python's functools.lru_cache works. Calling an async def does not run the function; it returns a coroutine, and the awaited result only appears after await. lru_cache does not await anything, so it saves that coroutine object as its return value. Reusing it means awaiting a coroutine that has already been consumed. The author also checked the copy of Lib/functools.py shipped with Python and confirmed the pure-Python implementation likewise stores the function's return value.
The suggested answer is to cache an asyncio Task instead of a value. A Task, unlike a coroutine, can be awaited many times, which is what makes sharing it work for parallel callers. Two details in the test carry weight: the sleep(0) deliberate yield lets a second call arrive before the first finishes, and there is no await between checking the dictionary and registering the Task, so within one event loop a second once cannot slip in and create a duplicate. Failed Tasks are removed, but only after confirming identity, so the next call retries rather than replaying the same exception.
The author is explicit about the boundary. Storing only the result is enough for sequential callers, but in parallel both callers see an empty cache and start work anyway. Keys should include the model name, generation settings, and input, and this sharing should be avoided where a fresh answer is wanted each time. Whether this pattern holds up in real LLM batch processing, where error and cancellation behavior is messier than the test script, appears to be the open question.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Meta, Walmart, Stripe and Sierra Technologies are creating the Personal Agent Protocol, introduced by Sierra c…
Mistral AI opened a public preview of Mistral Large 4, its 1.05 trillion-parameter MoE model nicknamed "Le Cho…

Reflection AI started early access to Beam, its first open-weights model, with 501 billion total and 23 billio…

Google began rolling out "Simple Guide" in Gemini Live on Android, letting users share camera or screen views…

Anthropic launched Claude for Google Workspace as a public beta for paid Claude plans, adding Claude to Google…

Anthropic released "Claude for Google Workspace" as a public beta, reachable from each tool's "Extensions(拡張機能…
