AIToday
Large Language ModelsAI Coding AssistantsQiita 機械学習Published: Oct 7, 2026, 13:00 JST

lru_cache breaks async calls on second await, dev finds

lru_cache breaks async calls on second await, dev finds

3 Key Points

  1. What happened

    Testing a fake LLM call on CPython 3.14.8, the developer saw the first await succeed and the second raise RuntimeError: "cannot reuse already awaited coroutine", because lru_cache saves a coroutine object, not its result.

  2. Why it matters

    The cache can hit and still be unusable, so teams caching async LLM calls may find parallel duplicate requests are not actually collapsed, while sequential-only code is fine if it stores the value after the await.

  3. What to watch

    His fix shares an asyncio Task per key, and the test shows two parallel "same" calls run once (calls: 1) and a failed call retries once (calls: 2), but this is for single-event-loop batches with no cancellation — resident services need limits and TTL.

WHO IT HITSPython developers and data/ML engineers who wrap async LLM or API calls in caches, especially those running batched parallel jobs, are the ones who would hit this second-call failure.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The article walks through a debugging session whose starting point is simple: the author added a cache and only the second call broke. The reason sits in how Python's functools.lru_cache works. Calling an async def does not run the function; it returns a coroutine, and the awaited result only appears after await. lru_cache does not await anything, so it saves that coroutine object as its return value. Reusing it means awaiting a coroutine that has already been consumed. The author also checked the copy of Lib/functools.py shipped with Python and confirmed the pure-Python implementation likewise stores the function's return value.

The suggested answer is to cache an asyncio Task instead of a value. A Task, unlike a coroutine, can be awaited many times, which is what makes sharing it work for parallel callers. Two details in the test carry weight: the sleep(0) deliberate yield lets a second call arrive before the first finishes, and there is no await between checking the dictionary and registering the Task, so within one event loop a second once cannot slip in and create a duplicate. Failed Tasks are removed, but only after confirming identity, so the next call retries rather than replaying the same exception.

The author is explicit about the boundary. Storing only the result is enough for sequential callers, but in parallel both callers see an empty cache and start work anyway. Keys should include the model name, generation settings, and input, and this sharing should be avoided where a fresh answer is wanted each time. Whether this pattern holds up in real LLM batch processing, where error and cancellation behavior is messier than the test script, appears to be the open question.

FAQ
Why does lru_cache fail with async functions?
Calling an async def returns a coroutine, and lru_cache stores return values without awaiting, so it keeps an already-used coroutine. The second call then raises RuntimeError: "cannot reuse already awaited coroutine".
What is the proposed fix for parallel duplicate calls?
The developer shares one asyncio Task per key, so two parallel identical calls collapse into a single execution. His test shows calls: 1 for the parallel pair and calls: 2 after a failed call was retried.
Where should this approach not be used?
It is meant for batches of finite size in a single event loop with no mid-flight cancellation. For long-running services, limits, TTL, and cancellation handling are needed.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleClaude for Google Workspace opens as public beta with approval-first editing