AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 3, 2026, 22:00 JST

AIエージェント pipeline: the real bottleneck is queue time

AIエージェント pipeline: the real bottleneck is queue time

3 Key Points

  1. What happened

    Operating an AIエージェント pipeline from GitHub Issue to pull request, the author found waiting time splits into four kinds, and adding more agents did not shorten lead time.

  2. Why it matters

    The slowdown looks like a queueing problem, not a model-speed problem, so tuning agent count is likely the wrong fix.

  3. What to watch

    The analysis is from one environment, and the caveat is that GitHub timestamps cannot separate waiting from work, so the four-way split depends on adding external timing records.

WHO IT HITSEngineering teams running agent-based coding pipelines get a concrete lesson: measure dispatch and queue time before buying more agents, since idle workers keep waiting while the dispatcher is occupied. The measurement prescription is aimed at whoever operates the pipeline's scheduling layer.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The account begins with a simple goal: figure out whether an AI-agent pipeline that opens pull requests from GitHub Issues actually got faster. The first attempt, summing time from Issue creation to close, failed because that span mixes detection, assignment, implementation, CI, review, merge, and close. The author then tried to split waiting from implementation using the first commit's timestamp, but the measured implementation time came out to a median of one minute, because this pipeline commits once at the end and immediately opens the pull request.

Treating the post-PR period as pure waiting was also wrong. In many pull requests the last commit came after the PR was opened, by gaps ranging from tens of minutes to three and a half hours, because review feedback generates new commits. The author concludes that only external timing records, capturing dispatch, implementation completion, and handoff, can separate the intervals, and that the first fix should be adding measurement, not improving the scheduler.

The author's four-way split of waiting time and the proposed response, a dispatcher that only assigns and a worker released at handoff, point to the same reading: the constraint appears to sit in flow management rather than model performance, though this reflects a single operating environment and the method's usefulness hinges on whether the added timing points are actually instrumented.

FAQ
Why can't you measure this from GitHub timestamps alone?
Because commits are made only at the end of implementation, not at the start, and review-response work gets mixed into the post-PR period. GitHub's timestamps lack information that separates waiting from actual implementation.
Does adding more agents shorten the wait?
No. If the dispatcher is busy, more agents do not shrink the dispatch wait. Throughput rises, but the per-issue experience does not change.
What are the four kinds of waiting time?
Detection wait from the polling interval, dispatch wait while the assigning side is occupied, dependency wait from running related Issues at once, and post-PR wait for CI, review, and merge.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleKaran Joshi extracts Muse files on everyone you know