AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 11, 2026, 10:00 JST

changelog-bot's Jev trick: pick the Why, don't write it

changelog-bot's Jev trick: pick the Why, don't write it

@nyaomaru's changelog-bot avoids asking an LLM to write a reason for a code change, and instead has Jev score candidate snippets pulled from the pull request, accepting only text the author already wrote.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The problem @nyaomaru ran into is that an LLM handed a pull request and asked "Why was this change made?" will happily summarize, paraphrase, combine sentences, and infer reasons that were never stated. For a changelog, that is the wrong kind of help. The pivot was to reclassify the task: not "How can AI write a good reason?" but "How can we find evidence that the reason already exists?" That reframing splits the work between deterministic code, which locates and normalizes candidate evidence, Jev, which evaluates it, and a renderer, which formats it.

This fits the project's broader v1 direction, which the author calls "deterministic-first": a complete changelog should be produced even if the AI is disabled or fails. AI enriches the changelog rather than owning it. The WHY engine is accordingly an optional enrichment step and does not generate the final Markdown itself, which stays with a single deterministic renderer. Phases 1 through 3 are already implemented, with the boundary between structured rendering and enrichment slated for further work.

The threshold question is being answered with measurement rather than intuition. The initial evaluation harness had 14 cases, mixing positives (where one candidate is clearly the reason) with negatives (where the correct answer is to select nothing), and after @johnnylemonny joined as a contributor, PR #212 expanded it to about 50 cases covering explicit rationale, implementation-only text, template noise, and multilingual pull requests. The author set one rule for the ground truth: it must be decided by humans independently of the current Jev threshold, so that labels are not simply whatever the model already scored highly.

FAQ
How does changelog-bot decide which PR sentence to use as the "Why"?
It first preprocesses the pull request to find candidate snippets using signals like section names such as "why" or "reason" and phrases like "because" or "to prevent". Jev then asks two questions per candidate — whether it explicitly states a reason, and whether that reason applies to the specific change — and only candidates passing both probability thresholds are accepted.
Does the AI rewrite the reason it finds?
No. The final "Why" is the original sentence from the pull request description, with Jev deciding which candidate to select but the PR author having written the actual text that appears in the changelog.
What happens if a pull request has no clear reason?
The tool adds no "Why" at all. The author considers "There is no reliable WHY in this PR" a valid and important outcome, preferring it over forcing a plausible-sounding reason.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleByte Transformers beat subword models as scale grows