AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 11, 2026, 22:00 JST

evilution kills 1,088 of 1,093 mutants in Rails AI test

evilution kills 1,088 of 1,093 mutants in Rails AI test

The author laid out design rules as a contract table in AGENTS.md, had AI write the Rails code and RSpec to match, and used the evilution mutation tester plus per-method CRAP scores to check test strength. In the final run, evilution killed 1,088 of 1,093 mutants, with 0 survivors.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The piece builds its contract table on Design by Contract, splitting rules into invariants, preconditions, and postconditions. It gives concrete evidence that invariants alone are not enough: a spec checking only that the total stays at 10 let a swap bug survive, where available_stock and reserved_stock were exchanged. Adding the postcondition that available drops by 3 and reserved rises by 3 killed that mutant, and adding boundary values just below, at, and just above the reservation quantity took the inventory model from 62 survivors out of 150 to all 150 killed.

On tooling, the author compared Ruby mutation testers directly. mutant's README restricts free use to public repositories, with $30 per developer per month or $250 per year for commercial use, so it was avoided for a private app. The archived mutest fork failed to install. evilution, under MIT, was chosen, and the author notes both evilution and the also-MIT mutineer only appeared in 2026, so changelogs and issues are worth reviewing before adoption.

The review division of labor is explicit: quality gates decide automatically whether the completion conditions are met, and humans check only whether each spec actually verifies the contract row named in its ID, whether surviving mutants are real test gaps, and whether disabled lines have genuine equivalence reasons. The author adds that stub usage in specs should be limited to two cases — failures unreachable from any real DB state, and external APIs — each with a written reason.

FAQ
Why is passing RSpec not enough to trust AI-written tests?
The author shows a User#adult? example where a spec covering ages 20 and 15 had full line coverage, but evilution reported 2 of 17 mutants survived. A mutant changing the age-18 boundary result passed because the tests never checked age 18 exactly.
Why did the author choose evilution over mutant or mutest?
The author says mutant's README limits free use to public repositories, requiring $30 per developer per month or $250 per year for commercial use. The MIT fork mutest was archived in 2020 and failed to bundle install.
What did the author find when checking evilution's equivalent-mutant judgments?
The author manually rewrote 4 mutants evilution had marked equivalent without running them; only 1 turned out genuinely equivalent. The author recommends reading the equivalent list in the JSON output even at a 100% score.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle's Playground builds games in minutes, Wired finds slop