
What happened
A student made granite-code:8b and granite3.2:8b write a TORCS racing AI in 13 parts, checked by Python test scripts; all 13 passed, but the assembled car could not finish one lap.
Why it matters
The tests only confirmed the code matched the student's specifications, so they froze his specification errors as correct behavior rather than catching them.
What to watch
The lesson hinges on whether the specification itself serves the goal — not just whether each part passes its test — and the article suggests placing a whole-goal test, such as whether the car can run a lap, from the start.
WHO IT HITSThis lands on developers and students who use small local LLMs (language models that run on your own PC) to generate code part by part, and on the reviewers who sign off on those parts. It warns that a passing test suite can hide a flawed specification, so teams need a separate whole-goal check.
Summaries like this, in your inbox every morning.
The article comes from a student who spent the spring of 2026 on an exchange year in the UK, taking an IBM client course that required teams to build a TORCS racing AI while using IBM Granite, IBM's LLM. The class setup put Granite in VS Code through the Continue.dev extension, with Ollama serving the model locally rather than through a cloud API, and the rules also required finishing IBM's online course and showing how Granite was used in a video and blog. The student's team decided in late March to split work by function, have Granite write each part, and have at least two people review any generated code before it was merged in.
What makes the story useful is the gap between mechanical checking and real performance. The workflow moved from manual prompting to a 435-line automated loop that called Ollama's API directly, with the model split between granite-code:8b for code and granite3.2:8b for rewriting prompts. Thirteen parts were defined and checked by per-part Python scripts, with retries, prompt rewrites after two failures, and a Plan B that split stubborn parts further after five failures. The parts passed, but the assembled car did not, because the tests only confirmed fidelity to the specifications, and the specifications contained the actual mistakes.
The article also draws a distinction between being required to use Granite and being allowed to use IBM Bob, IBM's coding agent that became generally available on April 28, 2026. Bob was not used for the racing AI itself but for repository cleanup, and the student describes it as capable enough to plan across multiple files and finish four of five tasks in five instructions, though he still preferred Claude Code. The broader point is about degrees of freedom: a fixed 8B model forces a team to first discover what the tool can do, while an agent that can route between several models is a different kind of companion. The outcome, and the value of the lesson, hinges on whether teams pair part-level tests with a whole-goal test from the outset.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Google is paying about 100 digital publishers for content used in AI Overviews, AI Mode, and Gemini, with paym…

A Qiita walkthrough trained a five-label car-damage classifier on Gemini Enterprise Agent Platform AutoML usin…

Lauren Tan says she shipped about 2,000 pull requests a month to production on the SpaceX AI Grok Bot team

At its September 29, 2026 DevDay, OpenAI announced more than 20 items, including dots, an agent running on GPT…

Alibaba's Qwen team open-sourced Qwen-Image-2.1 on September 20, 2026 — a 7B model generating 2048×2048 images…

Reading one spec-index file of 803 lines on 2026年9月29日, an external program returned about 798 tokens against…
