
What happened
The author built a Python pipeline using cairosvg that extracts SVG from LLM replies, statically checks format and forbidden elements like script, renders to a fixed 800x600 PNG, and measures foot-to-pedal gap on labeled parts. The demo scored a LIFTED foot at 0.218 wheel-diameters versus 0.005 for GOOD.
Why it matters
The render-based measurement correctly detected that the foot was lifted, even though the LIFTED version only used a transform and did not change the cy value. The gap is expressed as a ratio to the back wheel's apparent diameter, so artworks of different sizes can be compared.
What to watch
The threshold for how close counts as contact should not be used for judgment until calibrated on human-labeled samples. Also, when the part ID is missing or not visible, the system returns uncertain rather than guessing.
WHO IT HITSThis work matters to AI engineers and evaluators who score generated SVG outputs, as well as anyone building automated grading pipelines for visual tasks. It offers a way to make grading reproducible and to separate format compliance from image quality.
Summaries like this, in your inbox every morning.
The author notes that comparing LLMs by asking them to draw a pelican on a bicycle has become a common pastime. But once you have dozens of images, the question becomes who grades them and how. Visual inspection is tiring and the criteria drift from day to day. Even if you assign a total score of 8, you cannot explain the reasoning afterward. This pipeline tries to solve that by automating the parts that can be supported by evidence: format, rendering, and auxiliary geometric measurements.
The author tested whether counting DOM elements like circles can stand in for visual judgment. Drawing the same two wheels using circle, path, and use produced identical rendered pixels, but the circle counts were 2, 0, and 1 respectively. This shows that counting elements only looks at how the SVG is written, not at what is actually drawn on screen. The author therefore measures geometry from the rendered output instead, by removing an element and checking which pixels change.
The pipeline records judgments in four states: pass, fail, uncertain, and not_applicable. If any required check fails, the overall verdict is fail. If no check fails but uncertain remains, the verdict is uncertain and goes to review. When aggregating results, the author warns against dropping uncertain cases, which can inflate the pass rate. The true range of the pass rate includes uncertain cases, and reporting only the rate after dropping them gives a misleadingly better number. The author also notes that the grading design is based on ideas from a pelican grading article on the Folkbench blog, and that the author is involved in Folkbench's development. The measurement values in the article are execution results of the posted code or illustrative examples, not measurements of actual models or services.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
At Team ‘26 Europe, Atlassian unveiled Agentic Multiplayer Protocol, giving every AI agent an admin-assigned i…
System76's COSMIC desktop now requires contributors to declare: "I have not included any LLM (also known as AI…

Meta and enterprise AI startup Sierra are leading Walmart, Stripe, Shopify, Genesys, Instinct and Rocket in de…

Gartner forecasts that 70% of AI agents developed by vendors' forward-deployed engineers (FDE) will be abandon…

OpenAI published 372 mathematical results from an internal frontier model on GitHub, including improvements to…

Impress released Kamome Ashizawa's book on October 7 for 1,980 yen
