AIToday
Large Language ModelsAI Coding AssistantsQiita 機械学習Published: Oct 7, 2026, 19:01 JST

cairosvg pipeline scores LLM SVGs at 0.218 vs 0.005

cairosvg pipeline scores LLM SVGs at 0.218 vs 0.005

3 Key Points

  1. What happened

    The author built a Python pipeline using cairosvg that extracts SVG from LLM replies, statically checks format and forbidden elements like script, renders to a fixed 800x600 PNG, and measures foot-to-pedal gap on labeled parts. The demo scored a LIFTED foot at 0.218 wheel-diameters versus 0.005 for GOOD.

  2. Why it matters

    The render-based measurement correctly detected that the foot was lifted, even though the LIFTED version only used a transform and did not change the cy value. The gap is expressed as a ratio to the back wheel's apparent diameter, so artworks of different sizes can be compared.

  3. What to watch

    The threshold for how close counts as contact should not be used for judgment until calibrated on human-labeled samples. Also, when the part ID is missing or not visible, the system returns uncertain rather than guessing.

WHO IT HITSThis work matters to AI engineers and evaluators who score generated SVG outputs, as well as anyone building automated grading pipelines for visual tasks. It offers a way to make grading reproducible and to separate format compliance from image quality.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The author notes that comparing LLMs by asking them to draw a pelican on a bicycle has become a common pastime. But once you have dozens of images, the question becomes who grades them and how. Visual inspection is tiring and the criteria drift from day to day. Even if you assign a total score of 8, you cannot explain the reasoning afterward. This pipeline tries to solve that by automating the parts that can be supported by evidence: format, rendering, and auxiliary geometric measurements.

The author tested whether counting DOM elements like circles can stand in for visual judgment. Drawing the same two wheels using circle, path, and use produced identical rendered pixels, but the circle counts were 2, 0, and 1 respectively. This shows that counting elements only looks at how the SVG is written, not at what is actually drawn on screen. The author therefore measures geometry from the rendered output instead, by removing an element and checking which pixels change.

The pipeline records judgments in four states: pass, fail, uncertain, and not_applicable. If any required check fails, the overall verdict is fail. If no check fails but uncertain remains, the verdict is uncertain and goes to review. When aggregating results, the author warns against dropping uncertain cases, which can inflate the pass rate. The true range of the pass rate includes uncertain cases, and reporting only the rate after dropping them gives a misleadingly better number. The author also notes that the grading design is based on ideas from a pelican grading article on the Folkbench blog, and that the author is involved in Folkbench's development. The measurement values in the article are execution results of the posted code or illustrative examples, not measurements of actual models or services.

FAQ
How does the pipeline measure whether a foot is touching a pedal?
It removes the element with the specified ID, renders the SVG again, and sees which pixels changed. Then it measures the shortest distance from the visible pixels of one part to the other, dividing by the back wheel's apparent diameter.
What happens if the required part ID is missing from the SVG?
The contact measurement returns uncertain and no gap value. The pipeline does not guess; it records uncertain so that the result is not automatically counted as pass.
Can this pipeline replace human judgment for whether a bird is correctly riding a bike?
No. The author states the pipeline does not automate semantic judgment like whether the bird looks like it is riding correctly. That part still needs human review or a visual judge, and the pipeline provides the foundation to record their input and results.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI opens Decisions API beta, up to 10x faster than GPT-6 Luna