AIToday
Video GenerationZenn AI/MLPublished: Oct 10, 2026, 22:00 JST

3 models, one chinchilla, $1.12: all three failed alike

3 models, one chinchilla, $1.12: all three failed alike

Three video models — h3-max-turbo, veo3.1-lite and wan-3.0 — each got the same 1672×941 chinchilla image and the same prompt. All three failed the same three of five checks.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The test started from a practical question: which animal scenes earn views. Counting 1,279 YouTube animal videos, the author found median views varied by an order of magnitude — 53.1 million for drinking scenes, 21.25 million for sneezing, but 5.16 million for grooming. Grooming looked safe and its motion easy to see, so it was chosen. That choice is what later mattered.

The first round changed only the model, holding image, prompt, length and judging consistent, with Codex scoring five items. All three models passed camera lock and centering and failed the other three, and their failure locations matched. wan-3.0, the most expensive at $0.10 per second, was the only one to fail centering and changed the face mid-clip. The two cheaper models matched item by item, so price did not predict quality.

The author then separated model limits from subject difficulty by hiding the front paws and shrinking the motion to whisker twitches and blinks, dropping wan-3.0. Fusion, warping and disappearance did not reappear. What remained across both remaining models were thin lines — whisker count and length, and forehead and cheek stripes. These point to a shared weakness in current models, not a gap between them.

FAQ
Which models were tested?
h3-max-turbo, veo3.1-lite and wan-3.0, all accessed through fal.ai. They ran at $0.01, $0.03 and $0.10 per second respectively.
How much did the test cost?
The total was $1.12 (about 170 yen). The balance shown right after generation did not reflect the charge, which appeared later.
What broke in all three videos?
Camera lock and centering mostly passed, but natural motion, shape consistency and image quality failed in all three. The front paws fused, lost their outline or differed left to right.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleApate's 350,000 AI bots tie up phone scammers for hours