AIToday
Large Language ModelsVideo GenerationApple Machine LearningPublished: Sep 12, 2026, 01:00 JST2 min read

Apple's CapQuiz grades video captions by quiz answers

Apple's CapQuiz grades video captions by quiz answers

3 Key Points

  1. What happened

    Apple researchers introduced CapQuiz, a benchmark that scores video captions by how well they answer human-verified multiple-choice questions from the video, covering 10 question types across 24 video domains, plus a composite metric, CapF1.

  2. Why it matters

    Existing metrics match text against reference captions, which penalizes valid descriptions that differ in wording or visual focus. CapQuiz correlates significantly better with human judgments, and CapF1 combines factuality (CapP) and coverage (CapR) rather than collapsing them into one score.

  3. What to watch

    The test is whether reference-free, question-based scoring becomes a trusted way to compare video models, since it trades text matching for verifiable answers. Watch whether CapQuiz and CapF1 are adopted beyond this research setting.

WHO IT HITSAI researchers and teams building or evaluating video-capable models are the main audience, since the benchmark offers a way to compare captions without relying on reference text. Model developers preparing to ship video features may also use it to spot coverage and factuality gaps before release.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Video captioning has long been judged by comparing a model's generated text against a ground-truth reference. The Apple researchers behind CapQuiz argue this is a poor fit for video, because a single video can be described in many valid ways. Two captions can both be accurate while emphasizing different visual details, and under reference-matching that legitimate difference shows up as an error. The authors also note that existing assessments tend to be one-dimensional, without a fine-grained breakdown of caption quality.

Their alternative reframes the question from "does this caption look like the reference?" to "does this caption carry the information that matters?" That is what they call information fidelity: a good caption should cover salient visual content while remaining strictly factual. CapQuiz operationalizes this by turning each video into human-verified multiple-choice questions, spanning a hierarchical taxonomy of 10 question types across Descriptive and Inferential categories and 24 video domains. CapF1 then separates the score into two parts, CapP for factuality and CapR for coverage, so the two dimensions can be read independently.

The stated payoff is empirical: across extensive experiments, CapQuiz correlates significantly better with human judgments than existing metrics and offers interpretable insights into model performance. Whether that translates into a durable standard is likely to hinge on adoption outside this research setting, and on whether teams find the question-based signal easier to act on than a single similarity score.

FAQ
What does CapQuiz actually measure?
It measures information fidelity: how well a caption covers salient visual information while staying strictly factual. The benchmark scores captions by their utility in answering human-verified, fine-grained multiple-choice questions derived from the video.
What is CapF1 and how is it different from other metrics?
CapF1 is a composite metric that synthesizes CapP, which measures factuality, and CapR, which measures coverage. It is part of the same reference-free approach as CapQuiz.
Why not just compare captions to reference text?
Because video description is one-to-many: high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. The authors say CapQuiz correlates significantly better with human judgments than existing metrics.
Apple Machine LearningRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 4h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 4h ago
  • OpenAI agents hit RubyGems, undisclosed since May 12thSimon Willison's Weblog · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleHuawei stakes out AI optics rulebook with 7.2Tbps NPO