AIToday
AI Safety & AlignmentarXiv cs.CVPublished: Apr 29, 2026, 13:00 JST1 min read

XTC-Bench framework reveals unified multimodal models show weak cross-task consistency despite high individual performance

3 Key Points

  1. Researchers introduced XTC-Bench, a scene-graph-grounded evaluation framework that measures whether unified multimodal models (systems supporting both visual understanding and generation in a shared representation) maintain semantic consistency across tasks given a visual concept.

  2. The framework uses Continuous Cross-Task Agreement (CCTA), a fine-grained metric that quantifies semantic agreement between generation and understanding over matched atomic facts (objects, attributes, and relations), isolating internal consistency from standalone task accuracy.

  3. Experiments on eight open-source and one commercial unified models found that high generation or understanding performance does not imply strong cross-task alignment; architectural analysis shows consistency is governed by how tightly learning objectives are coupled across modalities, not by architectural unification alone.

Ask the AI about this article →

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • CrowdStrike Falcon Guardian Targets AI SecurityTop Companies AI · 2h ago
  • Palo Alto Networks pitches Authority-Aware DLP for AI agentsTop Companies AI · 2h ago
  • Music Publishers Sue Anthropic Over Lyrics CopyrightTop Companies AI · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI's missed revenue targets raise questions about ability to meet massive computing commitments exceeding $1.15 trillion