AIToday
Image GenerationApple Machine LearningPublished: Jul 18, 2026, 10:00 JST2 min read

Apple researchers introduce VICIS to test AI models on visual concept inference

Apple researchers introduce VICIS to test AI models on visual concept inference

Key takeaway

  • Apple researchers have introduced VICIS, a new task designed to test whether vision-language models can infer visual concepts from sets of example images.

  • Current state-of-the-art models perform poorly on this capability, often ignoring visual context or defaulting to biased outputs.

  • The researchers propose a training framework and architecture that learns to extract and apply visual concepts, demonstrating improved performance and generalization on synthetic and large-scale real-world image datasets.

3 Key Points

  1. What happened

    Apple researchers have introduced Visual Concept Inference from Sets (VICIS), a task that evaluates whether AI vision-language models can learn shared visual concepts from sets of example images and apply those concepts to new inputs. The researchers found that current state-of-the-art models perform poorly on this task, often ignoring visual context or producing biased outputs.

  2. Why it matters

    Vision-language models are widely used but struggle with reasoning from visual examples alone—a capability needed for tasks like generating variations of a concept shown through images. The work identifies a concrete gap in how these models learn from visual context rather than text, which could inform development of more capable AI systems for image understanding and generation.

  3. What to watch

    The researchers propose a training framework and architecture to address this gap. Their model was tested on synthetic data and large-scale ImageNet/WordNet datasets, and showed improved accuracy and diversity in outputs while generalizing to unseen concepts and modalities such as sketches.

Ask the AI about this article →

Context & Analysis

Vision-language models have become central to modern AI applications because they can follow complex textual instructions. However, this strength in text-based reasoning masks a fundamental weakness: the ability to reason from purely visual examples. The VICIS task targets this gap directly. Rather than relying on language descriptions, a model must observe visual patterns across multiple example images and extract the unifying concept—then apply it consistently to generate new images.

The poor performance of current state-of-the-art models on VICIS reveals that visual concept learning from examples is not an emergent property of scaling or instruction-following capability. The models either discount the visual examples in favor of textual biases or generate outputs that lack coherence with the provided context. This suggests that the architecture and training dynamics of today's vision-language models are not well-suited to this type of visual reasoning, a challenge the researchers address through a new training framework and architecture design.

FAQ

What is VICIS and what does it test?
VICIS (Visual Concept Inference from Sets) is a task that evaluates whether AI models can infer shared concepts from a small set of example images and apply those concepts to new inputs. Given a context set of images sharing a concept and a query image, the model must generate new images that preserve the context-defined concept while remaining consistent with the query.
How did current state-of-the-art models perform on this task?
State-of-the-art vision-language models performed poorly on VICIS, often ignoring the visual context or defaulting to biased generations.
What datasets did the researchers use to evaluate their approach?
Experiments were conducted on synthetic data and large-scale ImageNet/WordNet data, and the model was shown to generalize to unseen concepts and modalities such as sketches.
Apple Machine LearningRead Original Article

Get the latest Image Generation news every morning

For example, today's edition would include:

  • Why AI images feel 'cringey' to consumersITmedia AI+ · 4h ago
  • Disney concept art auction fetches $3.43MTop Companies AI · 2d ago
  • ESP32-P4 Reads Water Meter with AIr/robotics · 2d ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI boom pushes Prologis industrial real estate demand higher, but costs rise