
Apple researchers have introduced VICIS, a new task designed to test whether vision-language models can infer visual concepts from sets of example images.
Current state-of-the-art models perform poorly on this capability, often ignoring visual context or defaulting to biased outputs.
The researchers propose a training framework and architecture that learns to extract and apply visual concepts, demonstrating improved performance and generalization on synthetic and large-scale real-world image datasets.
What happened
Apple researchers have introduced Visual Concept Inference from Sets (VICIS), a task that evaluates whether AI vision-language models can learn shared visual concepts from sets of example images and apply those concepts to new inputs. The researchers found that current state-of-the-art models perform poorly on this task, often ignoring visual context or producing biased outputs.
Why it matters
Vision-language models are widely used but struggle with reasoning from visual examples alone—a capability needed for tasks like generating variations of a concept shown through images. The work identifies a concrete gap in how these models learn from visual context rather than text, which could inform development of more capable AI systems for image understanding and generation.
What to watch
The researchers propose a training framework and architecture to address this gap. Their model was tested on synthetic data and large-scale ImageNet/WordNet datasets, and showed improved accuracy and diversity in outputs while generalizing to unseen concepts and modalities such as sketches.
Ask the AI about this article →
Vision-language models have become central to modern AI applications because they can follow complex textual instructions. However, this strength in text-based reasoning masks a fundamental weakness: the ability to reason from purely visual examples. The VICIS task targets this gap directly. Rather than relying on language descriptions, a model must observe visual patterns across multiple example images and extract the unifying concept—then apply it consistently to generate new images.
The poor performance of current state-of-the-art models on VICIS reveals that visual concept learning from examples is not an emergent property of scaling or instruction-following capability. The models either discount the visual examples in favor of textual biases or generate outputs that lack coherence with the provided context. This suggests that the architecture and training dynamics of today's vision-language models are not well-suited to this type of visual reasoning, a challenge the researchers address through a new training framework and architecture design.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Recent controversies include Ajinomoto's official X account posting an AI-edited image and a restaurant menu s…

Heritage Auctions' sale of Disney art last weekend generated $3.43 million across 1,346 lots

A hobbyist built a system that reads a water meter using the Makerfabs ESP32-P4 board with autofocus camera

A developer created a tiny image generation model, a latent flow transformer with 12 layers, that runs fully o…

A new episode of an AI image benchmark compares 33 image models from 8 providers, including Meta Muse Image 1.…

Since August 13, Cara, an image-sharing app for artists, was hit by three major scrapes
