GAZE (Grounded Agentic Zero-shot Evaluation) lets medical vision-language models call viewer-level tools (zoom, windowing, contrast, edge detection) and retrieval tools backed by PubMed and Open-i, with structured outputs validated against a schema and full tool-call traces recorded for auditability.
On NOVA, a benchmark of 906 brain MRI cases covering 281 rare neurological conditions, GAZE reaches 58.2 mean average precision (mAP) at intersection-over-union (IoU) 0.3 for lesion localization and 34.9% Top-1 diagnostic accuracy under a joint protocol scoring captioning, diagnosis, and localization, without task-specific fine-tuning.
Tool use helps rare pathologies disproportionately: the fraction of cases with IoU > 0.3 rises from 17% to 58% for diagnoses with three or fewer examples versus 25% to 68% for common conditions (≥10 cases), with retrieval ablations revealing a model-dependent trade-off in which gains in diagnosis can coincide with losses in localization.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic updated the system prompt for Claude 5.1, adding a strict ban on reproducing song lyrics, poems, or…

The Pentagon added OpenAI's ChatGPT Mil and xAI's Grok for Government to its AI platform GenAI.mil, which prev…

Amazon Web Services (AWS) has started offering “AWS Cloud Quest 2.0,” a new version of its online game that te…

A job seeker named Christopher, after five unanswered AI interviews with recruiter 'Riley' from IT firm Everfo…

Sandisk says its NAND-based High Bandwidth Flash (HBF) technology can match HBM bandwidth while providing eigh…

World Labs unveiled Atlas, an omni-model trained on text, images, video, and 3D data that anchors every input…
