AIToday
Large Language ModelsarXiv cs.CVPublished: Mar 26, 2026, 13:00 JST1 min read

Multimodal AI System Combines Six Encoder Families Including Gemini 2.0 to Recognize Blended Emotions with 32% Accuracy

Multimodal AI System Combines Six Encoder Families Including Gemini 2.0 to Recognize Blended Emotions with 32% Accuracy

3 Key Points

  1. System uses late probability fusion of six encoder families: S4D-ViTMoE face encoder, Wav2Vec2 audio features, TimeSformer and VideoMAE body-language encoders, and Gemini Embedding 2.0

  2. Frozen Wav2Vec2 prosody layers (6-12) outperform fine-tuning approaches (0.207 vs 0.161 score) by avoiding irrelevant phonetic feature extraction

  3. Gemini Embedding 2.0 achieves competitive 0.320 ACCP accuracy using only 2 seconds of video input, marking first use of large multimodal models in emotion recognition

  4. Post-processing salience threshold varies significantly across data folds (0.05-0.43), indicating personalized expression styles require adaptive approaches

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 2h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 2h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleResearchers introduce Cluster-R1, a reasoning-based AI system that outperforms standard embedding models by autonomously interpreting user instructions to determine optimal text clustering without manual intervention.