
System uses late probability fusion of six encoder families: S4D-ViTMoE face encoder, Wav2Vec2 audio features, TimeSformer and VideoMAE body-language encoders, and Gemini Embedding 2.0
Frozen Wav2Vec2 prosody layers (6-12) outperform fine-tuning approaches (0.207 vs 0.161 score) by avoiding irrelevant phonetic feature extraction
Gemini Embedding 2.0 achieves competitive 0.320 ACCP accuracy using only 2 seconds of video input, marking first use of large multimodal models in emotion recognition
Post-processing salience threshold varies significantly across data folds (0.05-0.43), indicating personalized expression styles require adaptive approaches
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…

Nvidia CEO Jensen Huang said on Fox Business that AI is creating 'hundreds of thousands' of jobs, including in…

Israeli startup DataAgent Ltd