
What happened
AWS announced the general availability of TwelveLabs Marengo Embed 3.0 as an embedding model in Amazon Bedrock Knowledge Bases, a fully managed RAG service that ingests video (MP4, MOV), images (JPEG, PNG), and audio.
Why it matters
Previously, semantic search over video meant stitching together transcription, frame extraction, embedding models, vector databases, and synchronization logic; AWS says Managed MKB now handles segmentation, frame sampling, and transcription internally.
What to watch
Availability is limited to the US East (N. Virginia) and US West (N. California) Regions, so the test is whether AWS expands to more Regions. Embeddings generation with Marengo Embed 3.0 is charged at the standard Amazon Bedrock model invocation rate.
WHO IT HITSMedia, sports analytics, education, security, and retail teams that hold large video or image archives are the clearest beneficiaries, since they can run semantic search on media without assembling a transcription, frame-extraction, and vector-database pipeline. Enterprise IT and developer teams evaluating retrieval tools may also weigh a managed option against self-built stacks.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
AWS is positioning video and image files as a searchable data source for everyday business teams, not just machine-learning engineers. The article describes the old path as a chain of separate services: transcription, frame extraction, an embedding model, a vector database, and synchronization logic. By offering TwelveLabs Marengo Embed 3.0 inside Amazon Bedrock Knowledge Bases, AWS says that chain is replaced by a managed workflow. The walkthrough narrative uses a 10-minute clip of the 2022 FIFA World Cup final, ingests it, and runs a query like "show me the penalty kicks from this soccer match." The top results identify moments where penalty kicks were attempted.
That example matters because it shows the type of question the system answers: not keyword matching but a meaning-based search across a long video. The body also notes Marengo Embed 3.0 encodes video, audio, images, and text into a compact, storage-efficient 512-dimensional vector space, and that Managed MKB generates embeddings capturing visual, textual, speech, and audio signals. The console defaults to 4 seconds for both audio and video segmentation, a detail that may matter to teams tuning retrieval granularity. The article lists media, sports analytics, education, security, and retail as the industries that need this capability, and offers a sports example and a security-camera example.
Availability may be the main limitation for some readers. The article lists only the US East (N. Virginia) and US West (N. California) Regions, so teams outside those areas would need to wait or route around them. The article does not say whether additional Regions are planned. Whether this becomes a default tool for media-heavy organizations likely hinges on how well the managed pipeline handles their particular footage and how the per-retrieval pricing compares with their current approach.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…

Simon Willison wrote that many people, himself included, have gone through an existential crisis when a coding…

Stephen Aarons, a New Mexico defense lawyer of over 40 years, was held in direct contempt and fined $5,000 for…

Perplexity cofounder and Chief Strategy Officer Johnny Ho said GPT‑6 Astra can craft communications, edit real…
