
What happened
AWS published a code sample deploying a real-time endpoint for FLUX.2-klein-4B image generation and an asynchronous endpoint for Wan2.1-VACE-1.3B video generation from one AWS vLLM-Omni DLC on SageMaker AI.
Why it matters
Sharing one container across both models could simplify updates and consistency for teams running generative media, while separate endpoints let each model use the instance type and inference option that fits.
What to watch
The numbers AWS reports are single-run checkpoints, not benchmarks, so real throughput and latency will depend on prompts and instance availability. Watch the default 17 frames and 30 diffusion steps, and the note that four steps did not preserve source composition.
WHO IT HITSML platform engineers and generative media teams who run image and video models on AWS will likely find this sample useful for reducing serving-stack variation, though they still need to choose instance types and validate GPU memory per model.
Summaries like this, in your inbox every morning.
This post is Part 2 of an AWS series on specialized Deep Learning Containers. Part 1 used vLLM-Omni and SageMaker AI bidirectional streaming for real-time speech. Part 2 moves to generative media: vLLM-Omni extends vLLM beyond text to models that process or generate text, audio, images, and video through OpenAI-compatible APIs. AWS packages tracked vLLM-Omni releases into a DLC and adds routing middleware for SageMaker AI, so the same container image can serve different model families.
The sample deploys one pinned DLC image to two endpoints and switches models by setting SM_VLLM_MODEL. The image endpoint runs FLUX.2-klein and returns a base64-encoded PNG inline, while the video endpoint runs Wan VACE and writes the MP4 to Amazon S3. AWS says it chose asynchronous inference for video because generation is longer-running and carries an image-conditioned multipart payload, a workload choice rather than a rule for every video model.
The reported timings from one US East (N. Virginia) Region run are reproduction checkpoints, not benchmarks, so they should not be read as expected performance. Whether teams adopt this pattern will likely hinge on validating GPU memory and performance per model and on instance availability for ml.g6.xlarge and ml.g6e.xlarge, since the video endpoint here stays on a fixed instance type.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Adobe announced in September 2026 it is expanding its connector for Google Gemini and its plugin for Claude, s…

Comfy, the developer of the local image and video generation app ComfyUI, released Comfy Router on September 2…

ASICS built a generative AI system that creates on-foot styling images of its shoes without a photo shoot or m…

Synthesia, a digital avatar startup valued at $4 billion, made an interactive avatar for a TechCrunch reporter…

An undervalued-stock screen highlighted Adobe (ADBE) and Broadcom (AVGO) as AI-driven picks, flagging Adobe's…

Casio's G-SHOCK team used AI image generation alongside its own designers to develop a watch concept into a pr…
