
Tomofun, a Taiwan-headquartered pet-tech startup behind the Furbo Pet Camera, migrated its pet behavior detection inference workloads from GPU-based Amazon EC2 instances to EC2 Inf2 instances powered by AWS Inferentia2 to reduce costs while maintaining accuracy for always-on monitoring across hundreds of thousands of devices.
The BLIP vision-language model (a model that interprets images and generates text descriptions) was decomposed into three components—Image Encoder, Text Encoder, and Text Decoder—each compiled independently using torch_neuronx and combined into the inference pipeline without altering the original pretrained logic, using lightweight wrapper classes to standardize inputs and outputs.
The architecture uses Elastic Load Balancing and EC2 Auto Scaling groups across two layers: one for API servers and one dedicated to model inference on Inf2 instances, with Amazon CloudWatch monitoring latency, throughput, and error rates to maintain service-level objectives as demand shifts.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Amazon Web Services (AWS) has integrated its fully managed data warehouse service, Amazon Redshift, with Agent…

Visual Studio Code 1.135 now includes an experimental 'Rubber Duck' feature that lets developers request a sec…

Sonos announced a new app update with generative AI features, a new soundbar called the Beam Ultra, and its se…
