AIToday
Large Language ModelsGoogle AI BlogPublished: Apr 3, 2026, 07:00 JST1 min read

Google launches Flex and Priority tiers for Gemini API to give developers flexible options for managing costs versus response speed

Google launches Flex and Priority tiers for Gemini API to give developers flexible options for managing costs versus response speed

3 Key Points

  1. Google introduces two new inference tiers to the Gemini API: Flex for cost-optimized workloads and Priority for latency-sensitive applications

  2. The new tiers allow developers to balance their specific needs between API expenses and response time performance

  3. This expansion gives users more granular control over their Gemini API usage patterns and pricing models

Ask the AI about this article →

Google AI BlogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Walmart settles opioid claims for $50MTop Companies AI · 2h ago
  • Tim Cook's legacy hinges on Apple's AI betTop Companies AI · 2h ago
  • CrowdStrike Falcon Guardian Targets AI SecurityTop Companies AI · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleLegal battles against Tempus AI challenge the boundaries of genetic data collection and privacy rights in healthcare AI applications.