
What happened
Amazon announced rate limiting support for Bedrock AgentCore gateway, a managed AI gateway that routes traffic to tools, models, and agents. The feature enables per-user and per-group rate controls using request rates (RPS/RPM), token throughput (TPM), and concurrent connections (CPS).
Why it matters
Organizations can now prevent individual users or groups from monopolizing gateway capacity during traffic spikes. Rate limits are enforced at multiple layers—service quotas set the ceiling, while customer-defined rules (grouped by user role, identity, or target) ensure fair usage across callers and keep downstream services available.
What to watch
Configuration uses dimension keys (such as JWT role claims, IAM principals, or target names) paired with rate entries; a two-layer enforcement model (group-level and per-user limits combined with AND logic) prevents both cross-group and within-group capacity hoarding. Basic users in the example receive 100 RPM / 50 CPS per group and 20 RPM / 10 CPS individually; Advanced users receive 300 RPM / 150 CPS per group and 60 RPM / 30 CPS individually.
Summaries like this, in your inbox every morning.
Bedrock AgentCore gateway addresses a fundamental challenge in managed AI infrastructure: ensuring fair resource allocation when multiple callers (users, teams, or applications) share a single entry point. Prior to this update, organizations had no built-in mechanism to prevent a single heavy user from degrading service for others or to enforce different policies across user tiers. Rate limiting closes that gap by enabling fine-grained, identity-aware traffic shaping at the gateway layer.
The two-layer enforcement model—combining per-group and per-user limits—reflects a practical requirement. Group-level limits alone would allow dominant individuals to consume an entire group's quota, while per-user limits alone would not prevent one user group from starving another. By evaluating both independently with AND logic, the system protects both cross-group and within-group fairness. The use of JWT claims and IAM principals as dimension keys integrates seamlessly with existing identity providers (such as Microsoft Entra ID in the example) and role-based access control policies already enforced by Bedrock AgentCore, creating a cohesive identity and authorization story.
The distinction between token-based rate limits (for inference) and request-based limits (for all targets) reflects the different cost profiles of AI workloads: a streaming inference call that returns thousands of tokens consumes one request slot but many token slots, so both metrics are necessary. Service-managed quotas provide an account-level safety ceiling, ensuring no customer-defined policy can overwhelm the shared AWS infrastructure.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
On September 28, Nvidia announced the NVIDIA Open Agent Safety Platform, built from the open-source OpenShell…

Anthropic released Claude Sonnet 5.5 on September 28, the second model in its Claude 5.5 family, calling it fa…

Dentsu Group embedded generative AI in its accounting approval process and says it reached 99% accuracy on the…

Nvidia launched a platform Monday that it says can monitor AI agents and quarantine suspicious ones, following…

Nvidia announced a record $150 billion stock buyback, and CEO Jensen Huang said its cash generation gives it t…

A September 2026 guide organizes AI coding tools into three types — terminal/agent tools like Claude Code, AI-…
