AIToday
Large Language ModelsAI Business & IndustryAmazon AI BlogPublished: Aug 7, 2026, 04:01 JST

AWS adds rate limiting to Bedrock AgentCore gateway

AWS adds rate limiting to Bedrock AgentCore gateway

3 Key Points

  1. What happened

    Amazon announced rate limiting support for Bedrock AgentCore gateway, a managed AI gateway that routes traffic to tools, models, and agents. The feature enables per-user and per-group rate controls using request rates (RPS/RPM), token throughput (TPM), and concurrent connections (CPS).

  2. Why it matters

    Organizations can now prevent individual users or groups from monopolizing gateway capacity during traffic spikes. Rate limits are enforced at multiple layers—service quotas set the ceiling, while customer-defined rules (grouped by user role, identity, or target) ensure fair usage across callers and keep downstream services available.

  3. What to watch

    Configuration uses dimension keys (such as JWT role claims, IAM principals, or target names) paired with rate entries; a two-layer enforcement model (group-level and per-user limits combined with AND logic) prevents both cross-group and within-group capacity hoarding. Basic users in the example receive 100 RPM / 50 CPS per group and 20 RPM / 10 CPS individually; Advanced users receive 300 RPM / 150 CPS per group and 60 RPM / 30 CPS individually.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Bedrock AgentCore gateway addresses a fundamental challenge in managed AI infrastructure: ensuring fair resource allocation when multiple callers (users, teams, or applications) share a single entry point. Prior to this update, organizations had no built-in mechanism to prevent a single heavy user from degrading service for others or to enforce different policies across user tiers. Rate limiting closes that gap by enabling fine-grained, identity-aware traffic shaping at the gateway layer.

The two-layer enforcement model—combining per-group and per-user limits—reflects a practical requirement. Group-level limits alone would allow dominant individuals to consume an entire group's quota, while per-user limits alone would not prevent one user group from starving another. By evaluating both independently with AND logic, the system protects both cross-group and within-group fairness. The use of JWT claims and IAM principals as dimension keys integrates seamlessly with existing identity providers (such as Microsoft Entra ID in the example) and role-based access control policies already enforced by Bedrock AgentCore, creating a cohesive identity and authorization story.

The distinction between token-based rate limits (for inference) and request-based limits (for all targets) reflects the different cost profiles of AI workloads: a streaming inference call that returns thousands of tokens consumes one request slot but many token slots, so both metrics are necessary. Service-managed quotas provide an account-level safety ceiling, ensuring no customer-defined policy can overwhelm the shared AWS infrastructure.

FAQ
What metrics can be rate-limited on AgentCore gateway?
The gateway supports three types of rate limiting: request rates (measured in requests per second or requests per minute, applying to all target types), token rates (measured in tokens per minute for inference targets only, accounting for both input and output tokens), and connection rates (measured in connections per second for all target types, tracking how long each request holds an open connection).
How does the two-layer rate limiting model work?
Customer-defined rate limits are evaluated first; if a request passes, service quotas are evaluated. Additionally, organizations can combine per-group limits and per-user limits using AND logic—a request must satisfy both the group-level ceiling and the individual user's cap to proceed. For example, Basic users might share a 100 RPM group pool but each individual is capped at 20 RPM, preventing one user from consuming the entire group's quota.
What dimension keys can be used to define rate buckets?
Dimension keys include targetName, toolName, qualifiedModelId, JWT claims ($.context.jwt.<claim>), IAM principal ($.context.iam.principal), and IAM source identity ($.context.iam.sourceIdentity). These keys determine how the gateway groups incoming traffic into separate rate buckets; combinations of multiple keys enable more granular control.
Amazon AI BlogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Anthropic ships Claude Sonnet 5.5, 30%+ faster at same priceITmedia AI+ · 1h ago
  • Dentsu AI clears 99% on 20万件超 approvalsITmedia AI+ · 1h ago
  • Nvidia launches tool to quarantine rogue AI agentsSemafor Tech · 1h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI publishes country-level ChatGPT usage data