
What happened
Grab published how it moved over 500 internal agent services onto one in-house framework, LLM-Kit, cutting production connection from over two weeks to about one hour, with an OpenAI-compatible gateway bundling five providers and over 50 MCP servers fetched at runtime.
Why it matters
That gateway gives Grab one place to see model costs, swap providers, and gather credentials, so app code does not bake in a provider name; the framework supplies only the scaffolding, so teams still keep their own quality standards.
What to watch
The claimed two weeks to one hour gain hinges on Grab's existing infrastructure and unique scale, so a company with only five or ten internal AI features likely cannot justify the same central build; watch the count of shared services before copying it.
WHO IT HITSEnterprise platform and infrastructure teams at companies running many internal AI features will read this as a blueprint for centralizing model access; firms with only a handful of such services may find the central build cost outweighs the benefit.
Summaries like this, in your inbox every morning.
Grab's write-up centers on a split that is easy to miss. The inference loop itself, meaning the core logic where an AI reasons, calls tools, and answers, took a single afternoon to write. The surrounding production shell, covering where credentials live, how logs and traces are wired, service discovery registration, and evaluation, took two weeks. Grab's move was to stop letting each team decide those elements on its own. A form now returns a GitLab repository with secret management and service discovery already wired, plus an evaluation endpoint attached from the first commit. Notably, Grab did not invent a new abstraction layer or a dedicated language; it built scaffolding once around infrastructure the company already had.
What moved to the center is the calling interface. Model calls pass through one OpenAI-compatible gateway that bundles five providers and inserts credentials, so app code does not hard-code a provider name. Tools are not bundled in advance; they are fetched at runtime from over 50 MCP servers, a shared connection standard for letting AI use tools. Adding a new model or tool needs no redeploy, only a registration or config change. The layers above that, however, stayed with the teams. Five hundred services can share the same scaffolding, yet prompt quality and the criteria for what counts as improvement still exist once per team.
The article also notes a stated tension. Grab said it chose a framework over a platform to avoid locking teams in, then later wrote that at 500 services the problem had shifted from framework to platform. The write-up reads this sequence as deliberate: calling it a framework got 500 services onto it, and that scale then justified centralizing the shared pieces. The piece also cites a separate analyst's warning that agent lock-in is harder to unwind than API lock-in because it accumulates across four layers at once. For a company weighing whether to build the same thing, the article argues the first move is a single model calling interface, which alone covers cost visibility, provider swapping, and credential aggregation, while MCP registries and evaluation infrastructure can wait until enough services share them.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
FieldAI is reportedly seeking $700 million in funding at a valuation that could reach $10 billion, five times…
SpaceX finished its first full quarter as a public company on Sept

After OpenAI staff raised concerns on Slack, president Greg Brockman abandoned the second half of his $50 mill…

A practitioner listed five Japanese-language books he keeps re-opening, from '機械学習 100+ページ エッセンス' by Andriy Bu…

The article lays out the split in Claude Code — CLAUDE.md is the file the user writes with instructions and ru…

Cloudflare's Day 4 announcements made AI Search and the Cloudflare Basin data platform generally available, op…
