
EXO Labs says four Mac Studios with M5 Ultra chips can act as one AI cluster with 4.8TB/s memory bandwidth.
This enables large AI models to run locally at speeds comparable to cloud APIs.
Apple highlighted the work on its product page.
What happened
EXO Labs announced on August 25 that a cluster of four Mac Studios with Apple's new M5 Ultra chip achieves a total of about 4.8TB/s memory bandwidth. The company developed a low-latency RDMA (direct memory-to-memory data transfer) technology over Thunderbolt 5 in collaboration with Apple over one year.
Why it matters
This allows memory bandwidth to scale almost proportionally with the number of connected devices, letting large models like Kimi K3 and GLM-5.3 run locally at "API-like speeds". The M5 Ultra chip, announced the same day, features up to 36 CPU cores, up to 80 GPU cores, and 1.2TB/s of memory bandwidth per chip.
What to watch
The M5 Ultra Mac Studio starts at ¥949,800 (tax included), with pre-orders from August 25 and release on September 22. The 512GB memory model ships in late October.
Ask the AI about this article →
EXO Labs has been working with Apple for a year to develop a way to link Macs into an AI cluster. Their approach uses RDMA over Thunderbolt 5, which lets data move directly between memories without going through the OS or CPU. This is what allows the memory bandwidth of four machines to add up to roughly 4.8TB/s.
The M5 Ultra chip, also announced on the same day, provides up to 36 CPU cores and 80 GPU cores, with a memory bandwidth of 1.2TB/s per chip. By scaling bandwidth almost linearly with the number of devices, the cluster can run large language models like Kimi K3 and GLM-5.3 at what the company calls "API-like speeds" locally. Apple has also featured this capability on its Mac Studio product page, suggesting official recognition of the use case.
For business readers, the significance lies in running very large models on-premises rather than relying on cloud APIs. The main trade-off appears to be cost and hardware setup time, given the ¥949,800 starting price per machine. The late-October availability of the 512GB model suggests that maximum memory capacity may be the more relevant configuration for this type of workload.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
On August 25, 2026, OpenAI unveiled Jalapeño, its first custom inference chip, claiming 1.5–1.9x more AI work…

The Linux Foundation announced the contribution of TRACE (Trust, Runtime Attestation and Compliance Evidence)…

A new Stanford research paper warns that AI chatbots and agents may let advertising dollars influence what the…

The Wall Street Journal published an op-ed by billionaire investor Stanley Druckenmiller that was written with…

OpenAI unveiled benchmarks for its Jalapeño inference chip on Tuesday, showing it competes with Nvidia’s best…

Researchers from Bytedance Seed developed EdgeBench, a benchmark measuring how AIs improve on tasks over multi…
