AIToday
Large Language ModelsOpen-Source AIQiita 機械学習Published: Oct 4, 2026, 10:00 JST

Rai runs AI on CPU alone, no GPU or Python needed

Rai runs AI on CPU alone, no GPU or Python needed

3 Key Points

  1. What happened

    A developer published Rai, a small Rust engine that runs language models on an ordinary PC's CPU, using only the machine's built-in AVX2 instructions and no external libraries.

  2. Why it matters

    Because Rai bundles the whole calculation into one readable program, it can be used on machines that are not connected externally or in settings where all the code must be inspected, rather than relying on a chain of prepared parts.

  3. What to watch

    Rai's selling point is not speed, and llama.cpp is more mature on that front, so its practical value hinges on whether fewer parts and readable code matter more than raw performance for a given use case.

WHO IT HITSDevelopers and IT teams working on offline or security-sensitive machines, who need to run language models without pulling in large external libraries, are the ones most likely to benefit.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Running AI on an ordinary PC has usually meant preparing a GPU or stacking layers of ready-made parts such as PyTorch, which makes installation heavy and the inner workings hard to follow. Rai takes the opposite approach: its author wrote the calculation code by hand in Rust and relies only on AVX2, a calculation feature already present in the computer. That places it in contrast with llama.cpp, a well-known CPU-only tool that still borrows some external calculation libraries.

The models Rai runs are compressed to 4 bits, a technique called quantization that rounds the model's internal numbers to make them lighter. According to the author, this can reduce accuracy by 13–30% in some situations. The design favors having few parts and being able to read the whole program, rather than raw speed, which the author notes is where llama.cpp is more mature.

A practical limit is that Rai is fast on Intel and AMD PCs but falls onto a slow path on Macs with Apple Silicon. It was also only shared on Hacker News on October 2, 2026, so it is still a young tool. Whether it finds a following may depend on how much users in offline or security-sensitive settings value a small, fully readable engine over a faster but more layered one.

FAQ
What models can Rai run?
It supports a wide range of models, including Llama, Qwen, and Gemma. Models are run after being compressed to 4-bit through a process called quantization.
How fast is Rai?
A small TinyLlama (1.1 billion parameters) runs at about 22 tokens per second on an ordinary 4-core laptop. A larger 7-billion-parameter model runs at around 3 tokens per second.
What are the requirements to try Rai?
You need Rust 1.87 or later. Rai is fast on Intel and AMD PCs, but it runs slowly on Macs with Apple Silicon.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleMinecraft AI bot plan hits 月数円 with Jev, Gemini, Mineflayer