AIToday
Large Language ModelsOpen-Source AIGIGAZINE AIPublished: Oct 1, 2026, 16:00 JST

Hermes Agent's Qwen3.8 27B local AI runs on Intel Arc Pro B70

Hermes Agent's Qwen3.8 27B local AI runs on Intel Arc Pro B70

3 Key Points

  1. What happened

    After forcing Hermes Agent's backend to Vulkan with the command "hermes config set local_runtime.backend vulkan", GIGAZINE got Qwen3.8 27B running on 32GB of VRAM, and the smaller Q4_K_S version finished the article-summarization task without trouble.

  2. Why it matters

    The same summarization task had repeatedly failed on Solar Pro 4, so the local model handled it far more reliably.

  3. What to watch

    The test ran on a specific Intel Arc Pro B70 card, so results elsewhere may hinge on whether that GPU and the Vulkan setting are available. GIGAZINE says it will next try Bot Mode.

WHO IT HITSPeople running local AI agents on a single high-VRAM desktop GPU or workstation, and IT staff who must configure model backends, may find that the right quantization level and backend setting determine whether a task finishes quickly or fails.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

GIGAZINE had previously tasked Hermes Agent with fetching and summarizing its own articles, using models such as Nous Research's Hermes 4 70B and Solar Pro 4. Solar Pro 4 eventually managed the job, but only after struggling: it failed to launch the browser, switched methods, and opened the wrong articles along the way. The premise of this latest test was that 32GB of VRAM would allow Qwen3.8 27B to run at a far higher-quality quantization setting than before.

The test did not go smoothly at first. Hermes Agent loaded the model into ordinary RAM instead of the GPU, and the runtime listed its engine as "llama.cpp b10964 (cpu)" — the previous (cuda) label was gone. Settings also turned out to be profile-specific, so the Vulkan backend had to be set for the correct profile and the gateway restarted. Even after that, the standard 27B model used nearly 30GB of VRAM with over 4GB spilling into RAM, leaving inference at single-digit tokens per second.

The outcome suggests that for readers running AI agents on a single high-VRAM GPU, the usable speed may depend less on raw memory than on picking a quantization that fits entirely in VRAM and getting the backend configuration right. The author still plans to test Bot Mode next, which may show whether the same setup holds up for agentic browsing tasks.

FAQ
What did GIGAZINE have to change to make the local model use the GPU?
It set the backend to Vulkan with the command "hermes config set local_runtime.backend vulkan". Without that setting, Hermes Agent loaded the model onto the CPU instead of the GPU.
How did the Q4_K_S version compare with the standard 27B model?
The standard Qwen3.8 27B used Q4_K_M internally, and part of it spilled into regular RAM, so inference was slow. The Q4_K_S version used 26GB of VRAM and ran much faster.
Did the local model handle the GIGAZINE summarization task well?
Yes, Qwen3.8 27B completed the summary without any trouble, while Solar Pro 4 had repeatedly failed. On a follow-up request it started output in about 10 seconds.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleMetaview raises $60 million, launches Slack recruiting agent fillmore