
What happened
After forcing Hermes Agent's backend to Vulkan with the command "hermes config set local_runtime.backend vulkan", GIGAZINE got Qwen3.8 27B running on 32GB of VRAM, and the smaller Q4_K_S version finished the article-summarization task without trouble.
Why it matters
The same summarization task had repeatedly failed on Solar Pro 4, so the local model handled it far more reliably.
What to watch
The test ran on a specific Intel Arc Pro B70 card, so results elsewhere may hinge on whether that GPU and the Vulkan setting are available. GIGAZINE says it will next try Bot Mode.
WHO IT HITSPeople running local AI agents on a single high-VRAM desktop GPU or workstation, and IT staff who must configure model backends, may find that the right quantization level and backend setting determine whether a task finishes quickly or fails.
Summaries like this, in your inbox every morning.
GIGAZINE had previously tasked Hermes Agent with fetching and summarizing its own articles, using models such as Nous Research's Hermes 4 70B and Solar Pro 4. Solar Pro 4 eventually managed the job, but only after struggling: it failed to launch the browser, switched methods, and opened the wrong articles along the way. The premise of this latest test was that 32GB of VRAM would allow Qwen3.8 27B to run at a far higher-quality quantization setting than before.
The test did not go smoothly at first. Hermes Agent loaded the model into ordinary RAM instead of the GPU, and the runtime listed its engine as "llama.cpp b10964 (cpu)" — the previous (cuda) label was gone. Settings also turned out to be profile-specific, so the Vulkan backend had to be set for the correct profile and the gateway restarted. Even after that, the standard 27B model used nearly 30GB of VRAM with over 4GB spilling into RAM, leaving inference at single-digit tokens per second.
The outcome suggests that for readers running AI agents on a single high-VRAM GPU, the usable speed may depend less on raw memory than on picking a quantization that fits entirely in VRAM and getting the backend configuration right. The author still plans to test Bot Mode next, which may show whether the same setup holds up for agentic browsing tasks.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Google DeepMind launched Gemini 4 Argon for coding, enterprise knowledge work and cyber defense, claiming firs…

Bloomberg reports Amazon's delivery smart glasses shoot still images at intervals during walks, possibly thous…

The Information reports Google's "AI Contribution Pilot Program" pays about 100 digital publishers, including…

Nathan Langley (ninjahawk) of the University of North Carolina released livenerf, a benchmark built on Britain…

Huntress found attackers naming a Custom GPT "Plus 5.6" on the real chatgpt.com, sometimes reached via sponsor…

Metaview Labs Inc. closed $60 million in Series C funding led by Insight Partners, bringing total funding to $…