AIToday
Large Language ModelsAI Business & IndustryZenn AI/MLPublished: Oct 2, 2026, 10:00 JST

Ryzen 5 5600G runs uncensored Qwen3.6 at 66.57 pp512

Ryzen 5 5600G runs uncensored Qwen3.6 at 66.57 pp512

3 Key Points

  1. What happened

    A Zenn author published a CPU-only recipe that quantizes Qwen3.6-35B-A3B-uncensored-heretic, reporting llama.cpp scores of 66.57 (pp512) and 13.53 (tg128) on a Ryzen 5 5600G.

  2. Why it matters

    The recipe lets people with 32GiB of RAM run a 35B uncensored model on a CPU alone, which may help those programming in air-gapped or isolated environments.

  3. What to watch

    Whether the setup works well hinges on the RAM, command-line skill, and CPU involved, since the author warns MTP hurts prompt processing on CPU.

WHO IT HITSDevelopers and tinkerers with 32GiB of RAM and basic cmd.exe or bash skills, especially those in air-gapped environments, can now try a 35B uncensored model without a GPU.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The article is written for people who want to run an LLM on a CPU, have at least 32GiB of RAM, and know cmd.exe or bash, and the author notes it may help with programming in isolated environments such as submarines.

The author starts from byteshape/Qwen3.6-35B-A3B-Q4_K_S-4.22bpw as a reference and uses Unsloth's imatrix as the importance source, then builds a custom quantization that imitates ByteShape's Q4_K_S-4.22bpw. A Makefile is provided for the quantization steps, and the resulting models are Qwen3.6-35B-A3B-uncensored-heretic-BS.gguf, the IK ​​variant, and the CPU-only repacked variant.

The author avoids MTP because it trades prompt processing for token generation and is not worth it on CPU, and also notes that the VSCode BYOK feature took about 10 minutes for a hello while the Zoo Code extension returned in about a minute. Whether this recipe is worth following likely hinges on the reader's CPU, RAM, and tolerance for the command-line steps involved.

FAQ
What hardware do I need to run this?
You need at least 32GiB of RAM and want to know cmd.exe or bash. The author's measurements use a Ryzen 5 5600G with 2ch DDR4-3200.
Which engine should I use for the uncensored model?
The author quantizes llmfan46/Qwen3.6-35B-A3B-uncensored-heretic with Unsloth's imatrix file and an IK variant, then runs it with ik_llama.cpp. The IK variant is CPU-only and repacked.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAI CEOs sign AI safety accord letting labs self-police, Trump says