
A fork of 1386.ai has been adapted to run on ROCm (a GPU compute platform for AMD hardware), targeting the AMD Strix Halo APU. The original author trained a 235M-parameter model; this port enables training a 500M-parameter model on the 128 GB Strix Halo APU in a GMKTec Evo X2 mini PC.
PyTorch's ROCm backend required virtually no model-specific code changes for training. The pipeline now uses torch.compile for performance, includes a Dockerfile to simplify ROCm installation, and changed training workers from 2 to 0 (running on the main thread) because training could not start with workers enabled.
Training a 500M-parameter model on this hardware takes roughly three weeks at ~4,750 tokens/s. The author notes there is likely not much low-hanging fruit left for optimization without writing custom CUDA kernels or deeper fused-operator optimizations.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Walmart settled opioid dispensing claims for $50 million

Tim Cook's legacy as Apple CEO is now tied to the company's push into artificial intelligence, according to a…

CrowdStrike is introducing Falcon Guardian, its flagship solution for the AI Detection and Response (AIDR) cat…

AT&T, Dell Technologies, and AMD have announced OTel 2.0, the largest and best-performing open-source model bu…

AT&T's legal department built an in-house center of expertise called Legal Edge, described as an AI-first lega…

John Deere introduced JD, an AI assistant designed to help farmers manage and interpret their farm data, as re…
