
What happened
The article traces nn.Linear(784, 256) layer by layer, showing its 200704 weights plus 256 biases, roughly 785 KiB in FP32, flowing through CPU SIMD, GPU Tensor Core, FPGA DSP/BRAM, and ASIC MAC arrays.
Why it matters
The same mathematical parameters can look completely different in hardware, the article argues, so the real design question becomes where the 200704 weights sit and how they reach the MAC units, not just how fast multiply runs.
What to watch
The piece notes that scaling weight supply — 1000 MACs at 8 bits each need 8 Tbit/s at 1 GHz — is a bandwidth problem, so the test is whether reusing weights cuts memory traffic; it flags the FPGA speed-versus-area trade-off.
WHO IT HITSThis matters to hardware and AI accelerator engineers weighing CPU, GPU, FPGA, and ASIC targets, and to the teams that write FPGA or ASIC designs: their decisions on how many MAC units to build and where to store 200704 weights determine speed, power, and chip area.
Summaries like this, in your inbox every morning.
The article starts from a single line of Python, nn.Linear(784, 256), and follows it down through four hardware worlds. Along the way it notes that the layer computes y = xWᵀ + b, and that the dominant operation underneath neural networks is multiply-and-accumulate (MAC): 784 multiply-adds repeated for each of 256 outputs. On a CPU, the weights begin in DRAM and travel through L3, L2, and L1 caches into registers before SIMD/FMA instructions do the work — the CPU has no dedicated Linear circuit at all, it just computes a large matrix product.
On a GPU, the same weights move over PCIe or unified memory into HBM or GDDR, then down through L2 cache and shared memory to CUDA or Tensor Cores, which process the multiplication as tiled GEMM. With an FPGA, the article explains, you can build the Linear circuit itself: weights go into BRAM or URAM (about 1.61 Mbit for the weights if quantized to INT8), DSP blocks become MAC units, and unchosen parts of the for-loop can be unfolded into space or pipelined so a new data item arrives every clock. An ASIC pushes further, letting designers place Weight SRAM right beside the MAC array and even eliminate the weight data entirely by turning a fixed value like 37 into bit shifts and additions, so the information lives in circuit structure rather than memory.
The article's unifying point is that the same mathematical parameter W can exist as 32-bit floats in DRAM, as HBM matrix data, as BRAM bits or logic, as dedicated SRAM, or even as a physical quantity in analog compute-in-memory. What the design hinges on, it suggests, is weight supply: when 1000 MACs each need 8 bits per cycle, that is 8 Tbit/s at 1 GHz, so reuse and data flow such as Weight Stationary are likely to matter as much as raw multiply speed for anyone building an AI accelerator.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.