AIToday
Qiita 機械学習Published: Oct 3, 2026, 19:00 JST

One nn.Linear(784, 256) line, four chips: 200960 weights, four physical forms

One nn.Linear(784, 256) line, four chips: 200960 weights, four physical forms

3 Key Points

  1. What happened

    The article traces nn.Linear(784, 256) layer by layer, showing its 200704 weights plus 256 biases, roughly 785 KiB in FP32, flowing through CPU SIMD, GPU Tensor Core, FPGA DSP/BRAM, and ASIC MAC arrays.

  2. Why it matters

    The same mathematical parameters can look completely different in hardware, the article argues, so the real design question becomes where the 200704 weights sit and how they reach the MAC units, not just how fast multiply runs.

  3. What to watch

    The piece notes that scaling weight supply — 1000 MACs at 8 bits each need 8 Tbit/s at 1 GHz — is a bandwidth problem, so the test is whether reusing weights cuts memory traffic; it flags the FPGA speed-versus-area trade-off.

WHO IT HITSThis matters to hardware and AI accelerator engineers weighing CPU, GPU, FPGA, and ASIC targets, and to the teams that write FPGA or ASIC designs: their decisions on how many MAC units to build and where to store 200704 weights determine speed, power, and chip area.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The article starts from a single line of Python, nn.Linear(784, 256), and follows it down through four hardware worlds. Along the way it notes that the layer computes y = xWᵀ + b, and that the dominant operation underneath neural networks is multiply-and-accumulate (MAC): 784 multiply-adds repeated for each of 256 outputs. On a CPU, the weights begin in DRAM and travel through L3, L2, and L1 caches into registers before SIMD/FMA instructions do the work — the CPU has no dedicated Linear circuit at all, it just computes a large matrix product.

On a GPU, the same weights move over PCIe or unified memory into HBM or GDDR, then down through L2 cache and shared memory to CUDA or Tensor Cores, which process the multiplication as tiled GEMM. With an FPGA, the article explains, you can build the Linear circuit itself: weights go into BRAM or URAM (about 1.61 Mbit for the weights if quantized to INT8), DSP blocks become MAC units, and unchosen parts of the for-loop can be unfolded into space or pipelined so a new data item arrives every clock. An ASIC pushes further, letting designers place Weight SRAM right beside the MAC array and even eliminate the weight data entirely by turning a fixed value like 37 into bit shifts and additions, so the information lives in circuit structure rather than memory.

The article's unifying point is that the same mathematical parameter W can exist as 32-bit floats in DRAM, as HBM matrix data, as BRAM bits or logic, as dedicated SRAM, or even as a physical quantity in analog compute-in-memory. What the design hinges on, it suggests, is weight supply: when 1000 MACs each need 8 bits per cycle, that is 8 Tbit/s at 1 GHz, so reuse and data flow such as Weight Stationary are likely to matter as much as raw multiply speed for anyone building an AI accelerator.

FAQ
How many parameters does nn.Linear(784, 256) actually have?
It has 200704 weight values plus 256 bias values, for 200960 parameters in total. In FP32 that works out to about 785 KiB.
What is Weight Stationary in an AI chip?
It is a data flow where the weight stays inside the MAC unit and only the input values flow through. Keeping the weight in place reduces memory access, and ASICs can design this placement directly.
Can an FPGA really turn weights into circuit structure?
Yes, the article says. For fixed inference weights such as 37, the multiply can be rewritten as bit shifts and additions — (x<<5)+(x<<2)+x — so the data disappears and the value becomes part of the circuit itself.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Next articleAnthropic's Opus 5.5, OpenAI's Sol and Luna lead AI week