facebook pixel
Four architectures. Four different bets about what computation actually looks like. The gap between them is not clock speed. It is assumption. A CPU assumes the work is unpredictable. It spends its transistor budget on machinery for guessing: branch predictors, out-of-order execution engines, deep cache hierarchies, speculative loads. Only a small fraction of a modern core is arithmetic. The rest exists to keep that arithmetic fed when the instruction stream refuses to behave. Superb at serial, branch-heavy code. Mediocre at bulk mathematics. A GPU assumes the opposite. The work is regular, parallel and enormous. Thousands of simple lanes execute in lockstep under one instruction stream. Latency is not avoided, it is hidden: when one warp stalls on memory, another is swapped in almost immediately. An H100 carries 16,896 shader cores and roughly 3 TB/s of HBM3 bandwidth to keep them fed. A TPU narrows further. Its heart is a systolic array, a grid of multiply-accumulate cells through...

 6.5k

 169

 3

 6.5k

    Suggested Credits
    Tags, Events, and Projects