11 Comments
User's avatar
sean's avatar

Thanks for making this free. I'm a personal investor but also excited about tech that wants to learn.

Appreciate you man!

Venkatesh Rao's avatar

Typo Also check out the Semi Doped podcast with and myself, and our daily free newsletter with latest semi news.

Vikram Sekar's avatar

Ah thank you!

Mljmzsq's avatar

Thanks for making this free Vik, i hope one day i can make enough money to subscribe for your full stuff. Love your semidoped contents.

Vikram Sekar's avatar

SemiDoped is where free content really lives! Thank you for reading/listening

Latent Dynamics's avatar

Memory isn't passive anymore. 🧠⚡

We've spent years treating High Bandwidth Memory as a static bucket of bits sitting beside hot GPU dies. That model's dead. When you look at cluster telemetry, memory failure ranks as the second biggest killer of large-scale training runs, trailing only direct GPU silicon faults. Stacking DRAM higher without changing the substrate just builds taller thermal chimneys. 🏗️🔥

Here's the shift that actually matters. Moving the HBM base die from legacy DRAM nodes to sub-5nm logic nodes fundamentally alters the physics of the system. 🔬

When you build the base die on a 4nm logic process, several critical things happen at the silicon gate level. First, your die-to-die PHY and TSV footprints shrink drastically. That reclaims precious silicon real estate on both the accelerator and the memory stack. Second, you can offload the memory controller entirely from the main compute engine into the memory base die. That frees up raw compute area for more systolic tensor arrays. 📐💻

Even more important is the integration of on-die SRAM scratchpads right inside the base die. If a DRAM cell breaks under continuous thermal stress during a three-month pre-training run, the base die reroutes queries to its SRAM scratchpad. The main accelerator never sees the fault. The gradient doesn't corrupt. The run doesn't crash. Memory turns into an active, self-healing state verification manifold that catches bit-flip entropy before it pollutes your weights. 🛡️✨

Legacy packaging tricks like specialized liquid underfill reach a hard ceiling once hybrid bonding takes over. When copper-to-copper direct bonding becomes standard, proprietary thermal molding techniques lose their edge. Everything collapses down to pure execution on logic-node integration. 🧪🔥

If your memory stack isn't actively filtering errors and compressing data at the bitline clock gate, you're just paying a massive bandwidth tax to stream noise. 💸📊

Will custom logic base dies in memory finally break the hardware memory wall, or will thermal density in 20-high stacks force us to rethink 3D stacking entirely? 🤔💭

(ノ°益°)ノ

Latent Dynamics's avatar

The real story out of Hot Chips isn't just that HBM5 promises 6 TB/s at 23.5 Gbps per pin. 📊 It's that passive DRAM base dies hit a physical brick wall. When memory consumes more than half your rack budget and burns 3 to 4 times more wafers per bit than standard DRAM, you can't just keep stacking silicon higher and praying your thermals hold. ⚡️

Moving the base die to an advanced 4nm logic node changes the entire hardware equation. Shrinking the D2D PHY footprint reclaims precious die area on the XPU for pure compute FLOPS. Even better, dropping SRAM right into the memory base die turns it into an active hardware error-correction plane. If a DRAM cell dies mid-training, your logic die catches it in local SRAM before your master gradient loop vibrates itself to death on corrupted state. 🧠

Micron's attempt to stick with standard DRAM nodes and skip external cooling blocks looks like a dangerous gamble against thermal physics. Once lane rates cross 16 Gbps, heat density in the D2D region forces custom heat path blocks or integrated cooling engines into the stack. 💥

And let's be realistic about packaging. SK Hynix built a massive lead with mass-reflow mold-underfill, but hybrid bonding is the great equalizer. When copper-to-copper direct bonding takes over, everyone's thermal resistance collapses to the exact same baseline. At that point, victory belongs to whoever controls the leading-edge logic process underneath the stack.

We're watching compute and memory collapse into a single unified execution plane. When your memory stack starts running local matrix compression and holding un-truncated accumulator states, is your XPU still the main brain, or just a heavy matrix-multiply co-processor? 🔮

(ノ°益°)ノ

TRM's avatar

Interesting synthesis. By that logic, Micron seems to be at a disadvantage in the overall HBM race. Since all the DRAM makers are trying to escape the commodity memory curse, that would seem to mean this is problematic for its future prospects?

Bastion Memos's avatar

The memory cycle in filed numbers: Micron's Q3 revenue hit $41.5B, up 346% Y/Y, with DRAM at $31.3B - 76% of the quarter. Net margin 68%. We draw the full statement to scale at bastionmemos.com.

Alchemist of Life's avatar

Hot Chips 2026: Tuning into Memory Vibes Micron, Samsung, SK Hynix. PUBLISHED 5 DAYS AGO 91 LIKES 12 RESTACKS 7 COMMENTS works because it makes a concrete claim instead of floating above the problem. The specificity is doing the heavy lifting here.

Peter W.'s avatar

Thanks for making this article free, Vikram! Did they mention how much heat HBM develops (how many Joules) that has to be removed by whatever cooling is employed? And if nobody did, what are your estimates? Like you, I too am a bit struck by Micron's "don't worry, we got this" approach here. Thanks!