technology

Nvidia shelves rtx 50 super, races to 1.6 nm 'feynman' gpus

NVIDIA has quietly buried the annual mid-cycle refresh. No RTX 50 Super cards will appear this spring, breaking a cadence the market has treated as scripture since the GTX 16 series. Instead, Jensen Huang will step onto the GTC stage on 16 March 2026 and unwrap a die labeled Feynman: the first GPU architecture printed on TSMC’s 1.6 nm node, a geometry so tight that electrons are measured in dozens.

Why the calendar moved

Inside Santa Clara the math became brutal. A Blackwell shrink to 3 nm would deliver barely 8 % more frames per watt while board power already tickles 600 W. Rather than sell a hotter, pricier SKU to an audience that just shelled out for RTX 5090s, the firm is hoarding its wafers for a two-node leap that, on paper, lifts efficiency 23 % and doubles transistor density. The risk: leaving 2025 open to AMD’s RDNA 4 and Intel’s Battlemage refresh. The bet: no rival can match a 1.6 nm monster clocked above 3 GHz.

One kilowatt per socket

One kilowatt per socket

Feynman is not a single chip; it is a dual-reticle slab paired through Intel’s EMIB-T bridge, a packaging move that sidesteps TSMC’s overloaded CoWo-S lines. Each package can pull 1 000 W; stack two on a workstation board and you hit 2 kW before the monitor warms up. NVIDIA’s answer is a cold plate the size of a hard-back book and a mandatory 240 V outlet. Data-centre customers are being told to budget one extra kilowatt per rack slot—power budgets normally reserved for AI training nodes, not pixels.

The names after blackwell

The names after blackwell

Roadmap slides obtained by Coastal Code show the succession line: Blackwell → Rubin (late 2025) → Feynman (2028). Rubin will keep the 3/4 nm split and act as a stop-gap for hyperscalers. Feynman introduces TSMC’s A16 node with back-side power rails, a trick that shoves resistive loss off the critical path. The result: 1.2× frequency at equal voltage, or 0.85× voltage at equal clocks. Gamers care about frames; NVIDIA cares that every millivolt saved multiplies across a 50 000-GPU cluster burning midnight oil on Llama-10.

Memory moves to 36 Gbps GDDR8, stacking eight 3 GB dice per controller for 96 GB on the flagship card. The bus remains 512-bit; bandwidth tops 2.3 TB/s, enough to feed 200 TFLOPS of shader math without the cache misses that currently hobble RTX 4090-class silicon in path-traced titles.

Heat death and the electricity bill

Heat death and the electricity bill

Early Feynman prototypes at 1 800 MHz already flirt with 150 °C junction temps. The company’s fix is a vapor chamber milled directly into the package substrate, channeling coolant to within 0.2 mm of the hot spots. Board partners are testing external radiators that mount where 5.25-inch bays used to live—PC cases are becoming small chillers. Expect a $200 cooling tax on top of a $1 999 MSRP. For AI farms, NVIDIA will sell a bare OAM module; hyperscalers slot it into their own glycol loops.

What it means for your wallet today

What it means for your wallet today

With no Super refresh coming, RTX 5090 inventory stays frozen at $1 599. Retailers who expected a February price cut are now stuck holding chips that won’t be refreshed for three years. Used-market prices for RTX 4090s actually ticked up 7 % last week—gamers hedging against a 2025 vacuum. AMD, meanwhile, is whispering “$499” for its RDNA 4 top card, a spoiler aimed squarely at the gap NVIDIA just punched in its own line-up.

The cruel irony: Feynman’s 1.6 nm wafers will be so scarce that Huang is likely to prioritize Grace-Next CPUs and AI GPUs over GeForce. Gamers may drool at 200 TFLOPS, but most units will never leave the server room. By 2029 the graphic card you actually buy could be a cut-down Feynman die binned for 350 W, water-cooled and still slower than the data-centre variant.

NVIDIA just told the world to skip a generation. The question is whether the world will wait—or whether AMD and Intel finally land a punch while the champ reloads.