← Back to list

Thermal Punk: A Heterogeneous Architecture for Post-Silicon AI Compute

Abstract

lukas langsadl · 2026-03-05 21:37 · 0 claps · 13.9 min read
#ai #computes #ai-slop
Open on Medium ↗
Wiki topics: AI · AI · General TLS · Design Tools & Workflow 🏛️ · Architecture

Thermal Punk: A Heterogeneous Architecture for Post-Silicon AI Compute

Abstract

The global AI compute stack is a monoculture begging to be disrupted. One company (TSMC) fabricates >90% of the world’s sub-7nm chips, on an island 180 km from a hostile military. One vendor (Nvidia) owns 80–90% of the AI accelerator market, shipping GPUs that now draw 1,200 watts apiece and climbing. The IEA projects data center electricity consumption will hit 945 TWh by 2030 — more than Japan uses today. We are building the century’s most important technology on a single-point-of-failure supply chain and powering it by boiling the ocean faster.

This paper proposes an alternative. Not a silicon replacement. A heterogeneous hybrid architecture that assigns each computational task to the physical substrate whose native physics best matches the required mathematics: (1) a photonic compute plane for linear algebra at the speed of light; (2) an entropic processing layer that turns thermal noise from enemy to computational resource for sampling and optimization; (3) a neuromorphic routing fabric for sparse, event-driven activation management; and (4) a classical silicon control plane for orchestration and I/O.

No single alternative paradigm solves all problems. Photonic processors lack nonlinearity and density. Thermodynamic computers lack deterministic precision. Neuromorphic chips cannot do dense matrix math. But each solves one problem extraordinarily well, and the problems are complementary. Three of the four substrates fabricate at mature semiconductor nodes (45–130nm) — no EUV lithography, no TSMC dependency. This is a technical architecture, a geopolitical hedge, and a declaration that the future of computation need not be a monoculture.

Every claim is flagged: demonstrated, theoretically sound but undemonstrated, or speculative.

I. Photonic compute plane

The physics

Matrix-vector multiplication dominates AI workloads (>90% of FLOPs). Photonic processors execute it using optical interference, with near-zero energy for the multiply itself.

The architecture: a Mach-Zehnder interferometer (MZI) mesh network. Decompose matrix W = UΣV† via SVD. U and V† are unitary matrices implemented as cascaded MZI arrays (Reck et al. 1994, improved by Clements et al. 2016). Σ is a diagonal of attenuators. The product computes in a single pass of light through the mesh — sub-nanosecond latency per multiply, energy approaching zero for the optical operation itself.

Shen et al. (2017, Nature Photonics) demonstrated this: a 56-MZI programmable nanophotonic processor doing vowel recognition. They projected 100× speed and ~1,000× power efficiency over electronics for inference. The work proved the physics. It did not prove it scales.

The density problem (real and fundamental)

Telecom-band light at λ=1,550nm in silicon yields a diffraction-limited feature size of ~220nm — two orders of magnitude larger than a 3nm transistor. A single MZI occupies ~100–1,000 μm², giving photonic circuits 1⁰⁴–1⁰⁵ fewer components per mm² than electronic ICs. This is not an engineering gap. It is physics. Photonics compensates through wavelength-division multiplexing (16+ parallel computations on one waveguide), picosecond latency, and sub-picojoule energy per operation. The density disadvantage is real; the throughput and energy advantages are also real.

What’s shipping now

Lightmatter ($4.4B valuation, October 2024 Series D) builds Envise photonic processors with MEMS-based phase shifters achieving 8-bit precision, and the Passage M1000 — a multi-reticle photonic interposer exceeding 4,000 mm² with 114 Tbps optical bandwidth using 16-wavelength bidirectional DWDM. Celestial AI (acquired by Marvell for $3.25B upfront, December 2025, up to $5.5B with milestones) developed a “Photonic Fabric” claiming 25× bandwidth at 10× lower latency than co-packaged optics. Luminous Computing — once promising at $105M raised — laid off its photonics team in May 2023. The gap between physics demos and engineering reality eats startups. [Demonstrated: MZI mesh physics. Engineering phase: commercial-scale photonic AI acceleration.]

The quantum enhancement question (let’s be honest)

Here is where we stop stitching citations together and tell you what nobody has actually shown.

Classical photonic processors face a precision wall: shot noise. Encoding matrix elements as optical amplitudes, measurement precision scales as Δφ ~ 1/√N for N photons. For 8-bit precision (the current sweet spot for inference), this is fine. For 16-bit training accumulation, it becomes a fundamental barrier.

Caves (1981) showed that squeezed vacuum states can push interferometer sensitivity below the standard quantum limit. LIGO has operationalized this since 2019, using frequency-dependent squeezing with 300-meter filter cavities to increase observable volume by 60%. Giovannetti, Lloyd, and Maccone (2004, Science) established the Heisenberg limit: N entangled photons achieve Δφ ~ 1/N — a quadratic improvement over classical. Braunstein and van Loock (2005, Reviews of Modern Physics) provided the continuous-variable quantum computing framework. Arrazola et al. (2021, Nature) and Madsen et al. (2022) demonstrated Gaussian boson sampling at scale.

Now here’s the gap. LIGO operates in cryogenic ultra-high vacuum with 300-meter filter cavities, measuring a single phase parameter with exquisite control. We are proposing to apply similar quantum metrological techniques to thousands of simultaneous computations in a photonic mesh operating at 85°C in a data center. Nobody has demonstrated this. The physics of squeezing is sound. The physics of Heisenberg-limited computation is sound. The engineering path from LIGO’s single-parameter measurement to multi-mode, room-temperature, commercial-speed precision enhancement in a photonic matrix multiplier does not exist yet. It is not clear it can exist without cryogenics, which would destroy much of the energy and cost advantage.

What we actually propose: a photonic plane that operates primarily in classical mode at 4–8 bit precision (sufficient for inference). Quantum enhancement via squeezed-light injection is a research target for precision-critical operations (gradient accumulation, attention scores), not a baseline assumption. If it works, it’s transformative. If it doesn’t, the classical photonic plane is still viable for inference. We are not betting the architecture on it. [Demonstrated: squeezing in metrology. Theoretically sound: application to compute. Undemonstrated: everything in between.]

Fabrication (this is the strategic punchline)

Photonic circuits fabricate at 45–90nm on standard silicon photonics platforms. GlobalFoundries’ Fotonix platform operates at 90nm-class on 300mm wafers. Tower Semiconductor’s PH18 platform generates >$220M annual revenue in silicon photonics with 70% year-over-year growth. LIGENTEC and LioniX provide silicon nitride with <0.5 dB/m propagation loss. None of this requires EUV. None of it depends on TSMC. [Demonstrated and commercially shipping.]

II. Entropic processing layer

The insight that should make you angry

Modern AI’s most expensive frontier is probabilistic. Diffusion models generate images by iteratively denoising random noise. LLMs sample from probability distributions at every token. Bayesian networks maintain uncertainty estimates. All of this requires stochasticity — and all of it currently simulates stochasticity on deterministic hardware.

This is insane. We are spending 3,500× the thermodynamic minimum energy per switching event (Landauer’s kT ln(2) ≈ 2.87 × 10⁻²¹ J at room temperature; a 22nm Intel transistor dissipates ~10 fJ) to generate deterministic signals, then spending additional energy to make them random again. Bérut et al. (2012, Nature) verified Landauer’s principle experimentally. Bennett (1973) proved computation can be logically reversible, avoiding the penalty entirely. We have known this for fifty years. We are still doing it backwards.

P-bits: the physics does the sampling for free

Camsari, Faria, Sutton, and Datta (2017, Physical Review X) formalized probabilistic bits (p-bits): stochastic magnetic tunnel junctions with energy barriers of order ~kT, fluctuating between states via thermal noise at room temperature. The output follows P(mᵢ = 1) = σ(I₀ + Σⱼ Jᵢⱼmⱼ) — exactly the conditional probability in Gibbs sampling on a Boltzmann machine. A network of interacting p-bits samples from the Boltzmann distribution naturally. No pseudo-random number generation. No MCMC approximation. The physics does the sampling.

Borders et al. (2019, Nature) demonstrated eight spintronic p-bits performing integer factorization at room temperature. Singh et al. (2024, Nature Communications) showed CMOS+sMTJ integration for probabilistic inference — asynchronous, no clock required. Energy: ~2 fJ per random bit, replacing ~10,000 CMOS transistors at two orders of magnitude less energy. The Tohoku group (Fukami et al., 2025) has extended this to Gaussian p-bits (g-bits) for continuous distributions, enabling hardware-native Gaussian-Bernoulli Boltzmann Machines directly applicable to diffusion model acceleration.

The diffusion connection is not metaphorical

Sohl-Dickstein et al. (2015) titled their paper “Deep Unsupervised Learning using Nonequilibrium Thermodynamics” because that is literally what it describes: forward diffusion destroys structure by adding noise (thermodynamic equilibration); the learned reverse process generates structure from noise (time-reversed Langevin dynamics). Ho et al. (2020) made it practical. The thermodynamic interpretation was always the foundation, not the metaphor.

Thermodynamic hardware makes this literal. Extropic (founded 2022, $14.1M seed, ex-Google quantum researchers) builds Thermodynamic Sampling Units using subthreshold CMOS where thermal noise is the computation. Their Denoising Thermodynamic Models chain energy-based models that denoise data directly in hardware. They claim ~10,000× energy reduction for image generation — a number that is currently based on their own benchmarks and requires independent verification. Normal Computing (ex-Google Brain) unveiled the first thermodynamic computer prototype in January 2024, publishing results in Nature Communications (2025): an 8-unit stochastic processing unit demonstrating Gaussian sampling and matrix inversion.

Hinton and Sejnowski (1983) built Boltzmann machines on this exact mathematical framework. Hinton won the 2024 Nobel Prize for it. The entropic processing layer closes the circle: the mathematics Hinton simulated on deterministic silicon can now run natively on hardware where Boltzmann’s distribution is physical fact, not computational approximation. [Demonstrated: p-bit physics at small scale. Demonstrated: thermodynamic computing prototypes. Undemonstrated: integration at AI-relevant scale. Speculative: energy efficiency claims beyond small benchmarks.]

III. Neuromorphic routing fabric

The right job for the wrong chip

Neuromorphic computing has spent three decades failing to become general-purpose AI hardware. Good. We don’t want it for general-purpose AI hardware. We want it for the one thing it does better than any other substrate: spending energy only when information changes.

Intel’s Loihi 2 (128 neuromorphic cores, up to 1M neurons per chip) achieves 52× energy reduction over a Jetson Nano GPU on CIFAR-10 inference and 3,500× reduction on vibration anomaly detection. Davies et al. (2018, IEEE Micro) documented >1,000× superior energy-delay product over CPU solvers. The Hala Point system scales to 1,152 chips and 1.15 billion neurons. SpiNNaker 2 (22nm FDSOI) claims 78× efficiency over GPUs for event-based processing. BrainScaleS-2 (65nm CMOS, Heidelberg) emulates neural dynamics at 1,000× biological real-time. The DYNAP-SE2 (Zurich, Richter et al. 2024) runs at sub-milliwatt power.

But none of these can do dense matrix multiplication competitively. The software ecosystems are immature. Training spiking networks is harder than training conventional ANNs. This is well-documented failure, and we are not proposing to fix it.

What we actually want: MoE routing

Mixture-of-Experts is now the dominant paradigm for frontier models. DeepSeek-V3: 671B total parameters, 37B active per token, routing through 256 fine-grained experts with top-8 selection. The routing function is a lightweight linear layer (~1M parameters) followed by softmax and top-K — a sparse, low-dimensional classification decision on the critical path. Experts cannot compute until the router decides which ones to activate.

This routing decision must be ultra-low latency, ultra-low power, and adaptive. Neuromorphic hardware offers all three. Spike-timing-dependent plasticity (STDP) enables routing functions that adapt online to shifting data distributions without backpropagation. Frequently used pathways strengthen; unused pathways prune. The biological analogue is homeostatic plasticity — which neuromorphic hardware implements natively.

The neuromorphic fabric handles three functions: MoE router decisions, sparse activation management (tracking which portions of other processors are in use), and cache hierarchy management (predicting access patterns via STDP-learned temporal correlations). For all three, energy scales with information change — exactly where neuromorphic excels.

[Demonstrated: neuromorphic energy efficiency on sparse workloads. Architecturally sound: application to MoE routing. Speculative: STDP-based adaptive routing for production AI workloads.]

IV. Classical silicon control plane

The system still needs conventional silicon for orchestration, error correction, external I/O, and compilation. This is commodity hardware — advanced nodes for single-thread performance but without the extreme area demands of GPU arrays. The control plane implements a compiler targeting all four substrates using MLIR’s extensible dialect system, with custom dialects for photonic operations (unitary decomposition, WDM scheduling), entropic operations (energy landscape specification, Boltzmann sampling), and neuromorphic operations (spike-train encoding, STDP rules). It maintains a real-time cost model for each substrate and dynamically routes computation to minimize a composite objective.

The precedent: CPU+GPU (CUDA), CPU+FPGA (Intel Stratix), AMD APUs. In each case the classical processor orchestrated; it never did the bulk compute. Same here. [Demonstrated for CPU+GPU+FPGA. Extension to four exotic substrates: requires significant compiler development.]

V. Integration architecture

This is where it gets hard

Every domain boundary costs energy: E/O conversion at 0.7–5 pJ/bit, O/E at 0.5–2 pJ/bit, ADC at 1–10 mW per GHz, spike encoding/decoding at neuromorphic boundaries. A full roundtrip costs 2–10 pJ/bit — comparable to DRAM access energy. If the computation-to-communication ratio is too low, the heterogeneous architecture loses to a GPU that never converts between physical domains.

The architecture therefore imposes coarse-grained partitioning: photonic execution handles entire matrix multiplies (millions of MACs per domain crossing); entropic processing handles complete sampling chains (hundreds of Gibbs steps per invocation); neuromorphic routing makes per-layer decisions, not per-token. This is not instruction-level heterogeneity. It operates at the granularity of neural network layers.

Interconnect technology is maturing fast. Ayar Labs’ TeraPHY (3rd gen, 2025) delivers 8 Tbps bidirectional as a UCIe-compliant optical chiplet on GlobalFoundries 45nm. Intel’s OCI chiplet demonstrated 4 Tbps at OFC 2024. Lightmatter’s Passage M1000 claims >200 Tbps total I/O. UCIe 3.0 (ratified August 2025) specifies up to 1.3 TB/s per mm of die edge at 0.01 pJ/bit in advanced 3D packaging. These are on commercial roadmaps.

Physical integration: a multi-layer chiplet assembly on a silicon interposer. Photonic chiplets (GlobalFoundries 45nm) for optical I/O and compute. Spintronic p-bit arrays (MRAM-compatible) as separate chiplets via electrical UCIe. Neuromorphic chiplets (45–65nm, asynchronous) connected via Address-Event Representation protocol. Classical control die on the interposer. Thermal zoning is critical: photonic micro-ring resonators shift ~0.1 nm/°C and need thermal isolation from high-power CMOS. The entropic layer operates at ambient temperature by design — thermal noise is the resource. [Demonstrated: photonic + CMOS chiplet integration. Speculative and unprecedented: full four-substrate chiplet integration.]

VI. Fabrication diversity as geopolitical strategy

The strategic argument in one table:

Substrate Node required EUV needed? Available fabs Photonic compute 45–90nm No GlobalFoundries, Tower, IMEC/UMC, LIGENTEC, LioniX Entropic/spintronic 28–65nm MRAM-compatible No Research fabs globally; MRAM lines at Samsung, GF, TSMC mature nodes Neuromorphic 22–180nm No SpiNNaker 2 at 22nm FDSOI (GF); BrainScaleS at 65nm; DYNAP-SE2 at 180nm Classical control 3–7nm (small die area) Yes, but low volume Any advanced foundry

Three of four substrates can be manufactured in dozens of fabs across the US, EU, Japan, South Korea, and Southeast Asia. A mature-node fab costs $1–5B versus $20–50B for a leading-edge EUV fab. The capital barrier drops by an order of magnitude. The European Chips Act (€43B) explicitly provisions photonic ICs, advanced packaging, and neuromorphic chips as strategic technologies. GlobalFoundries has invested $700M specifically in silicon photonics at its Malta, NY facility. Tower Semiconductor is investing $650M to triple silicon photonics shipments by mid-2026.

The concentration of AI hardware manufacturing in a single geopolitical flashpoint is a civilizational risk. An architecture that distributes fabrication across multiple substrates, nodes, and geographies is inherently more resilient than one that concentrates it in three fabs on one island. [Demonstrated: fab capabilities for each substrate independently. Does not exist today: integrated supply chain for heterogeneous AI compute.]

VII. What could kill this

Interface losses compound faster than optimism. A photonic matrix multiply costs ~0.1 pJ/bit for the optical operation. E/O + O/E conversion costs 2–10 pJ/bit. The interface dominates the energy budget by 20–100×. If computation-to-communication ratios are insufficient, the heterogeneous system loses to monolithic GPUs that never cross domain boundaries. The minimum viable granularity is probably entire transformer layers or larger.

Precision is a non-negotiable. Photonic processors currently achieve 4–8 bit effective precision. Modern inference uses INT4/FP8, which is within reach. Training requires BF16/FP32 accumulation, which is not. Quantum precision enhancement is theoretically sound but experimentally undemonstrated for computing — and the LIGO-to-data-center gap is enormous. Entropic computing outputs are stochastic by nature; converting to deterministic results requires statistical aggregation over multiple samples, adding latency. For applications where stochasticity is the output, this is fine. For exact answers, the entropic layer is the wrong tool.

The software ecosystem gap dwarfs the hardware challenge. CUDA has 4+ million developers. The entire neuromorphic ecosystem has maybe a few thousand. There is no compiler that can partition a PyTorch model across photonic, entropic, neuromorphic, and silicon substrates. MLIR provides the theoretical framework. Writing the actual compiler passes represents years of focused engineering. History shows software ecosystems are harder to build than hardware: OpenCL was technically superior to CUDA in generality and lost the ecosystem battle.

The brutal GPU counterfactual. The H100 delivers ~1.4 × 1⁰¹² FLOP/J. CMOS has roughly 200× headroom before hitting fundamental limits. Nvidia delivers 2–3× per generation on 2-year cycles. A 10–15 year development timeline means competing against GPUs that are themselves 10–30× more efficient. The architecture bets the wall is real and approaching — power per GPU doubled in one generation (H100 700W → Blackwell Ultra 1,400W). If Moore’s Law finds another decade of miracles, this arrives too late. [All risks grounded in current engineering reality.]

VIII. Research roadmap

Phase 1 (~€2–4B): Individual substrates at AI-relevant scale. 64×64 photonic matrix multiply at >6-bit precision with WDM. 100K+ p-bit array for practical Boltzmann sampling. Neuromorphic MoE router at Mixtral-scale. MLIR dialects for each substrate.

Phase 2 ( ~€5–10B): Pairwise integration. Photonic + silicon (leveraging Ayar Labs/Intel trajectory). Entropic + silicon (following Extropic/Normal Computing roadmap). End-to-end transformer inference on photonic matmul with neuromorphic routing. Full heterogeneous compiler stack.

Phase 3 (~€10–20B): Four-substrate chiplet integration. Benchmark against contemporary GPUs on inference throughput/watt and training convergence/watt. Multi-fab supply chain across ≥3 geographies.

Total: €17–34B . For context: TSMC’s Arizona expansion costs $65B. Hyperscalers are spending $736B on AI compute over 2025–2026. The European Chips Act allocates €43B. ITER has consumed €20B+. This is large by academic standards and small relative to the problem.

Conclusion

The history of computing is a history of heterogeneity suppressed by convenience, then restored by necessity. Early computers mixed vacuum tubes, magnetic cores, and relay logic. The microprocessor collapsed everything into silicon. For decades the monoculture worked. Now it’s reaching limits — thermodynamic, economic, geopolitical — and pressure for diversification is building from every direction.

Light interferes — use that for linear algebra. Magnets fluctuate — use that for sampling. Neurons spike sparsely — use that for routing. Silicon switches reliably — use that for control. The mathematics of modern AI decomposes naturally along these physical boundaries. The fabrication requirements map onto different tiers of the semiconductor supply chain, distributing risk rather than concentrating it.

We do not claim this architecture will work. We claim it could work. The physics underlying each component is demonstrated. The integration challenges are formidable but bounded by engineering, not fundamental law. The alternative — scaling a single-vendor, single-fab, single-physics approach until it breaks or boils the planet — is not a plan. It is an addiction.

Build the machine the mathematics demands.

Key References (abbreviated; full bibliography available on request)

[1] Shen et al. (2017) “Deep learning with coherent nanophotonic circuits” Nature Photonics 11, 441–446. [2] Caves (1981) “Quantum-mechanical noise in an interferometer” Physical Review D 23, 1693. [3] Aasi et al. (2013) “Enhanced sensitivity of the LIGO gravitational wave detector” Nature Photonics 7, 613–619. [4] Giovannetti, Lloyd, Maccone (2004) “Quantum-Enhanced Measurements” Science 306, 1330–1336. [5] Braunstein & van Loock (2005) “Quantum information with continuous variables” Rev. Mod. Phys. 77, 513. [6] Arrazola et al. (2021) “Quantum circuits with many photons” Nature 591, 54–60. [7] Madsen et al. (2022) “Quantum computational advantage with a programmable photonic processor” Nature 606, 75–81. [8] Ozawa et al. (2019) “Topological photonics” Rev. Mod. Phys. 91, 015006. [9] Landauer (1961) “Irreversibility and Heat Generation” IBM J. Research 5, 183–191. [10] Bennett (1973) “Logical Reversibility of Computation” IBM J. Research 17, 525–532. [11] Bérut et al. (2012) “Experimental verification of Landauer’s principle” Nature 483, 187–189. [12] Camsari et al. (2017) “Stochastic p-bits for Invertible Logic” Physical Review X 7, 031014. [13] Borders et al. (2019) “Integer factorization using stochastic magnetic tunnel junctions” Nature 573, 390–393. [14] Singh et al. (2024) “CMOS plus stochastic nanomagnets” Nature Communications 15, 2685. [15] Hinton & Sejnowski (1983) “Optimal Perceptual Inference” IEEE CVPR. [16] Sohl-Dickstein et al. (2015) “Deep Unsupervised Learning using Nonequilibrium Thermodynamics” ICML. [17] Ho, Jain, Abbeel (2020) “Denoising Diffusion Probabilistic Models” NeurIPS. [18] Whitelam (2025) “Generative Thermodynamic Computing” arXiv:2506.15121. [19] Coles et al. (2023) “Thermodynamic AI and the Fluctuation Frontier” arXiv. [20] Indiveri & Liu (2015) “Memory and Information Processing in Neuromorphic Systems” Proc. IEEE 103, 1379–1397. [21] Davies et al. (2018) “Loihi: A Neuromorphic Manycore Processor” IEEE Micro 38, 82–99. [22] Reck et al. (1994) “Experimental realization of any discrete unitary operator” PRL 73, 58. [23] Clements et al. (2016) “Optimal design for universal multiport interferometers” Optica 3, 1460. [24] Chowdhury et al. (2023) “A full-stack view of probabilistic computing with p-bits” IEEE JXCDC. [25] Richter et al. (2024) “DYNAP-SE2” Neuromorphic Computing and Engineering. [26] Pernot-Borràs et al. (2022) “BrainScaleS-2 Accelerated Neuromorphic System” Frontiers in Neuroscience. [27] Kay, Date, Schuman (2020) “Neuromorphic graph algorithms” OSTI Technical Report. [28] Nazeer et al. (2023) “Language Modeling on a SpiNNaker 2 Neuromorphic Chip” arXiv:2312.09084. [29] Lattner et al. (2021) “MLIR: Scaling compiler infrastructure” IEEE CGO. [30] Lee et al. (2023) “Probabilistic computing with NbOₓ memristors” Nature Communications 14, 7349.

This paper integrates peer-reviewed physics, commercially reported data (verified against primary sources as of March 2026), and speculative architecture. The quantum photonics section is the most speculative and the authors know it. Rigorous criticism welcome — preferably over drinks.


메타데이터
post_id
c94bae8f0ffb
slug
thermal-punk-a-heterogeneous-architecture-for-post-silicon-ai-compute-c94bae8f0ffb
url
https://medium.com/@apoage/thermal-punk-a-heterogeneous-architecture-for-post-silicon-ai-compute-c94bae8f0ffb
canonical_url
https://medium.com/@apoage/thermal-punk-a-heterogeneous-architecture-for-post-silicon-ai-compute-c94bae8f0ffb
author_url
https://medium.com/@apoage
status
ok
fetched_at
2026-06-11 05:11:55