Key Moments

What If We Stopped Using GPUs? | YC Paper Club

Y CombinatorY Combinator
Science & Technology6 min read81 min video
Oct 2, 2026|12,417 views|168|8
Save to Pod
TL;DR

AI compute is hitting efficiency plateaus, pushing researchers to explore radical alternatives like optical and brain-inspired computing, but widespread adoption faces significant hurdles.

Key Insights

1

The cost of Gigaflops must decrease significantly to reach human brain power consumption levels (around 20 watts).

2

Optical computing offers near-instantaneous, low-power computation but struggles with data storage and non-linear activation functions.

3

Neuromorphic computing aims to mimic the brain's architecture using analog circuits, facing challenges in replicating the brain's dynamic learning and plasticity.

4

The human brain learns from as few as two examples, a stark contrast to the vast datasets required for current AI models.

5

Current AI models deviate significantly from the brain's biological mechanisms, often optimizing for silicon implementation rather than true biological fidelity.

6

A key challenge in optical computing is the expensive conversion between digital and optical domains, with 99.99% of 'jewels' being burned in the process.

The plateau of current AI hardware and the need for alternatives

The current paradigm of AI development is heavily reliant on the tight coupling of transformer architectures, backpropagation, and GPUs. This has led to significant advancements, but compute efficiency is beginning to plateau. The speaker highlights that the cost per Gigaflop needs to drastically reduce to even approach the energy efficiency of the human brain, which operates on approximately 20 watts. This fundamental limitation necessitates exploring entirely different approaches to building the machines that will power future AI. The historical shift in GPU development, from focusing on FLOPS (floating-point operations per second) to prioritizing memory and bandwidth, is noted as a response to the increasing demands of transformer models, particularly their attention mechanisms which scale quadratically.

Optical computing: The promise of light-speed, low-power computation

Optical computing leverages light (photons) for computation, offering significant advantages in speed and energy efficiency over electronics. Photons experience vastly less loss and offer much higher bandwidth compared to electrons. This makes light ideal for data transmission, with nearly all data in AI data centers and intercontinental communication relying on optical fibers. The core idea is to process information using light directly. In electronics, operations like matrix-vector multiplication involve charging wires and using transistors, consuming energy proportional to N-squared for N inputs/outputs. Optical systems, however, can modulate light directly on a number of modulators, allowing it to propagate through a designed weight mask. This process modifies the light's path without significant energy consumption. The output can then be collected on detectors. Companies like Lightmatter are already demonstrating optical accelerators that match GPU energy efficiency. However, a major hurdle is the conversion between digital electronic data and optical signals, which is expensive and inefficient. Additionally, while passive optics excel at linear operations, implementing non-linear activation functions, crucial for AI, requires additional complex mechanisms. The researchers propose carefully selecting use cases and algorithms where optical computing can offer a tangible advantage, such as in image generation using diffusion models by programming light propagation to mimic the diffusion process.

Neuromorphic computing: Mimicking the brain's efficiency and architecture

Neuromorphic computing seeks to emulate the brain's organizational principles in electronic circuits. The brain, a marvel of evolutionary engineering, performs complex tasks like language, reasoning, and creativity using only about 20 watts. Key features being explored include the tight integration of memory and computation, where synaptic connections hold memory and weights that directly influence calculations. Communication occurs through 'spikes,' which are non-linear and analog yet produce discrete digital outputs. This hybrid nature is difficult to replicate in purely digital systems. The brain's ability to learn and adapt continuously throughout life, even for tasks it wasn't evolutionarily prepared for, is a significant point of inspiration. Carver Mead coined the term 'neuromorphic engineering' in the late 1980s, aiming to apply these principles to chip design. Modern AI, while inspired by the brain (e.g., large networks of interconnected neurons), has largely deviated from biological fidelity, optimizing instead for silicon implementation. Researchers are exploring ways to integrate memory and computation more closely, as seen in companies like DMatrix, and to develop event-based computing that mimics the brain's continuous sensory input and spike generation. The concept of 'Spiking Neural Networks' (SNNs) and analog computing are active areas of research.

Moving beyond backpropagation: Alternative learning algorithms

The reliance on backpropagation, a cornerstone of deep learning, is also questioned. The speaker suggests that backpropagation may be phased out within a decade. The core hypothesis is that the cost of Gigaflops needs to fall dramatically. The speaker's PhD work focused on alternative optimization methods, exploring non-gradient-based learning procedures. They experimented with training a billion-parameter LSTM model using various methods and found that SPSA (Stochastic Parallel Search Algorithm), or finite difference methods, proved effective. SPSA offers a way to efficiently optimize loss landscapes that are flat or non-convex, outperforming backpropagation in certain scenarios by avoiding local minima. A key aspect is that it doesn't require gradients, making it potentially compatible with non-differentiable hardware like optical systems or biological neurons.

The Soma architecture: A fragmented mixture of experts

Inspired by the brain's cortical columns and a desire for alternative optimizers, the speaker proposes 'Soma.' This architecture involves fragmenting large datasets (like Common Crawl) and training small 'expert' models (e.g., 41k parameters) on distributed GPUs. At test time, these experts can be combined and routed, similar to Mixture-of-Experts (MoE) but as a mixture of models rather than just a feed-forward layer. A crucial finding is that the gradient estimation error in this optimizer worsens linearly with model size, making it impractical for very large models. However, by fragmenting the model, the noise in gradient estimation is reduced, allowing for efficient training with fewer updates. This approach shows promise for scaling by leveraging the concept of scaling laws across dimensions.

Challenges and potential of biological computing with living neurons

The possibility of using living brain cells for computation is explored, exemplified by training brain cells on a chip to play the game Doom. The approach treats the brain cells as an input-output system. Stimuli, such as game states (pixels, ammo, health), are encoded into electrical pulses delivered to the neuronal culture via an electrode array. These pulses are transmitted and processed by the neurons, mimicking a leaky integrate-and-fire model. The resulting neural activity is then decoded into game actions. While simpler games like Pong were manually coded, the complexity of Doom necessitated an end-to-end learning approach using reinforcement learning (PPO) to learn both the encoding of stimuli and the decoding of actions. A critical point emphasized is the avoidance of 'over-fitting' the decoder, which would undermine the learning process of the neurons themselves. The feedback mechanism is carefully designed to guide the neurons towards desired behaviors by adjusting stimulation frequencies and amplitudes based on game outcomes and 'surprise' factors, aiming to reduce entropy within the system.

The limitations of current approaches and the path forward

Despite the exciting possibilities, significant challenges remain across all these alternative compute paradigms. Optical computing struggles with data storage, non-linearities, and the conversion overhead. Neuromorphic computing faces difficulties in replicating the brain's dynamic learning and plasticity on static silicon. Biological computing with living neurons is in its nascent stages, with questions about scalability and true 'intelligence' versus sophisticated pattern matching. The speaker concludes that while current AI is powerful, it doesn't truly mirror the brain's efficiency or fundamental operating principles. The future likely lies in a co-design of hardware and algorithms, carefully considering the physics and constraints of new substrates, whether optical, neuromorphic, or biological. Patience and perhaps a bit of luck will be needed to unlock the next generation of AI compute.

Common Questions

Current GPUs are primarily focused on computational efficiency (FLOPS) and are hitting limitations in memory capacity and bandwidth. As AI models become more complex, particularly transformers, the demand for memory and speed increases significantly, outstripping the gains in raw processing power.

Topics

Mentioned in this video

More from Y Combinator

View all 634 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free