Menu
What If We Stopped Using GPUs? | YC Paper Club

What If We Stopped Using GPUs? | YC Paper Club

Y Combinator

229 views • 22 hours ago Save 70 min 10 min read

Video Summary

The pursuit of artificial intelligence is hitting fundamental compute limits, prompting a search for radically new approaches beyond current hardware architectures. While deep learning has advanced significantly, the "hardware-software co-adaptation" has led to diminishing returns in compute efficiency, particularly with the rise of transformer models demanding immense memory bandwidth. This has spurred interest in alternative computing paradigms, including optical computing, neuromorphic systems, and novel optimization algorithms that move away from traditional backpropagation.

Optical computing offers potential for near-instantaneous, low-power computation by using light instead of electrons, though challenges remain in data conversion and non-linear operations. Neuromorphic computing aims to mimic the brain's energy efficiency and learning capabilities by integrating memory and computation, using event-driven "spikes" rather than continuous signals. Meanwhile, researchers are exploring non-gradient-based optimization methods like SPSA, which can handle non-differentiable loss landscapes more effectively than backpropagation. One promising direction involves "sharded optimization" where models are broken into smaller, specialized "experts," potentially overcoming the scaling limitations of current optimizers and paving the way for more efficient AI.

Short Highlights

  • The Limits of Current AI Compute: Deep learning's reliance on hardware-software co-adaptation, especially for transformers, faces diminishing returns in compute efficiency and memory bandwidth.
  • Optical Computing Promise: Harnessing light for computation offers near-instantaneous, low-power processing, bypassing electronic limitations, though data conversion and non-linearities are challenges.
  • Neuromorphic Computing Inspiration: Mimicking the brain's energy efficiency and integrated memory-computation through "spiking" event-driven systems is a key focus.
  • Alternative Optimization Algorithms: Moving beyond backpropagation, methods like SPSA show promise for handling complex, non-differentiable problems.
  • Sharded Optimization: Breaking down models into specialized "experts" could overcome scaling limitations of current optimizers and improve efficiency.
  • Biocomputing for AI: Using biological neural networks, like brain cells on chips, to perform AI tasks is an experimental frontier, treating computation as a dynamic systems problem.
  • Hardware-Algorithm Co-Design: The future of AI compute likely lies in tightly integrating hardware capabilities with algorithmic design, rather than forcing models onto existing architectures.

Key Details

The Compute Bottleneck in AI [00:00:00]

  • The current trajectory of AI development is hitting fundamental compute limitations.
  • Hardware and software co-adaptation, particularly for transformer architectures, is yielding diminishing returns in efficiency.
  • The demand for memory capacity and bandwidth has surged with models like GPT-3, outstripping previous efficiency gains.

    "i think that we're so far down this rabbit hole of hardware software co-adaptation or more specifically hardware architecture and optimizer co-adaptation"

The Rise of Transformers and Memory Demands [00:02:00]

  • Nvidia's focus shifted from compute efficiency (flops per joule) to memory capacity and bandwidth after 2020, driven by transformer models.
  • The "attention" mechanism in transformers, with its N-squared complexity, is a likely driver for increased memory demands.
  • Compute efficiency improvements have plateaued in the last two years, signaling a need for new approaches.

    "these transformers demand more memory capacity and bandwidth than ever before"

The Brain as a Model for Efficiency [00:04:00]

  • The human brain operates on approximately 20 watts, performing complex tasks with incredible efficiency.
  • Achieving human-level AI may require a significant reduction in the cost per gigaflop.
  • Current AI hardware often focuses on training and inference chips, but largely remains tied to transformer architectures.

    "the human brain draws like 20 watts — that's like a light bulb — and it can do insane amounts of things"

Rethinking Backpropagation and Neural Computation [00:06:00]

  • The brain is massively feed-forward, contrasting with backpropagation, which requires backward passes.
  • The "weight transport problem" in backpropagation highlights the challenge of ensuring weights remain identical during forward and backward passes.
  • The brain utilizes massive inhibition within cortical columns, allowing assemblies to learn independently.

    "the brain is massively feed forward and so any argument that i've tried to that i've listened to that the brain does back propagation where the would require that these neurons that this neuron here fires forward and then it would have to fire backwards"

The Promise of Optical Computing [00:08:00]

  • Optical communication is highly efficient due to photons' low loss and high bandwidth compared to electrons.
  • The idea is to leverage light not just for communication but for computation itself, potentially eliminating energy loss from electronic conversions.
  • Photons do not interact, allowing for massively parallel computation within a given space.

    "if it provides such a low loss and high bandwidth why don't we process information with photons as well"

Challenges in Optical Computing [00:10:00]

  • Converting digital data to analog and then optical domains, and back again, incurs significant energy costs.
  • Programming and calibrating optical devices requires active energy expenditure.
  • Passive optics primarily support linear operations, while non-linear activation functions, crucial for AI, require additional mechanisms.

    "one of the challenges is that we pay this large price for converting digital information to analog and then optical domain"

Optical Diffusion-Based Image Generation [00:12:00]

  • A proposed use case involves programming light propagation for diffusion-based image generation, using optics as a denoising inference unit.
  • This approach aims to create an application-specific, passive device where light physics performs the diffusion steps.
  • The system iterates through a series of transparencies that modulate light's phase to diffract and filter information, gradually denoising an image.

    "our proposition here is that the using a general purpose electronic hardware for this task creates a big redundancy"

Optical Compute Scaling and Energy Efficiency [00:14:00]

  • Optical systems exhibit power-law scaling similar to digital neural networks as parameters increase.
  • Adding more parameters in optical systems is expected to lead to more favorable energy consumption scaling.
  • Early comparisons show significant energy advantages for optical systems in generating single images, assuming passive, fixed weights.

    "we could also show that by adding these extra parameters we can in a uh expectable way the performance improves in terms of the image quality"

Neuromorphic Computing: Mimicking the Brain [00:16:00]

  • Neuromorphic computing aims to borrow the brain's organizing principles for designing circuits and chips.
  • Key principles include intertwined memory and computation, event-driven communication via "spikes," and adaptive structure/function through learning.
  • A spiking neuron can be modeled as a leaky capacitor that fires when accumulated charge exceeds a threshold.

    "at the highest level it is can we borrow the brain's organizing principles somehow in designing circuits and chips and things that we use to do computation"

Digital vs. Biological Brain Models [00:18:00]

  • Modern neural networks, while brain-inspired, often deviate from biological plausibility, using techniques like backpropagation optimized for silicon.
  • The brain's hardware and software are deeply intertwined and co-evolved, a concept difficult to replicate on different computing substrates.
  • Neuromorphic computing is still largely in R&D, seeking the right balance between biological inspiration and physical implementation.

    "we've already kind of gone off off course and and and we're doing things native to the manifestation of these neural networks in silicon rather than than kind of staying staying true to you know faithful to the brain"

Emerging Neuromorphic Approaches [00:20:00]

  • Current research explores commingling memory and computation, often in purely digital systems.
  • Edge computing applications leverage spiking networks for on-device learning and event-driven computation.
  • New physical substrates like photonics and exotic electronic devices are being explored for their unique analog properties and non-linearities.

    "the idea is is there something we can take away from that property in order to um to to make ais far more energy efficient"

Alternative Optimization Algorithms: SPSA [00:22:00]

  • The speaker's PhD work focuses on alternative optimizers, moving beyond backpropagation.
  • SPSA (Simultaneous Perturbation Stochastic Approximation) is a zero-order optimization method that does not require gradients.
  • SPSA excels at optimizing non-Lipschitz and non-differentiable loss landscapes, unlike backpropagation which can get stuck in local minima.

    "the learning update rule that worked the best is something called spsa that pretty much no one knows about except james ball and his students"

Sharded Optimization: Mixture of Assemblies (SOMA) [00:24:00]

  • SOMA proposes sharding models into smaller "experts" (e.g., LSTMs) trained on clustered data.
  • This approach caps the gradient noise associated with model size, making training large models feasible with zero-order methods.
  • It leverages scaling laws across perturbations, batch size, and model sharding for improved performance.

    "what if we just take our common crawl data set cluster it with tf idf and some svd thing cluster the data set and then ship a parkette file to a gp to n gpus all throughout the world"

Biocomputing: Brain Cells Playing Doom [00:26:00]

  • Cortical Labs has placed human brain cells on chips and developed algorithms to enable them to play games like Pong and Doom.
  • The approach treats the brain cells as a dynamic system, encoding game states into electrical stimuli and decoding cell "spikes" into game actions.
  • For complex tasks like Doom, end-to-end learning of both encoding and decoding is necessary, as manual mapping is infeasible.

    "we want to treat it as kind of like an input output machine where some set of stimuli can you know produce some set of intelligence that allows us to get some form of output that's useful computationally"

Challenges and Feedback in Biocomputing [00:28:00]

  • Training requires stochastic stimulation (e.g., beta sampling) due to the inability to backpropagate through the biological substrate.
  • Feedback mechanisms are crucial; in this case, a silicon critic predicts "surprise" to guide learning, balancing positive and negative reinforcement.
  • Avoiding decoder overfitting is critical to ensure the brain cells, not the silicon, are learning.

    "it's very easy to cheat performance out of these systems all you have to do is you know massively oversize your decoder run back propagation on it and suddenly you have brain cells that do anything play anything"

Scalability and Future of Biocomputing [00:30:00]

  • The scalability of biocomputing hinges on achieving frontier-level intelligence in cells and serving it to billions.
  • Distributed computing with biological substrates is a significant unknown.
  • While current methods differ from natural sensory inputs, future evolution might lead to more similar interfaces or entirely new forms of intelligence.

    "we fundamentally don't know how to do this yet so you know kind of a shitty answer to that would be i don't know how scalable it is"

The Interplay of Hardware, Software, and Materials [00:32:00]

  • The rapid iteration cycles of software and hardware demand a co-design approach.
  • The choice between iterating models in software first or designing around hardware depends on the specific insights and constraints.
  • Future compute may involve more exotic materials beyond silicon, such as carbon-based substrates, to better mimic or surpass biological capabilities.

    "i think it just kind of depends on on where your where your angle comes from"

Noise, Reproducibility, and Novel Computing Paradigms [00:34:00]

  • The role of noise and reproducibility is a key consideration, contrasting with the deterministic nature of traditional silicon computing.
  • Thermodynamic computing and Boltzmann machines explore leveraging noise for computation.
  • Ensuring reproducibility of noise, particularly for zero-order methods like SPSA, remains a significant challenge.

    "the intention is for the computation to either be reproducible or extremely close to reproducible because of parallelism sort of stuff like that like little edge cases"

The Future Landscape of AI Compute [00:36:00]

  • The field is exploring diverse avenues: optical computing for speed and efficiency, neuromorphic systems for brain-like processing, and novel optimizers for better training.
  • Co-designing hardware and algorithms is paramount, leveraging the unique strengths of each substrate.
  • The ultimate goal is to overcome current compute limitations and unlock new frontiers in AI capabilities, potentially leading to forms of intelligence fundamentally different from our own.

    "i think that the future is going to be a lot more about the co-design of hardware and algorithms"

Other People Also See