Menu
The Desktop Frontier — Ahmad Osman, Osmantic

The Desktop Frontier — Ahmad Osman, Osmantic

AI Engineer

1,773 views yesterday

Video Summary

This presentation discusses the rapid advancement of local and open-source AI models, highlighting the increasing capability density and efficiency gains. The speaker predicts that within roughly 18 months, models equivalent to GLM 5.2 will run on a single RTX 5090 with 32GB of VRAM, a significant leap from current hardware requirements. This trend is attributed to architectural innovations and efficiency improvements, leading to smaller, more powerful models that can rival cloud-based solutions. The presentation also touches on the economic benefits and user control offered by owning local AI hardware.

Short Highlights

  • AI models are becoming significantly more efficient, with capabilities rapidly improving.
  • Predictions suggest GLM 5.2 class intelligence will run on a single RTX 5090 within 18 months.
  • Smaller, more efficient models are outperforming older, larger ones.
  • Local AI offers benefits like user control, cost savings, and customization.

Key Details

The Desktop Frontier: Where We Started [00:17]

  • The presentation focuses on the evolution of local and open-source AI models.
  • It aims to show how far these models have progressed from their initial stages.

    "it's basically about where we started and how far we've come with local and open source models."

Predictions for Future AI Capabilities [00:52]

  • A prediction is made that within approximately 18 months, GLM 5.2 class intelligence will run on a single RTX 5090 with 32GB of VRAM.
  • This is considered a conservative estimate, with the possibility of reaching this milestone even faster.

    "within roughly 18 months we are going to have the equivalent of GLM 5.2 class intelligence running on a single RTX 5090 with 32 GB of VRAM."

Shrinking Gap Between Frontier and Open Source Models [01:17]

  • Historically, the focus was on developing larger and larger models.
  • While a gap between frontier and open-source models will likely always exist, it is expected to shrink significantly.
  • The efficiency of models is improving exponentially.

    "there will always be a gap but that gap um will shrink and the efficiency of the models will get exponentially better."

Impact Per Parameter and Hardware Footprint [01:51]

  • The concept of "impact per parameter" is introduced as a key metric.
  • This considers the model's capabilities, hardware footprint, and resource requirements compared to previous iterations.
  • Examples show models like Llama 2 being surpassed by smaller, more efficient models like Qu 3.5.

    "what capability are we talking about? What could a model do? Um, what footprint like hardware footprint did it have last year in comparison to now?"

Advancements in Local Model Performance [03:00]

  • A year ago, local models struggled with tasks like code generation and context length.
  • Models like GLM 4.5 required significant hardware (multiple RTX 3090s or an RTX Pro 6000).
  • Current models achieve similar or better performance with much less hardware, like a single RTX 3090 or 1590.

    "Now that footprint for hardware is not needed anymore. All that you need is a single RTX 1390 1590 and you have something much more capable, much more intelligent."

Densing Law and Capability Density [04:33]

  • The trend is not about small models beating big models, but newer, efficient models surpassing older ones.
  • This phenomenon is referred to as "densing law" in literature, indicating increased capability density.
  • Models are delivering significantly more intelligence for their size and resource requirements.

    "It's not that small models are beating big models. It's that newer, more efficient models are beating older, less efficient ones."

Current Frontier Models and Local Accessibility [05:09]

  • GLM 5.2 is highlighted as a leading model with 744 billion parameters, supporting up to 1 million contexts.
  • It can be run on high-end hardware like a GGX station with eight RTX Pro 6000s.
  • Some benchmarks show GLM 5.2 outperforming GPT 5.5, indicating the competitiveness of local and open-source models.

    "That's something that you like a GX station is something that you can sit under your desk and it's running this kind of frontier intelligence."

The Rise of Sovereign AI and User Control [07:44]

  • The ability to run GPT40 quality models on an iPhone signifies a major shift.
  • This raises questions about investing in "sovereign AI" for individuals and businesses.
  • Owning local hardware provides control, optimization, and cost savings compared to cloud subscriptions.

    "Why wouldn't you want to be in control of the models that you're on? Why wouldn't you want to make sure that nothing gets taken away from you?"

Evolution of Open Source Models and Hardware Value [09:08]

  • Open-weight models like Mistral 7B and Llama 3 have shown significant progress.
  • The performance gains have been dramatic, with smaller models now outperforming much larger predecessors.
  • The value of existing hardware, like RTX 3090s, is expected to increase as models become more efficient.

    "So we've come a long Okay. Uh we had that we had mixed trial 8 by 7B which you know everybody knows is an MOE."

Future Outlook and Hardware Investment [15:33]

  • The question is posed whether current hardware purchases will increase in value as models become more efficient.
  • The speaker advises against funding external data centers and encourages owning hardware for long-term control and cost-effectiveness.
  • The potential of hardware like DGX stations is considered for future AI model execution.

    "So why are you funding other people to build data centers so that you can subscribe to them and pay subsidized tokens and then later on get those subsidies are going to go away and you're not going to be able to run those models and they will have so many limitations."

Other People Also See