Even by Generative AI standards, the rate and scale of changed introduced by Chinese startup DeepSeek in the past week has been extraordinary. Called by some the “Temu of ChatGPTs,” DeepSeek’s disruptive approach to generative AI has wiped out one trillion dollars (!) off the AI industry and is forcing established players to rethink their strategies.
Here’s why it matters—and what it means for your business.
Focus On: DeepSeek - Faster, Cheaper, and Open-Source
DeepSeek’s recently released open-source model, DeepSeek-R1, is making headlines for its exceptional performance on reasoning and mathematical benchmarks. Incredibly, the start-up developed this model in just two months at a cost of under $6 million—a fraction of what industry giants like OpenAI reportedly spend.
There is no silver bullet here, rather a combination of techniques that allowed DeepSeek to achieve this result, let’s explore them:
Optimised Compute Usage: Rather than relying on massive supercomputing clusters, DeepSeek leveraged existing model architectures and maximised training efficiency. Reports suggest they used approximately 3 million GPU hours, compared to Meta’s 60-70 million GPU hours for similar training efforts.
Mixture-of-Experts (MoE) Architecture: Unlike traditional dense models, where all parameters are activated simultaneously, MoE engages only a fraction of the total model during inference (the phase when an AI model generates answers). This significantly reduces computational demands while maintaining performance.
Distillation and Transfer Learning: DeepSeek refined its models using advanced distillation techniques, training smaller, more efficient models based on knowledge extracted from larger, more complex systems. This reduces training costs while maintaining competitive performance levels, although it is also raising questions on whether they distilled third party models as well.
Leveraging Open-Source Innovations: Rather than building everything from scratch, DeepSeek incorporated state-of-the-art developments from the open-source AI community, adapting and refining existing breakthroughs, effectively cutting corners.
Optimised Hardware Utilisation: While OpenAI and other major players purchase and deploy massive fleets of GPUs, DeepSeek took a more frugal approach, focussing on maximising the utility of each GPU in their cluster. They designed their infrastructure to run on widely available, cost-efficient hardware, significantly cutting cost.
Training Efficiency through Data Selection: DeepSeek employed highly refined data selection strategies, ensuring that only the most relevant, high-quality data was used for training. This further enhanced their cost-effectiveness while maintaining high performance.
DeepSeek’s approach signals a shift toward greater capital efficiency in the generative AI arms race, which is why leading chip makers, notably Nvidia, took the biggest hit from the news.
What does this mean for the AI ecosystem
DeepSeek’s emergence highlights the growing role of smaller players in shaping the generative AI landscape. But it also raises big questions, some not entirely new:
Economic Pressure on Big Tech: The efficiency of DeepSeek’s model could disrupt profit assumptions for hyperscalers and chipmakers like Nvidia, as companies look at novel ways to train and fine tune their models without the capex investments previously seen as a necessity.
Open Source vs Proprietary Models: DeepSeek’s MIT licensing framework allows companies to adapt and commercialise its models, in stark contrast to the closed ecosystems of larger players. It also showed how companies can take from multiple open-source models and iterate into highly sophisticated models.
Potential Price Wars: With DeepSeek pricing their models at 20-40x lower than OpenAI’s equivalents, AI industry-wide pricing structures may be forced to shift, driving down costs globally.
Key Takeaways for Business Leaders
Reassess AI Investments: DeepSeek’s cost efficiency highlights the importance of optimising AI budgets. Evaluate whether your organisation can achieve similar outcomes by leveraging open-source models or adopting new techniques.
Prepare for Commoditisation: As generative AI becomes more accessible, businesses must focus on leveraging these tools for product innovation, customer experience, and new revenue streams.
Monitor Emerging Players: The DeepSeek story is a reminder that disruption often comes from unexpected sources. Keep an eye on smaller competitors and the open-source community for ideas that could give you an edge.
Be Open for Business: The shift toward more efficient AI models suggests a possible transformation in the market, where AI capabilities become widespread and more accessible, ultimately driving faster innovation cycles. Clients will want greater access to data to create value leveraging their own IP, walled gardens might not work as well as they did in the past.
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.