Product Edu

Inside DeepSeek v3: Training One of the Largest AI Models on 15 Trillion Tokens

Discover the intricate training journey of DeepSeek v3, one of the largest language models, trained on a colossal dataset of 15 trillion tokens. This includes a diverse range of multilingual, domain-specific, scientific, and conversational data to ensure comprehensive knowledge coverage. Learn about the architectural strategies that address challenges such as latency and computational expense, making DeepSeek v3 a powerhouse in AI.
00:00 Introduction to DeepSeek v3
00:36 Training Journey: Scale to Precision