
Breaking New Ground with DeepSeek R1
Excited to share the results from our latest experiment with the DeepSeek R1 Distilled LLaMA 8B model with Dr. Farhan Ahmed Karim (Supervisor) and his Students Shayan Baig, Abdul Rehman and Syed Saad Akhtar.
Running on dual Radeon RX 7900 XTX GPUs, we tested two variants:
Q8_0 (8-bit): 0.152s per token — speed & efficiency at its finest.
BF16 (16-bit): 2.2s per token — precision meets performance.
DeepSeek R1 continues to impress with its balance of speed and accuracy, paving the way for high-performance AI in production. Big thanks to the DeepSeek team for making this breakthrough possible.