DeepSeek's newest model activates just 8 billion of its 552 billion mixture-of-experts parameters per token — a new Causal ...