view article Article The Transformers modeling backend serves DeepSeek-R1 at native speed hmellor • 14 days ago • 2
view article Article The Transformers modeling backend is now as fast as native vLLM hmellor • Jul 17 • 1