Peaked at #8 Off the board
SGLang Optimized for Nvidia Rubin
SGLang, with NVIDIA, optimized attention, MoE and speculative verification kernels on early Rubin hardware, speeding Kimi K3 inference up to 20% and end-to-end by 5.9%.
Original post
We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference.
Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware.
Highlights:
• Up to 20% faster FP8 MLA at batch 1 / 128K context
• 20% faster KDA verification, with bitwise-identical output
• 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step
SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
Full results and engineering details 👉 https://t.co/xGwlR0ax4L
Top replies on X
Sign in to see 2 top replies from X
From @sgl_project, @NVIDIAAI and others. Spam removed, with English and Chinese translations.
Discussion 0
Sign up Sign in to join the discussion
No comments yet. Start the conversation.