NoFOMO
Search 中EN Sign in Sign up

Peaked at #8 Off the board

SGLang Optimized for Nvidia Rubin

SGLang, with NVIDIA, optimized attention, MoE and speculative verification kernels on early Rubin hardware, speeding Kimi K3 inference up to 20% and end-to-end by 5.9%.

Rank over time

Original post

SGLang @sgl_project 10.7K followers

We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference. Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware. Highlights: • Up to 20% faster FP8 MLA at batch 1 / 128K context • 20% faster KDA verification, with bitwise-identical output • 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU. Full results and engineering details 👉 https://t.co/xGwlR0ax4L
137likes 22reposts 16replies 25.4Kviews

View on X Save

Top replies on X

Sign in to see 2 top replies from X

From @sgl_project, @NVIDIAAI and others. Spam removed, with English and Chinese translations.

Sign up free Have an account? Sign in

Discussion 0

No comments yet. Start the conversation.

Suggest a source

Is there a first-hand source we're missing, or a topic we should watch? Tell us. Once approved, everyone's board covers it.

@username or profile link