NoFOMO
搜索 中EN 登录 注册

最高第 8 名 已下榜

SGLang 适配英伟达 Rubin

SGLang与英伟达合作,在早期Rubin硬件上优化注意力、MoE与投机验证内核,Kimi K3推理最高加速20%,端到端提速5.9%。

名次变化

原推

SGLang @sgl_project 1.1万 粉丝

We brought SGLang to NVIDIA Vera Rubin and accelerated Kimi K3 inference. Working closely with @NVIDIA, we optimized attention, MoE, and speculative verification kernels on early-access Rubin hardware. Highlights: • Up to 20% faster FP8 MLA at batch 1 / 128K context • 20% faster KDA verification, with bitwise-identical output • 5.9% end-to-end inference speedup from MoE tail fusion, removing 276 kernel launches per decode step SGLang also powers rollouts for Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU. Full results and engineering details 👉 https://t.co/xGwlR0ax4L
137点赞 22转发 16回复 2.5万浏览

在 X 上查看 收藏

X 上的热门回复

登录后查看 X 上的 2 条热门回复

来自 @sgl_project, @NVIDIAAI 等,已过滤广告,附中英文翻译。

免费注册 已有账号?登录

讨论 0

还没有人讨论,来说第一句。

推荐信源

觉得我们漏了哪个一手信源,或者该关注什么话题?告诉我们。审核通过后,所有人的榜单都会覆盖到它。

@用户名,或者主页链接