NoFOMO
搜索 中EN 登录 注册

最高第 3 名 已下榜

PixelUMM开源代码与模型

作者团队宣布开源PixelUMM的代码与模型,并放出Hugging Face论文页面和模型链接。

名次变化

要点

  • NVIDIA 研究人员发布了 PixelUMM,这是一个统一的多模态模型,从架构中完全移除了 VAE 和 ViT,直接在原始像素空间处理图像和视频任务。
  • PixelUMM 在单个无编码器模型中同时处理视觉理解和生成。
  • PixelUMM 的代码和模型已开源,并提供了 Hugging Face 论文、代码和模型的链接。

要点和反应摘要由 AI 依据本页推文整理,请以原推为准。 我们怎么用 AI

原推

Cong Wei @CongWei1230 1,182 粉丝

Thanks for sharing! We open-sourced code and models. 🤗Huggingface Paper: https://t.co/BrdK3sXeiy 💻 Code: https://t.co/S9ouFlnHWg 🤗 Model: https://t.co/DKW5T7E6xI引用 @robotsdigest: NVIDIA researchers just released PixelUMM, a unified multimodal model that fundamentally changes how models process visual data. Groundbreaking highlights: Entirely removes VAEs and ViTs from the architecture Operates directly in raw pixel space for both image and video tasks Handles both visual understanding and generation within a single encoder-free model
524点赞 66转发 6回复 3.3万浏览

在 X 上查看 收藏

讨论 0

还没有人讨论,来说第一句。

推荐信源

觉得我们漏了哪个一手信源,或者该关注什么话题?告诉我们。审核通过后,所有人的榜单都会覆盖到它。

@用户名,或者主页链接