最高第 3 名 已下榜
PixelUMM开源代码与模型
作者团队宣布开源PixelUMM的代码与模型,并放出Hugging Face论文页面和模型链接。
要点
- NVIDIA 研究人员发布了 PixelUMM,这是一个统一的多模态模型,从架构中完全移除了 VAE 和 ViT,直接在原始像素空间处理图像和视频任务。
- PixelUMM 在单个无编码器模型中同时处理视觉理解和生成。
- PixelUMM 的代码和模型已开源,并提供了 Hugging Face 论文、代码和模型的链接。
要点和反应摘要由 AI 依据本页推文整理,请以原推为准。 我们怎么用 AI
原推
Thanks for sharing! We open-sourced code and models.
🤗Huggingface Paper: https://t.co/BrdK3sXeiy
💻 Code: https://t.co/S9ouFlnHWg
🤗 Model: https://t.co/DKW5T7E6xI引用 @robotsdigest: NVIDIA researchers just released PixelUMM, a unified multimodal model that fundamentally changes how models process visual data.
Groundbreaking highlights:
Entirely removes VAEs and ViTs from the architecture
Operates directly in raw pixel space for both image and video tasks
Handles both visual understanding and generation within a single encoder-free model
讨论 0
注册 登录 后参与讨论
还没有人讨论,来说第一句。