NVIDIA releases PixelUMM on Hugging Face, a Qwen3-8B-based image-to-text model
nvidia · hf · 2026-10-02
NVIDIA has released PixelUMM on Hugging Face, an image-to-text multimodal understanding model fine-tuned from Qwen/Qwen3-8B. It supports image understanding, video understanding, and video-text-to-text tasks, licensed under a custom 'other' license (region: us).
Related event: NVIDIA Releases PixelUMM, Encoder-Free Unified Multimodal Model(3 posts)→
More from Models
- Fine-tuned Qwen3 4B on AWS beats Claude Sonnet 4.6 at lower cost — wonderwomancode · 2026-10-02
- Yacine claims Claude Opus 5.5's ML R&D classifier keeps blocking his work — yacineMTB · 2026-10-02
- Dev: AI NPCs need local inference or 100x cheaper compute to reach mass-market games — rickasaurus · 2026-10-02
- theo ships Slopalytics, a better dashboard for Artificial Analysis data with T3 Code usage stats — 0xkarasy · 2026-10-02
- ChatGPT Pro 500 plan hits £445 (~$589) in the UK — koltregaskes · 2026-10-02
- Open-Source Coding Agent Z-Engine Delegates Micro-Decisions to a System 1 Model — arshadbarves · 2026-10-02