PixelUMM

Model on Hugging Face · Code · Paper

Encoder-free unified multimodal model from NVIDIA: one decoder-only Transformer (Qwen3-8B backbone, 15.2B params) that reads and writes raw pixels, with no VAE and no vision encoder. This demo runs the default S8-F22-R05 checkpoint. Weights are released under the NVIDIA One-Way Noncommercial License (research/evaluation only).

256 768
256 768
10 100
1 10
1 10
Examples