PixelUMM
Model on Hugging Face · Code · Paper
Encoder-free unified multimodal model from NVIDIA: one decoder-only Transformer (Qwen3-8B backbone, 15.2B params) that reads and writes raw pixels, with no VAE and no vision encoder. This demo runs the default S8-F22-R05 checkpoint. Weights are released under the NVIDIA One-Way Noncommercial License (research/evaluation only).
256 768
256 768
10 100
1 10
1 10
Examples
16 96
10 50
1 10
1 15
Larger resolutions with many frames run on a full-size GPU and use more quota. The official Cosmos guardrail checks are not run in this demo.
Examples
16 1024
0 1.5
Examples
| Image | Question |
|---|
16 1024
0 1.5
Frames are sampled at 1 fps (up to 96 frames, at most 448x448 each).