
Description
A research model from NVIDIA and the University of Waterloo that performs image and video tasks directly in pixel space. It removes the standard Variational Autoencoder (VAE) step, opting for a single-decoder architecture to handle both generation and understanding of visual media.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.