Data Filtering for Generative Video Pre-training | Field Notes by Linum
- Image and video models have gotten a lot better over the last few years, even though the internals of these models haven't changed much since Stable Diffusion 3.
- Of course, there have been small variants like the auto-regressive diffusion that GPT-Image popularized.
- But at a high level, it's pretty much all flow matching with a transformer backbone and a v-prediction objective.
Unverified
- Image and video models have gotten a lot better over the last few years, even though the internals of these models haven't changed much since Stable Diffusion 3.
- Of course, there have been small variants like the auto-regressive diffusion that GPT-Image popularized.
- But at a high level, it's pretty much all flow matching with a transformer backbone and a v-prediction objective.
Sources: Linum