Training Text-to-Image Models 3.6× Faster
- TL;DRLinum v2 was bottlenecked by the enormous size of its attention context window.
- A 720p, 5 second clip cost a whopping 110K tokens.
- To put that in perspective, LLMs see samples with fewer than 8K tokens for 97% of their pretraining.
Unverified
- TL;DRLinum v2 was bottlenecked by the enormous size of its attention context window.
- A 720p, 5 second clip cost a whopping 110K tokens.
- To put that in perspective, LLMs see samples with fewer than 8K tokens for 97% of their pretraining.
Sources: Linum