Pushing the Limits of KV Cache Compression · zartbot
- TL;DR When DeepSeek-V4.1 Flash was released, I thought it might just be a post-training iteration version...
- but after using it for a while, I found it reached nearly 420 Tokens/s in speed, and then Cui said all DeepSeek-V4 Pro models would be taken offline...
- suddenly I felt this was no small matter...
Unverified
- TL;DR When DeepSeek-V4.1 Flash was released, I thought it might just be a post-training iteration version...
- but after using it for a while, I found it reached nearly 420 Tokens/s in speed, and then Cui said all DeepSeek-V4 Pro models would be taken offline...
- suddenly I felt this was no small matter...
Sources: Github