Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
- I gave Qwen3.8's MTP drafter another 69.2 MiB of precision.
- Throughput fell from 50.44 to 37.02 tokens per second.
- That result sums up the whole experiment: the best local inference setup is rarely made from the individually "best" parts.
Unverified
- I gave Qwen3.8's MTP drafter another 69.2 MiB of precision.
- Throughput fell from 50.44 to 37.02 tokens per second.
- That result sums up the whole experiment: the best local inference setup is rarely made from the individually "best" parts.
Sources: Piszczek