Qwen/Qwen3.8-2.4T-A95B · Hugging Face
- This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
- These artifacts are compatible with vLLM, SGLang, TokenSpeed, etc.
- The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
Confirmed
- This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
- These artifacts are compatible with vLLM, SGLang, TokenSpeed, etc.
Unverified
- The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
Sources: Huggingface