Smaller, faster, safer: running Kimi and GLM at scale
- Workers AI runs inference for some of the best open models in the world on GPUs in Cloudflare data centers close to your users.
- Two of the most capable, and most demanding, are Moonshot's Kimi K-series and Z.ai's GLM.
- They are large, long-context, mixture-of-experts models, and they are wonderful to use.
Unverified
- Workers AI runs inference for some of the best open models in the world on GPUs in Cloudflare data centers close to your users.
- Two of the most capable, and most demanding, are Moonshot's Kimi K-series and Z.ai's GLM.
- They are large, long-context, mixture-of-experts models, and they are wonderful to use.
Sources: Cloudflare