Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, Xinference, Ollama, vLLM and CoderAI (September 2026)
- This is the reference version of a survey I did for my own project, written so it is useful to someone who is not running CoderAI.
- If you have one or more machines with GPUs and want an OpenAI-compatible endpoint in front of them, these are the self-hosted orchestrators that exist in September 2026, what each one actually does across machines, and which one to pick for which situation.
- Star counts are from the GitHub API on 2026-09-20; feature cells are from the projects' own README and docs, linked at the end.
Unverified
- This is the reference version of a survey I did for my own project, written so it is useful to someone who is not running CoderAI.
- If you have one or more machines with GPUs and want an OpenAI-compatible endpoint in front of them, these are the self-hosted orchestrators that exist in September 2026, what each one actually does across machines, and which one to pick for which situation.
- Star counts are from the GitHub API on 2026-09-20; feature cells are from the projects' own README and docs, linked at the end.
Sources: Nexlab