Neocloud
Together AI
Serverless open-model APIs and GPU clusters from one shop, rent the fleet or just the tokens.
Best known for per-token APIs serving open models (Llama, DeepSeek, Qwen) at low prices, but also rents dedicated H100/H200/B200 clusters. A natural home when you want to graduate from tokens to your own fine-tuned deployment without changing vendors.
GPU pricing on Together
On-demand list per GPU-hour, last reviewed 2026-09. Methodology.
Strengths
- Token API ↔ dedicated GPU continuum
- Strong open-model serving stack
- Research pedigree (FlashAttention lineage)
Watch for
- Cluster minimums for best pricing