Neocloud

Together AI

Serverless open-model APIs and GPU clusters from one shop, rent the fleet or just the tokens.

Best known for per-token APIs serving open models (Llama, DeepSeek, Qwen) at low prices, but also rents dedicated H100/H200/B200 clusters. A natural home when you want to graduate from tokens to your own fine-tuned deployment without changing vendors.

Visit Together ↗

GPU pricing on Together

On-demand list per GPU-hour, last reviewed 2026-09. Methodology.

GPU $/GPU-hr vs cheapest Notes
H100 SXM $2.39 cheapest
H200 $3.15 1.3×
B200 $5.50 2.3×

Strengths

  • Token API ↔ dedicated GPU continuum
  • Strong open-model serving stack
  • Research pedigree (FlashAttention lineage)

Watch for

  • Cluster minimums for best pricing

Alternatives

CoreWeave

The flagship neocloud: Kubernetes-native GPU fleets built for training clusters, now public company.

Lambda

The ML researcher's default: one-click clusters, sane pricing, no quota theater.

RunPod

Community and secure cloud tiers, per-second billing, and serverless GPU endpoints.