CUDA cores vs Tensor Cores
8 min read · updated 2026-09
Spec sheets lead with CUDA core counts because the number is big and familiar. For AI it is close to a decoy. The units doing your matrix math are the Tensor Cores, and which generation of them a GPU carries decides what precisions it speaks, which frameworks love it, and what a fair price for it looks like.
Two kinds of worker
A CUDA core is a general-purpose arithmetic unit: one multiply or add per value per clock, thousands of them working in lockstep groups. They run anything you can express as parallel math: physics, image pipelines, sorting, the glue logic of a neural network (activations, normalization, sampling). Flexible, and for AI, slow.
A Tensor Core is a matrix engine. Instead of one value at a time, it consumes a small tile of matrix A, a tile of matrix B, and produces their product in one operation, hundreds of multiply-accumulates per core per clock. Since a transformer forward pass is, by compute volume, almost entirely matrix multiplication, this specialization is worth an order of magnitude: the H100 manages about 67 teraFLOPS of FP32 through its CUDA cores and about 990 teraFLOPS of BF16 through its Tensor Cores.
The division of labor during inference: Tensor Cores burn through the weight matrices; CUDA cores handle softmax, layer norm, rotary embeddings, sampling, and data movement bookkeeping. Both matter, but when you buy FLOPS for AI, you are buying Tensor Cores.
The generations, and why they gate what you can run
| Gen | Architecture | New precisions | GPUs in our database |
|---|---|---|---|
| 1st | Volta (2017) | FP16 | V100 |
| 2nd | Turing (2018) | INT8, INT4 | T4 |
| 3rd | Ampere (2020) | TF32, BF16, sparsity | A100 80GB, A100 40GB, A10, RTX 3090 |
| 4th | Hopper / Ada (2022-23) | FP8, Transformer Engine | H100, H200, L40S, L4, RTX 4090 |
| 5th | Blackwell (2024-25) | FP4 | B200, RTX 5090 |
Each generation is really a statement about precision. Lower precision means fewer bytes per weight, and since inference throughput is bound by bytes moved, each step down roughly doubles effective serving speed on the same memory system. That is the practical meaning of the table:
- 3rd gen (A100): BF16 made large-model training numerically sane, which is why Ampere trained the first LLM generation. No FP8: an A100 serves 8-bit models through INT8 quantization instead, a well-trodden path in vLLM and TensorRT, and the reason the A100 remains a serving bargain rather than an antique.
- 4th gen (H100, RTX 4090): native FP8 plus the Transformer Engine, which auto-manages per-layer precision. This is the single biggest H100-vs-A100 capability gap, bigger than the raw FLOPS difference: FP8 halves bytes per token. The full decision math is in A100 vs H100.
- 5th gen (B200, RTX 5090): FP4 with much smarter scaling. Quality at 4-bit is workload-dependent, but where it holds, one B200 serves what recently took a rack slice.
Matching cores to models
- Serving open LLMs (Llama, Qwen, Mistral, DeepSeek): Tensor Core generation sets your precision ceiling; memory bandwidth sets your speed. A100-class 3rd gen with INT8 or 4-bit weight quantization is the value play; 4th gen and up when native FP8 and latency targets justify the premium.
- Fine-tuning: BF16 is the workhorse precision, available from 3rd gen on. This is why fine-tuning is the A100's strongest event, no FP8 needed.
- Diffusion and image models: compute-hungrier per byte than LLMs, so Tensor FLOPS matter more relative to bandwidth. Ada cards (L40S, 4090) punch above their weight here.
- Embeddings, rerankers, Whisper, classic ML: nearly any Tensor-Core-bearing card is overkill; buy the cheapest reliable memory and ignore this whole article.
How to read a spec sheet now
Skip the CUDA core count. Read, in order: memory capacity (does the model fit: calculator), memory bandwidth (how fast it serves), Tensor throughput at your target precision (can it run FP8/FP4, and how hard), then price (the database tracks all four). A GPU is a memory system with a matrix engine attached; the marketing numbers are ranked in almost exactly the reverse of what matters.