CUDA cores vs Tensor Cores

8 min read · updated 2026-09

Spec sheets lead with CUDA core counts because the number is big and familiar. For AI it is close to a decoy. The units doing your matrix math are the Tensor Cores, and which generation of them a GPU carries decides what precisions it speaks, which frameworks love it, and what a fair price for it looks like.

Two kinds of worker

A CUDA core is a general-purpose arithmetic unit: one multiply or add per value per clock, thousands of them working in lockstep groups. They run anything you can express as parallel math: physics, image pipelines, sorting, the glue logic of a neural network (activations, normalization, sampling). Flexible, and for AI, slow.

A Tensor Core is a matrix engine. Instead of one value at a time, it consumes a small tile of matrix A, a tile of matrix B, and produces their product in one operation, hundreds of multiply-accumulates per core per clock. Since a transformer forward pass is, by compute volume, almost entirely matrix multiplication, this specialization is worth an order of magnitude: the H100 manages about 67 teraFLOPS of FP32 through its CUDA cores and about 990 teraFLOPS of BF16 through its Tensor Cores.

The division of labor during inference: Tensor Cores burn through the weight matrices; CUDA cores handle softmax, layer norm, rotary embeddings, sampling, and data movement bookkeeping. Both matter, but when you buy FLOPS for AI, you are buying Tensor Cores.

The generations, and why they gate what you can run

GenArchitectureNew precisionsGPUs in our database
1stVolta (2017)FP16V100
2ndTuring (2018)INT8, INT4T4
3rdAmpere (2020)TF32, BF16, sparsityA100 80GB, A100 40GB, A10, RTX 3090
4thHopper / Ada (2022-23)FP8, Transformer EngineH100, H200, L40S, L4, RTX 4090
5thBlackwell (2024-25)FP4B200, RTX 5090

Each generation is really a statement about precision. Lower precision means fewer bytes per weight, and since inference throughput is bound by bytes moved, each step down roughly doubles effective serving speed on the same memory system. That is the practical meaning of the table:

Matching cores to models

How to read a spec sheet now

Skip the CUDA core count. Read, in order: memory capacity (does the model fit: calculator), memory bandwidth (how fast it serves), Tensor throughput at your target precision (can it run FP8/FP4, and how hard), then price (the database tracks all four). A GPU is a memory system with a matrix engine attached; the marketing numbers are ranked in almost exactly the reverse of what matters.

Related