What Is AI Inference? Training vs Inference, Speed and Cost Explained
Inference is the moment a trained AI model is actually used. It is where the cost, the latency and the hardware decisions live.
Key takeaways
- Inference is using a trained model to produce an output. Every chatbot reply, image generation and transcription is an inference. Training builds the model; inference runs it.
- Training happens once and costs millions. Inference happens billions of times, and over a model's life it is where most of the compute and money go.
- For language models, inference has two phases: prefill (reading your prompt, compute-bound) and decode (writing the answer one token at a time, memory-bound).
- Generation speed is capped by a simple ratio: memory bandwidth ÷ model size. That is why memory, not raw compute, drives hardware choices for local AI.
- Inference can run in the cloud, on your own server, or on your laptop. The model and the math are the same; privacy, cost structure and latency are not.
Inference in plain English
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- 30-day refund
A machine-learning model is, in the end, a very large set of numbers called weights. Training is the process of finding good values for those numbers by showing the model enormous amounts of data and nudging the weights after each mistake. Inference is what comes after: the weights are frozen, you give the model a new input, and it computes an output. The model learns nothing during inference. It applies what it already knows.
The word comes from statistics, where to infer is to draw a conclusion from evidence. In AI it has narrowed to mean "running the model." When a provider talks about inference cost, an inference API or an inference chip, they mean the business of serving answers.
Training vs inference
| Training | Inference | |
|---|---|---|
| Goal | Find the weights | Use the weights |
| How often | Once per model version | Every single request |
| Hardware | Thousands of data-center GPUs with fast interconnect | Anything from a GPU cluster to a phone |
| Numeric precision | 16-bit or mixed, to keep gradients stable | 8-bit or 4-bit is usually enough |
| Memory needed | Weights + gradients + optimizer state: several times the model size | Weights + working memory: about the model size |
| What you optimize | Total time and cost to reach a quality target | Latency per request, throughput, cost per token |
| Who does it | A few dozen labs | Everyone |
Between the two sits fine-tuning: a short, cheap round of extra training that adapts an existing model. It is training, not inference, but techniques like LoRA bring it within reach of a single consumer GPU.
The two phases of LLM inference
When you send a prompt to a language model, two very different things happen in sequence.
| Prefill | Decode | |
|---|---|---|
| What it does | Processes the whole prompt in one parallel pass and fills the KV cache | Produces the answer one token at a time, each needing a full pass through the model |
| Limited by | Compute (arithmetic throughput) | Memory bandwidth (how fast weights can be read) |
| You experience it as | The wait before the first word | The speed at which text streams |
| Typical speed on a consumer GPU | Hundreds to thousands of tokens per second | 10 to 150 tokens per second |
| Helped by | More GPU cores, Flash Attention, prompt caching | Faster memory, smaller or quantized models, MoE architectures, speculative decoding |
This split explains most surprises people have with local models. A long document makes you wait a long time before anything appears, then streams at normal speed: that is prefill. A huge model streams slowly even on a powerful GPU: that is decode hitting the memory wall. The attention side of prefill is covered in what is Flash Attention.
The metrics that matter
| Metric | Meaning | Good enough for |
|---|---|---|
| Time to first token (TTFT) | Delay between sending the prompt and seeing the first word | Under 1 s feels instant; several seconds is normal on long prompts |
| Tokens per second | Generation speed for one user | 10+ to read along comfortably; 30+ for agents and coding tools |
| Throughput | Total tokens per second across all concurrent users | What a serving team optimizes; see vLLM |
| Cost per million tokens | What an API charges, or what your hardware and power amortize to | Compare with the cost calculator |
Latency and throughput pull in opposite directions. Batching many users together raises throughput and lowers cost per token, but each individual user waits a little longer. A single-user desktop setup optimizes the first; an API provider optimizes the second. Units and benchmarks are explained in LLM tokens per second and what is a token.
What sets inference speed on your hardware
To generate one token, the processor must read every active weight once. So the ceiling on generation speed is memory bandwidth divided by the size of the weights. Taking Qwen 3 8B at 4-bit, 5 GB in our catalog:
| Hardware | Memory bandwidth | Theoretical ceiling for a 5 GB model |
|---|---|---|
| Desktop CPU, dual-channel DDR5 | ≈ 80 GB/s | ≈ 16 tok/s |
| Apple M4 (base) | 120 GB/s | ≈ 24 tok/s |
| RTX 4060 | 272 GB/s | ≈ 54 tok/s |
| RTX 5070 Ti | 896 GB/s | ≈ 180 tok/s |
| RTX 5090 | 1,792 GB/s | ≈ 360 tok/s |
Upper bounds: bandwidth ÷ model size. Real systems land below the ceiling because of overhead and the growing KV cache. Model size from the BestLLMfor catalog.
Three levers move that ceiling, and they are the whole toolbox of inference optimization:
- Make the model smaller in memory. Quantization stores weights in 4 or 8 bits instead of 16. See LLM quantization explained.
- Read fewer weights per token. Mixture-of-experts models activate only a fraction of their parameters for each token, so a 30B MoE can generate as fast as a 3B dense model while needing the memory of a 30B.
- Produce more than one token per pass. Speculative decoding drafts several tokens with a small model and verifies them in one pass of the large one.
All of this only applies if the model fits in fast memory in the first place. When it spills from VRAM into system RAM, speed drops by 5 to 20 times. That is the subject of what is VRAM and why VRAM matters more than TFLOPS.
Where inference runs
| Option | You pay | Strengths | Trade-offs |
|---|---|---|---|
| Cloud API | Per token | Largest models, zero setup, scales instantly | Data leaves your control; variable bill; rate limits |
| Your own server (on-prem or rented GPU) | Hardware or hourly rental | Control, predictable cost at volume, custom models | You operate it |
| Your desktop or laptop | Hardware once, then electricity | Private, offline, no per-token cost | Limited by your VRAM; smaller models |
| On-device (phone, NPU) | Nothing extra | Instant, private, works offline | Small models only; see what is an NPU |
The same open-weight model can run in all four places. That portability is the practical meaning of "open weights," and it is why inference, not training, is the part of AI that individuals and small teams can own. To start, follow how to run an LLM locally. Model sizes used on this page come from the BestLLMfor catalog, open through our public API (CC BY 4.0) and MCP server. For a formal treatment of serving efficiency, the vLLM paper is a readable entry point, and the llama.cpp and Ollama projects are where most local inference actually happens.
Frequently asked questions
What is the difference between training and inference?
Training finds a model's weights by learning from data; it is done once, on large GPU clusters. Inference uses those frozen weights to answer new inputs; it happens on every request and can run on anything from a server to a phone.
Does a model learn during inference?
No. The weights do not change. A chatbot that seems to remember earlier messages is simply re-reading them in its context window; nothing is stored in the model itself.
Why is inference expensive?
Each generated token requires a full pass through billions of weights, and popular services produce trillions of tokens. Over a model's lifetime, serving it typically costs far more than training it did.
Do I need a GPU for AI inference?
Not strictly. Small models run on a CPU at a few tokens per second. A GPU or an Apple Silicon Mac is 5 to 20 times faster because its memory bandwidth is much higher, which is what generation speed depends on.
What is an inference engine?
The software that loads model weights and executes them efficiently on your hardware. llama.cpp, Ollama (built on it), vLLM, TensorRT-LLM and MLX are inference engines, each optimized for a different situation.
What is time to first token?
The delay between sending a prompt and receiving the first token of the answer. It is dominated by prefill, so it grows with prompt length, and it is what makes a system feel responsive or sluggish.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.