BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
Guide · 2026-09-20

What Is AI Inference? Training vs Inference, Speed and Cost Explained

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Inference is the moment a trained AI model is actually used. It is where the cost, the latency and the hardware decisions live.

By Mohamed Meguedmi·Last updated 2026-09-20·8 min read·Tested on Windows, macOS, Linux

Key takeaways

  • Inference is using a trained model to produce an output. Every chatbot reply, image generation and transcription is an inference. Training builds the model; inference runs it.
  • Training happens once and costs millions. Inference happens billions of times, and over a model's life it is where most of the compute and money go.
  • For language models, inference has two phases: prefill (reading your prompt, compute-bound) and decode (writing the answer one token at a time, memory-bound).
  • Generation speed is capped by a simple ratio: memory bandwidth ÷ model size. That is why memory, not raw compute, drives hardware choices for local AI.
  • Inference can run in the cloud, on your own server, or on your laptop. The model and the math are the same; privacy, cost structure and latency are not.

Inference in plain English

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • 30-day refund

A machine-learning model is, in the end, a very large set of numbers called weights. Training is the process of finding good values for those numbers by showing the model enormous amounts of data and nudging the weights after each mistake. Inference is what comes after: the weights are frozen, you give the model a new input, and it computes an output. The model learns nothing during inference. It applies what it already knows.

The word comes from statistics, where to infer is to draw a conclusion from evidence. In AI it has narrowed to mean "running the model." When a provider talks about inference cost, an inference API or an inference chip, they mean the business of serving answers.

Training vs inference

TrainingInference
GoalFind the weightsUse the weights
How oftenOnce per model versionEvery single request
HardwareThousands of data-center GPUs with fast interconnectAnything from a GPU cluster to a phone
Numeric precision16-bit or mixed, to keep gradients stable8-bit or 4-bit is usually enough
Memory neededWeights + gradients + optimizer state: several times the model sizeWeights + working memory: about the model size
What you optimizeTotal time and cost to reach a quality targetLatency per request, throughput, cost per token
Who does itA few dozen labsEveryone

Between the two sits fine-tuning: a short, cheap round of extra training that adapts an existing model. It is training, not inference, but techniques like LoRA bring it within reach of a single consumer GPU.

The two phases of LLM inference

When you send a prompt to a language model, two very different things happen in sequence.

PrefillDecode
What it doesProcesses the whole prompt in one parallel pass and fills the KV cacheProduces the answer one token at a time, each needing a full pass through the model
Limited byCompute (arithmetic throughput)Memory bandwidth (how fast weights can be read)
You experience it asThe wait before the first wordThe speed at which text streams
Typical speed on a consumer GPUHundreds to thousands of tokens per second10 to 150 tokens per second
Helped byMore GPU cores, Flash Attention, prompt cachingFaster memory, smaller or quantized models, MoE architectures, speculative decoding

This split explains most surprises people have with local models. A long document makes you wait a long time before anything appears, then streams at normal speed: that is prefill. A huge model streams slowly even on a powerful GPU: that is decode hitting the memory wall. The attention side of prefill is covered in what is Flash Attention.

The metrics that matter

MetricMeaningGood enough for
Time to first token (TTFT)Delay between sending the prompt and seeing the first wordUnder 1 s feels instant; several seconds is normal on long prompts
Tokens per secondGeneration speed for one user10+ to read along comfortably; 30+ for agents and coding tools
ThroughputTotal tokens per second across all concurrent usersWhat a serving team optimizes; see vLLM
Cost per million tokensWhat an API charges, or what your hardware and power amortize toCompare with the cost calculator

Latency and throughput pull in opposite directions. Batching many users together raises throughput and lowers cost per token, but each individual user waits a little longer. A single-user desktop setup optimizes the first; an API provider optimizes the second. Units and benchmarks are explained in LLM tokens per second and what is a token.

What sets inference speed on your hardware

To generate one token, the processor must read every active weight once. So the ceiling on generation speed is memory bandwidth divided by the size of the weights. Taking Qwen 3 8B at 4-bit, 5 GB in our catalog:

HardwareMemory bandwidthTheoretical ceiling for a 5 GB model
Desktop CPU, dual-channel DDR5≈ 80 GB/s≈ 16 tok/s
Apple M4 (base)120 GB/s≈ 24 tok/s
RTX 4060272 GB/s≈ 54 tok/s
RTX 5070 Ti896 GB/s≈ 180 tok/s
RTX 50901,792 GB/s≈ 360 tok/s

Upper bounds: bandwidth ÷ model size. Real systems land below the ceiling because of overhead and the growing KV cache. Model size from the BestLLMfor catalog.

Three levers move that ceiling, and they are the whole toolbox of inference optimization:

  • Make the model smaller in memory. Quantization stores weights in 4 or 8 bits instead of 16. See LLM quantization explained.
  • Read fewer weights per token. Mixture-of-experts models activate only a fraction of their parameters for each token, so a 30B MoE can generate as fast as a 3B dense model while needing the memory of a 30B.
  • Produce more than one token per pass. Speculative decoding drafts several tokens with a small model and verifies them in one pass of the large one.

All of this only applies if the model fits in fast memory in the first place. When it spills from VRAM into system RAM, speed drops by 5 to 20 times. That is the subject of what is VRAM and why VRAM matters more than TFLOPS.

Where inference runs

OptionYou payStrengthsTrade-offs
Cloud APIPer tokenLargest models, zero setup, scales instantlyData leaves your control; variable bill; rate limits
Your own server (on-prem or rented GPU)Hardware or hourly rentalControl, predictable cost at volume, custom modelsYou operate it
Your desktop or laptopHardware once, then electricityPrivate, offline, no per-token costLimited by your VRAM; smaller models
On-device (phone, NPU)Nothing extraInstant, private, works offlineSmall models only; see what is an NPU

The same open-weight model can run in all four places. That portability is the practical meaning of "open weights," and it is why inference, not training, is the part of AI that individuals and small teams can own. To start, follow how to run an LLM locally. Model sizes used on this page come from the BestLLMfor catalog, open through our public API (CC BY 4.0) and MCP server. For a formal treatment of serving efficiency, the vLLM paper is a readable entry point, and the llama.cpp and Ollama projects are where most local inference actually happens.

Frequently asked questions

What is the difference between training and inference?

Training finds a model's weights by learning from data; it is done once, on large GPU clusters. Inference uses those frozen weights to answer new inputs; it happens on every request and can run on anything from a server to a phone.

Does a model learn during inference?

No. The weights do not change. A chatbot that seems to remember earlier messages is simply re-reading them in its context window; nothing is stored in the model itself.

Why is inference expensive?

Each generated token requires a full pass through billions of weights, and popular services produce trillions of tokens. Over a model's lifetime, serving it typically costs far more than training it did.

Do I need a GPU for AI inference?

Not strictly. Small models run on a CPU at a few tokens per second. A GPU or an Apple Silicon Mac is 5 to 20 times faster because its memory bandwidth is much higher, which is what generation speed depends on.

What is an inference engine?

The software that loads model weights and executes them efficiently on your hardware. llama.cpp, Ollama (built on it), vLLM, TensorRT-LLM and MLX are inference engines, each optimized for a different situation.

What is time to first token?

The delay between sending a prompt and receiving the first token of the answer. It is dominated by prefill, so it grows with prompt length, and it is what makes a system feel responsive or sluggish.

Recommended hardware

A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.

Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Did this guide help?

Found an error or have feedback? Let us know — it helps everyone who reads this guide.