Optimize thermals on inference
For continuous inference, the most cost-effective thermal lever is to cap GPU power: according to a measurement published on a RTX 3090, going from 350 W to 250 W costs less than 3% of generation speed, while below that threshold speed collapses. First check what is actually limiting the card (temperature or power) with nvidia-smi before touching the fans, voltage, or thermal paste.
A machine that answers queries all day heats up differently from a video game: constant load, heavily used memory, and an enclosed case. This page explains how to tell whether your card is throttling, which setting to try first, what the power limit actually costs in tokens per second, and when maintenance (dust, thermal paste, case) is the real answer.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Why thermals matter for an inference server
A GPU protects its chip: when the temperature or power exceeds what its firmware allows, it lowers its clock speeds. This mechanism is called throttling. The nvidia-smi documentation distinguishes its causes: a software temperature-related slowdown (“sw thermal slowdown”), a more severe hardware slowdown that reduces clock speeds by a factor of 2 or more (“hw thermal slowdown”), and a power limit (“sw power cap”). If all these reasons are shown as inactive, the clock speeds run as high as possible.
The threshold isn’t the same on every card. NVIDIA defines it as “GPU Slowdown Temp,” the temperature at which the hardware starts optimizing clock speeds to cool down, and exposes it in nvidia-smi. So don’t rely on a general figure like “83 °C”: read your card’s value. Throttling causes lower speeds and, if the fans compensate, more noise; over the long term, a card that runs hot all the time ages faster, but we don’t have a reliable figure to cite on this point.
#Measure before tuning: what the card really tells you
The starting point is determining whether the card is throttling and why. Two commands are enough on a NVIDIA card. The first continuously displays the temperature, power, core frequency, and utilization; the second details the throttling reasons and the temperature thresholds for your model.
Run these commands during a generation lasting several minutes, not while idle. Two readings matter. If the core frequency drops while the temperature rises and a thermal reason is active, it's a cooling problem. If the core stays high but a power limit is active, the card is limited by its power consumption: lowering the limit further won't have a noticeable effect on generation speed, but it will reduce heat.
| Observation during a long generation | Interpretation | First action |
|---|---|---|
| Temperature stable below the throttling threshold, frequency stable | No thermal issues | Change nothing; optionally cap power to reduce noise |
| Temperature reaches the threshold, frequency drops, thermal throttling kicks in | Insufficient cooling | Case airflow, dust, fan curve |
| Low frequency, active “power cap” reason, normal temperature | Power limit reached | Normal; raise the limit only if throughput is insufficient |
| Speed difference between the first and tenth exchange | Gradual warm-up | Measure under a sustained load, not at startup |
#The power limit: the setting with the biggest payoff
Capping power is simpler than adjusting the voltage curve, and reversible. For memory-bound inference, the loss is small up to a certain threshold, then becomes abrupt. The runaihome blog, which measured a RTX 3090 with a dense 27-billion-parameter model, reports that from 350 W to 250 W the card loses less than 3% of its speed for 100 W less power, and that at 200 W the core frequency collapses and speed drops by 35%.
| Limit | Measured effect on generation |
|---|---|
| 350 W (default) | Reference |
| 300 W | Less than 3% of the reference |
| 250 W | 31.7 tokens/s, less than 3% below the reference |
| 200 W | 20.6 t/s, or 35% less: the threshold you must not cross |
For a RTX 4090, the same blog cites Tom's Hardware, according to which limiting power to 70% (about 315 W) retains about 94% of performance. These figures come from secondary sources and a specific card: do not apply them as-is to yours. The reliable method is to measure your throughput during a long generation, lowering the limit in 25- to 50-W increments and stopping before the drop.
The nvidia-smi documentation specifies that the value must be between the minimum and maximum limits reported by the card, and that the command requires root privileges. After a reboot, check that the limit is still applied: if it has reverted to the default value, create a systemd unit to reapply it at startup.
Enable it with sudo systemctl enable --now nvidia-power-limit.service.
#Fan curve: ramp up early rather than abruptly late
The factory curve prioritizes silence: slow fans up to a high temperature, then a ramp. For a sustained load, a curve that starts earlier keeps the card cooler at comparable noise levels because it avoids speed spikes. On Windows, MSI Afterburner adjusts the curve; on Linux, LACT is a graphical AMD, NVIDIA, and Intel GPU control application that includes fan control.
| Temperature | Fan speed |
|---|---|
| 50 °C | 40 % |
| 60 °C | 55 % |
| 70 °C | 75 % |
| 78 °C | 90 % |
| Throttling threshold below 5 °C | 100 % |
These values are a starting point, not a manufacturer recommendation. Set the final step based on your card’s “GPU Slowdown Temp” value, and judge the result by the noise you can tolerate in the room. On macOS, Apple manages the fans: there is no curve adjustment.
#Undervolting: finer-tuned, longer to validate
Undervolting means achieving the same clock speed with lower voltage, and therefore less heat. In Windows, configure it in MSI Afterburner (the voltage-frequency curve); in Linux, the common approach is to lock frequencies with the nvidia-smi option that sets the core frequency range. The benefit depends on the card: every chip has its own headroom, and an overly aggressive setting causes errors or crashes later, not immediately.
- 01Start with a power limit that worksStart with the power limit: it's the safest gain. Add an undervolt only if the card remains too hot or too noisy.
- 02Set a maximum frequencyLock the core frequency below the maximum boost frequency, for example with nvidia-smi -lgc followed by a frequency range in MHz.
- 03Test for a long timeRun continuous generation for at least 30 minutes. Instability appears as CUDA errors, driver crashes, or incoherent output.
- 04CompareMeasure speed and temperature before and after on the same model and prompt, then keep the setting only if the speed loss is negligible.
- 05Cancel if neededReset the frequencies (nvidia-smi -rgc) or restart to restore the factory settings.
#Case and airflow: the improvement is measured in degrees
- Airflow path
- A clear front-to-back airflow path, with fresh air directed toward the card's fans. Cabling that obstructs airflow costs you degrees.
- Filters and dust
- A clogged filter reduces airflow: clean it regularly, and blow dust out of the card's heatsink with compressed air, keeping the fans from spinning so they do not overspeed.
- Clearance around the card
- Two cards mounted tightly together in a multi-GPU setup heat up the first one; space them apart or add a side fan.
- Environment
- A room at 30 °C versus 20 °C adds those ten degrees to the card: a server in a closed closet is a thermal problem before it is a card problem.
To evaluate a setting accurately, keep the protocol constant: use the same model, the same long prompt, and the same generation time, changing only one parameter at a time. Record the peak temperature, core frequency, and average throughput at the end of generation, not the beginning, because a cool card always produces better numbers than the same card after twenty minutes. Repeat the measurement at the same ambient temperature; otherwise, the comparison is skewed by the room, not the setting.
#Repasting: a last resort, with its risks
After several years, the thermal paste and pads on a used card may have dried out. The symptom is a higher temperature than when you bought it under the same load, despite a clean case. Taking apart a graphics card voids the warranty, can damage the pads and chip, and requires care. Only do it after ruling out simpler causes: dust, airflow, power limit, and ambient temperature. If you're not sure, take it to a service provider; compare the cost with that of a replacement card.
#Laptops and Macs: the levers are different
On a laptop, you can’t adjust the voltage or thermal paste. The useful lever is the power mode. Apple describes two modes: Low Power mode, which, on macOS Sequoia 15.1 and later, also reduces fan noise on models with active cooling, and High Power mode, which lets the fans spin faster to maintain performance under very intensive workloads, with additional noise. For a fanless MacBook Air, the only levers are a heat-dissipating stand, AC power, and a smaller model.
- Choose your GPU for local AI
- 32 GB AI PC build
- Multi-GPU setup with llama.cpp
- Source: nvidia-smi documentation
- Source: power-limit measurements (runaihome)
- Source: Apple, Mac power modes
#Frequently asked questions
At what temperature does a graphics card start throttling?+
Does limiting power slow inference down significantly?+
How can you tell if your GPU is throttling?+
Should you replace your card's thermal paste?+
Undervolting or power limiting: which should you choose?+
Does a MacBook have thermal settings for local AI?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.