Intermediate 11 minPerformance

Optimize thermals on inference

Direct response

For continuous inference, the most cost-effective thermal lever is to cap GPU power: according to a measurement published on a RTX 3090, going from 350 W to 250 W costs less than 3% of generation speed, while below that threshold speed collapses. First check what is actually limiting the card (temperature or power) with nvidia-smi before touching the fans, voltage, or thermal paste.

A machine that answers queries all day heats up differently from a video game: constant load, heavily used memory, and an enclosed case. This page explains how to tell whether your card is throttling, which setting to try first, what the power limit actually costs in tokens per second, and when maintenance (dust, thermal paste, case) is the real answer.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why thermals matter for an inference server

A GPU protects its chip: when the temperature or power exceeds what its firmware allows, it lowers its clock speeds. This mechanism is called throttling. The nvidia-smi documentation distinguishes its causes: a software temperature-related slowdown (“sw thermal slowdown”), a more severe hardware slowdown that reduces clock speeds by a factor of 2 or more (“hw thermal slowdown”), and a power limit (“sw power cap”). If all these reasons are shown as inactive, the clock speeds run as high as possible.

The threshold isn’t the same on every card. NVIDIA defines it as “GPU Slowdown Temp,” the temperature at which the hardware starts optimizing clock speeds to cool down, and exposes it in nvidia-smi. So don’t rely on a general figure like “83 °C”: read your card’s value. Throttling causes lower speeds and, if the fans compensate, more noise; over the long term, a card that runs hot all the time ages faster, but we don’t have a reliable figure to cite on this point.

i
Decode and prompt: two different sensitivities
Text generation is limited by memory bandwidth, so it is relatively insensitive to core frequency. Prompt processing is compute-limited, so it is more sensitive to any drop in frequency. A server receiving long documents or running agent loops will suffer more from a power cap than an interactive chat will.

#Measure before tuning: what the card really tells you

The starting point is determining whether the card is throttling and why. Two commands are enough on a NVIDIA card. The first continuously displays the temperature, power, core frequency, and utilization; the second details the throttling reasons and the temperature thresholds for your model.

Monitor under load (one line per second)
nvidia-smi --query-gpu=timestamp,temperature.gpu,power.draw,clocks.current.graphics,clocks.current.memory,utilization.gpu --format=csv -l 1
Reasons for slowdown and temperature thresholds
nvidia-smi -q -d PERFORMANCE
nvidia-smi -q -d TEMPERATURE
AMD cards
rocm-smi --showtemp --showpower --showuse

Run these commands during a generation lasting several minutes, not while idle. Two readings matter. If the core frequency drops while the temperature rises and a thermal reason is active, it's a cooling problem. If the core stays high but a power limit is active, the card is limited by its power consumption: lowering the limit further won't have a noticeable effect on generation speed, but it will reduce heat.

What to read, what to conclude
Observation during a long generationInterpretationFirst action
Temperature stable below the throttling threshold, frequency stableNo thermal issuesChange nothing; optionally cap power to reduce noise
Temperature reaches the threshold, frequency drops, thermal throttling kicks inInsufficient coolingCase airflow, dust, fan curve
Low frequency, active “power cap” reason, normal temperaturePower limit reachedNormal; raise the limit only if throughput is insufficient
Speed difference between the first and tenth exchangeGradual warm-upMeasure under a sustained load, not at startup

#The power limit: the setting with the biggest payoff

Capping power is simpler than adjusting the voltage curve, and reversible. For memory-bound inference, the loss is small up to a certain threshold, then becomes abrupt. The runaihome blog, which measured a RTX 3090 with a dense 27-billion-parameter model, reports that from 350 W to 250 W the card loses less than 3% of its speed for 100 W less power, and that at 200 W the core frequency collapses and speed drops by 35%.

Power limit and generation speed, RTX 3090 (measurement published by runaihome)
LimitMeasured effect on generation
350 W (default)Reference
300 WLess than 3% of the reference
250 W31.7 tokens/s, less than 3% below the reference
200 W20.6 t/s, or 35% less: the threshold you must not cross

For a RTX 4090, the same blog cites Tom's Hardware, according to which limiting power to 70% (about 315 W) retains about 94% of performance. These figures come from secondary sources and a specific card: do not apply them as-is to yours. The reliable method is to measure your throughput during a long generation, lowering the limit in 25- to 50-W increments and stopping before the drop.

Cap power (Linux, adapt to your card)
# Lire la limite par défaut, minimale et maximale
nvidia-smi -q -d POWER

# Appliquer 250 W sur la carte 0 (nécessite les droits root)
sudo nvidia-smi -i 0 -pl 250

The nvidia-smi documentation specifies that the value must be between the minimum and maximum limits reported by the card, and that the command requires root privileges. After a reboot, check that the limit is still applied: if it has reverted to the default value, create a systemd unit to reapply it at startup.

/etc/systemd/system/nvidia-power-limit.service
[Unit]
Description=Limite de puissance GPU
After=multi-user.target

[Service]
Type=oneshot
ExecStart=/usr/bin/nvidia-smi -i 0 -pl 250
RemainAfterExit=yes

[Install]
WantedBy=multi-user.target

Enable it with sudo systemctl enable --now nvidia-power-limit.service.

#Fan curve: ramp up early rather than abruptly late

The factory curve prioritizes silence: slow fans up to a high temperature, then a ramp. For a sustained load, a curve that starts earlier keeps the card cooler at comparable noise levels because it avoids speed spikes. On Windows, MSI Afterburner adjusts the curve; on Linux, LACT is a graphical AMD, NVIDIA, and Intel GPU control application that includes fan control.

Example curve for a sustained workload (validate on your card)
TemperatureFan speed
50 °C40 %
60 °C55 %
70 °C75 %
78 °C90 %
Throttling threshold below 5 °C100 %

These values are a starting point, not a manufacturer recommendation. Set the final step based on your card’s “GPU Slowdown Temp” value, and judge the result by the noise you can tolerate in the room. On macOS, Apple manages the fans: there is no curve adjustment.

#Undervolting: finer-tuned, longer to validate

Undervolting means achieving the same clock speed with lower voltage, and therefore less heat. In Windows, configure it in MSI Afterburner (the voltage-frequency curve); in Linux, the common approach is to lock frequencies with the nvidia-smi option that sets the core frequency range. The benefit depends on the card: every chip has its own headroom, and an overly aggressive setting causes errors or crashes later, not immediately.

  1. 01
    Start with a power limit that works
    Start with the power limit: it's the safest gain. Add an undervolt only if the card remains too hot or too noisy.
  2. 02
    Set a maximum frequency
    Lock the core frequency below the maximum boost frequency, for example with nvidia-smi -lgc followed by a frequency range in MHz.
  3. 03
    Test for a long time
    Run continuous generation for at least 30 minutes. Instability appears as CUDA errors, driver crashes, or incoherent output.
  4. 04
    Compare
    Measure speed and temperature before and after on the same model and prompt, then keep the setting only if the speed loss is negligible.
  5. 05
    Cancel if needed
    Reset the frequencies (nvidia-smi -rgc) or restart to restore the factory settings.
!
Undervolting is not guaranteed
Unlike a power limit, an overly aggressive undervolt can make the system intermittently unstable. On an unattended server, stick to the power limit alone unless you’ve tested it for several hours.

#Case and airflow: the improvement is measured in degrees

Airflow path
A clear front-to-back airflow path, with fresh air directed toward the card's fans. Cabling that obstructs airflow costs you degrees.
Filters and dust
A clogged filter reduces airflow: clean it regularly, and blow dust out of the card's heatsink with compressed air, keeping the fans from spinning so they do not overspeed.
Clearance around the card
Two cards mounted tightly together in a multi-GPU setup heat up the first one; space them apart or add a side fan.
Environment
A room at 30 °C versus 20 °C adds those ten degrees to the card: a server in a closed closet is a thermal problem before it is a card problem.

To evaluate a setting accurately, keep the protocol constant: use the same model, the same long prompt, and the same generation time, changing only one parameter at a time. Record the peak temperature, core frequency, and average throughput at the end of generation, not the beginning, because a cool card always produces better numbers than the same card after twenty minutes. Repeat the measurement at the same ambient temperature; otherwise, the comparison is skewed by the room, not the setting.

#Repasting: a last resort, with its risks

After several years, the thermal paste and pads on a used card may have dried out. The symptom is a higher temperature than when you bought it under the same load, despite a clean case. Taking apart a graphics card voids the warranty, can damage the pads and chip, and requires care. Only do it after ruling out simpler causes: dust, airflow, power limit, and ambient temperature. If you're not sure, take it to a service provider; compare the cost with that of a replacement card.

#Laptops and Macs: the levers are different

On a laptop, you can’t adjust the voltage or thermal paste. The useful lever is the power mode. Apple describes two modes: Low Power mode, which, on macOS Sequoia 15.1 and later, also reduces fan noise on models with active cooling, and High Power mode, which lets the fans spin faster to maintain performance under very intensive workloads, with additional noise. For a fanless MacBook Air, the only levers are a heat-dissipating stand, AC power, and a smaller model.

#Frequently asked questions

FAQ
At what temperature does a graphics card start throttling?+
It depends on the model. NVIDIA defines a “GPU Slowdown Temp” in nvidia-smi—the temperature at which the hardware adjusts frequencies to cool down. The nvidia-smi -q -d TEMPERATURE command displays your card’s value. Do not rely on a general threshold such as 83 °C without checking it.
Does limiting power slow inference down significantly?+
Not much, up to a threshold. A measurement published on a RTX 3090 indicates less than 3% performance loss from 350 W to 250 W during generation, followed by a 35% drop at 200 W. Prompt processing, which is compute-bound, is more sensitive. Measure on your card with your workload.
How can you tell if your GPU is throttling?+
Start a long generation and monitor the core frequency with nvidia-smi. If it drops while the temperature rises, check nvidia-smi -q -d PERFORMANCE: it indicates whether a thermal or power reason is active. If all reasons are inactive, the frequencies are running as high as possible.
Should you replace your card's thermal paste?+
Rarely. First rule out dust, case airflow, ambient temperature, and the power limit. Disassembly voids the warranty and poses risks to the pads and chip. It is justified only if temperatures have risen significantly on a clean card, after everything else has been checked.
Undervolting or power limiting: which should you choose?+
Start with the power limit: it is simple, reversible, and carries no instability risk. Undervolting requires lengthy testing and can cause intermittent crashes, especially on an unattended server. Add it only if the card remains too hot or noisy with the power limit alone.
Does a MacBook have thermal settings for local AI?+
Not often. Apple offers power modes: Power Saver, and High Power on certain models with faster fans but more noise. On a fanless MacBook Air, place the machine on a stand that dissipates heat, plug it into power, and choose a smaller model for long generations.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.