Advanced 9 minAnnouncements and updates

Intel Crescent Island: 160 GB for inference LLM

Direct response

Crescent Island is a data-center GPU announced by Intel on October 14, 2025, equipped with 160 GB of LPDDR5X memory in the reference configuration (up to 480 GB from some partners), designed for inference in air-cooled enterprise servers. It is neither a consumer product nor a card you can install in a personal PC: samples are not expected until the second half of 2026, and no price or commercial release date has been announced.

Crescent Island is the codename for Intel's inference GPU for data centers, unveiled in late 2025, with one main selling point: lots of memory that is cheaper and less power-hungry than the HBM used by competing accelerators. This article explains what Intel has confirmed, what remains to be verified, where this memory choice stands relative to accelerators already on the market, and why this is not a GPU for your personal PC.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What Intel announced

Intel presented Crescent Island on October 14, 2025, at the OCP Global Summit: a data center GPU with 160 GB of LPDDR5X memory, built on a Xe3P microarchitecture optimized for performance per watt. Intel explicitly says it is targeting “tokens-as-a-service” providers and inference workloads in air-cooled enterprise servers rather than liquid-cooled ones.

i
Inference only
Intel positions Crescent Island for inference, not model training. LPDDR5X memory, which costs less than HBM, is the choice that makes this positioning possible.

Samples intended for customers are expected in the second half of 2026. Intel has not announced any pricing, a consumer commercial availability date, or a confirmed OEM partner at this stage.

#Memory capacity isn't throughput

160 GB (up to 480 GB according to some partners cited by Tom’s Hardware) lets you load large models entirely into GPU memory, without splitting them across multiple cards. But capacity says nothing about memory bandwidth, which is the factor that actually limits token-by-token generation throughput. LPDDR5X is structurally slower in bandwidth than the HBM used by high-end accelerators: Intel is making an explicit trade-off between capacity/cost and maximum throughput.

#LPDDR5X versus the HBM in today's accelerators

Intel publishes no bandwidth figures for Crescent Island, so it is impossible to calculate a theoretical tokens/s ceiling as we can for the site's other coverage. What can be measured, however, is the bandwidth gap between memory families on accelerators already on the market — a useful order of magnitude for understanding the tradeoff Intel is making, and for avoiding comparisons between a future Crescent Island spec sheet and an H100 or MI300X as if they belonged to the same throughput category.

Memory bandwidth: current HBM accelerators (to put the gap with LPDDR5X in context)
AcceleratorMemoryAdvertised bandwidth
NVIDIA H100 SXM80 GB HBM33.35 TB/s
AMD Instinct MI300X192 GB HBM35.3 TB/s
Intel Crescent Island (reference)160 GB LPDDR5XNot published by Intel

Without an official figure for Crescent Island, the only verifiable reference comes from LPDDR5X implementations already sold elsewhere: the NVIDIA DGX Spark, with 128 GB of LPDDR5X, advertises 273 GB/s, or about 0.27 TB/s—more than ten times less than the MI300X. A server GPU like Crescent Island could use a wider bus and therefore deliver a higher figure, but until Intel publishes it, you should assume a gap of this order of magnitude, not parity with HBM. That's the price of abundant capacity and lower cost: Crescent Island targets workloads where loading a large model in its entirety matters more than delivering peak throughput, not replacing an H100 or MI300X for every workload.

#Crescent Island: the current state

Confirmed / unconfirmed / to be verified in the initial tests
PointStatusSource
160 GB of LPDDR5X, Xe3P architectureConfirmed by IntelIntel press release, 14/10/2025
Target: inference, air-cooled servers, tokens-as-a-serviceConfirmed by IntelIntel press release
Customer samples in the second half of 2026Confirmed by IntelIntel press release
Configurations up to 480 GB from some partnersReported by the technical pressTom's Hardware
Actual memory bandwidth, tokens/sNot measuredFirst independent benchmarks
Price, sale date, OEM availabilityUnknownWhat not to anticipate

#Who it's for, and what it's for

Crescent Island targets data center operators and inference service providers, not individuals or small teams building a development machine. The server form factor, air cooling designed for density, and explicit “tokens-as-a-service” focus point to an infrastructure product, comparable in role to a card intended for large-scale LLM serving rather than a workstation.

For a company considering internal LLM inference in the medium term, the value will come from a better cost-per-GB-of-memory and cost-per-watt ratio than current HBM GPUs—a calculation that can only be made once the actual price and performance are published.

This positioning reflects a broader trend seen among several manufacturers: offering abundant but slower memory (LPDDR) instead of scarce but very fast memory (HBM) for inference workloads where loading the entire model matters more than peak throughput. It is the same tradeoff found, at a different scale, in consumer unified-memory APUs such as the Ryzen AI Max+ 395 (LPDDR5X, 256 GB/s): more capacity per dollar spent, in exchange for lower peak throughput than a dedicated HBM GPU.

In practice, an operator serving several models or several requests in parallel (“tokens-as-a-service”) often benefits more from distributing the load across many simultaneous requests, where per-request latency matters less than aggregate throughput and cost per token served. This is exactly the economic calculation that Crescent Island's generous LPDDR5X capacity is designed to optimize, unlike a use case requiring the lowest possible latency for a single request, where HBM's raw bandwidth remains decisive.

#Drivers and frameworks: what is still missing

No source consulted currently confirms certified support for vLLM, SGLang, or a popular LLM inference runtime for Crescent Island specifically. The positive point is that vLLM already officially maintains an XPU backend for Intel GPUs, with supported hardware listed as “Intel Data Center GPU, Intel ARC GPU” in its own installation documentation. This is the most likely integration path for Crescent Island once samples are distributed, but nothing guarantees day-one support for this specific chip before its actual release. Intel also maintains an OpenVINO backend for llama.cpp on its current GPUs, reinforcing this coherent software trajectory without guaranteeing it for Crescent Island.

!
Don’t confuse it with a consumer GPU
Crescent Island is not a gaming GPU or a card you can install in a personal computer. Any comparison of price or availability with a consumer card would be misleading until Intel has published either one.
→
What to check in the first tests
Actual memory bandwidth (no figures published by Intel to date), tokens/s measured on reference open models (Llama, Qwen), price per card and per GB, the actual maturity of vLLM and SGLang on this specific chip, and actual availability from cloud or OEM providers.

#Buy now or wait?

Need for enterprise LLM inference today
Current HBM GPUs (H100, MI300X) or existing multi-GPU solutions remain the only measurable and deliverable options; Crescent Island is neither available for purchase nor priced in terms of cost or performance to date.
Infrastructure project looking ahead to 2027
Follow the first independent benchmarks of Crescent Island once the samples are released, especially actual tokens/s and cost per GB.
Personal or small-team use
This GPU is not relevant to this use case: prioritize a consumer GPU configuration or a unified-memory APU already documented on this site.
Decision based on the announced 480 GB
This figure corresponds to a maximum configuration reported by the press for certain partners, not the 160 GB reference configuration confirmed by Intel.
Frequently asked questions
Is Crescent Island available for purchase?+
No. Intel expects customer samples in the second half of 2026, presented on October 14, 2025, at the OCP Global Summit. No price, commercial release date, or confirmed OEM partner has been announced to date: this is a preparation project, not an available product.
Can Crescent Island replace a consumer GPU for local AI?+
No. It’s a data-center GPU designed for air-cooled enterprise servers, targeting token-as-a-service providers—not a card compatible with a personal PC. For local AI on a workstation, a consumer GPU or a unified-memory APU remains the right choice.
Does 160 GB of LPDDR5X mean faster models?+
That means larger models can be loaded entirely into memory, not necessarily that they are faster: LPDDR5X offers significantly less bandwidth than the HBM3 in high-end accelerators (3.35 to 5.3 TB/s for H100 and MI300X). Intel is making a deliberate trade-off between capacity, cost, and maximum throughput.
Does Crescent Island go up to 480 GB?+
The reference configuration Intel announced on October 14, 2025, has 160 GB of LPDDR5X. The 480 GB figure, with a 350 W TDP under air cooling, comes from configurations reported by the technical press (Tom's Hardware) for certain board partners, not from Intel's own published reference product specifications — do not present it as the standard configuration.
Will Crescent Island be compatible with vLLM or llama.cpp?+
Probably vLLM: its official XPU backend already supports Intel data center GPUs in its documentation. Intel also maintains OpenVINO support for llama.cpp. However, no source confirms certified support for this specific chip before its actual release.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.