Intel Crescent Island: 160 GB for inference LLM
Crescent Island is a data-center GPU announced by Intel on October 14, 2025, equipped with 160 GB of LPDDR5X memory in the reference configuration (up to 480 GB from some partners), designed for inference in air-cooled enterprise servers. It is neither a consumer product nor a card you can install in a personal PC: samples are not expected until the second half of 2026, and no price or commercial release date has been announced.
Crescent Island is the codename for Intel's inference GPU for data centers, unveiled in late 2025, with one main selling point: lots of memory that is cheaper and less power-hungry than the HBM used by competing accelerators. This article explains what Intel has confirmed, what remains to be verified, where this memory choice stands relative to accelerators already on the market, and why this is not a GPU for your personal PC.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What Intel announced
Intel presented Crescent Island on October 14, 2025, at the OCP Global Summit: a data center GPU with 160 GB of LPDDR5X memory, built on a Xe3P microarchitecture optimized for performance per watt. Intel explicitly says it is targeting “tokens-as-a-service” providers and inference workloads in air-cooled enterprise servers rather than liquid-cooled ones.
Samples intended for customers are expected in the second half of 2026. Intel has not announced any pricing, a consumer commercial availability date, or a confirmed OEM partner at this stage.
#Memory capacity isn't throughput
160 GB (up to 480 GB according to some partners cited by Tom’s Hardware) lets you load large models entirely into GPU memory, without splitting them across multiple cards. But capacity says nothing about memory bandwidth, which is the factor that actually limits token-by-token generation throughput. LPDDR5X is structurally slower in bandwidth than the HBM used by high-end accelerators: Intel is making an explicit trade-off between capacity/cost and maximum throughput.
#LPDDR5X versus the HBM in today's accelerators
Intel publishes no bandwidth figures for Crescent Island, so it is impossible to calculate a theoretical tokens/s ceiling as we can for the site's other coverage. What can be measured, however, is the bandwidth gap between memory families on accelerators already on the market — a useful order of magnitude for understanding the tradeoff Intel is making, and for avoiding comparisons between a future Crescent Island spec sheet and an H100 or MI300X as if they belonged to the same throughput category.
| Accelerator | Memory | Advertised bandwidth |
|---|---|---|
| NVIDIA H100 SXM | 80 GB HBM3 | 3.35 TB/s |
| AMD Instinct MI300X | 192 GB HBM3 | 5.3 TB/s |
| Intel Crescent Island (reference) | 160 GB LPDDR5X | Not published by Intel |
Without an official figure for Crescent Island, the only verifiable reference comes from LPDDR5X implementations already sold elsewhere: the NVIDIA DGX Spark, with 128 GB of LPDDR5X, advertises 273 GB/s, or about 0.27 TB/s—more than ten times less than the MI300X. A server GPU like Crescent Island could use a wider bus and therefore deliver a higher figure, but until Intel publishes it, you should assume a gap of this order of magnitude, not parity with HBM. That's the price of abundant capacity and lower cost: Crescent Island targets workloads where loading a large model in its entirety matters more than delivering peak throughput, not replacing an H100 or MI300X for every workload.
#Crescent Island: the current state
| Point | Status | Source |
|---|---|---|
| 160 GB of LPDDR5X, Xe3P architecture | Confirmed by Intel | Intel press release, 14/10/2025 |
| Target: inference, air-cooled servers, tokens-as-a-service | Confirmed by Intel | Intel press release |
| Customer samples in the second half of 2026 | Confirmed by Intel | Intel press release |
| Configurations up to 480 GB from some partners | Reported by the technical press | Tom's Hardware |
| Actual memory bandwidth, tokens/s | Not measured | First independent benchmarks |
| Price, sale date, OEM availability | Unknown | What not to anticipate |
#Who it's for, and what it's for
Crescent Island targets data center operators and inference service providers, not individuals or small teams building a development machine. The server form factor, air cooling designed for density, and explicit “tokens-as-a-service” focus point to an infrastructure product, comparable in role to a card intended for large-scale LLM serving rather than a workstation.
For a company considering internal LLM inference in the medium term, the value will come from a better cost-per-GB-of-memory and cost-per-watt ratio than current HBM GPUs—a calculation that can only be made once the actual price and performance are published.
This positioning reflects a broader trend seen among several manufacturers: offering abundant but slower memory (LPDDR) instead of scarce but very fast memory (HBM) for inference workloads where loading the entire model matters more than peak throughput. It is the same tradeoff found, at a different scale, in consumer unified-memory APUs such as the Ryzen AI Max+ 395 (LPDDR5X, 256 GB/s): more capacity per dollar spent, in exchange for lower peak throughput than a dedicated HBM GPU.
In practice, an operator serving several models or several requests in parallel (“tokens-as-a-service”) often benefits more from distributing the load across many simultaneous requests, where per-request latency matters less than aggregate throughput and cost per token served. This is exactly the economic calculation that Crescent Island's generous LPDDR5X capacity is designed to optimize, unlike a use case requiring the lowest possible latency for a single request, where HBM's raw bandwidth remains decisive.
#Drivers and frameworks: what is still missing
No source consulted currently confirms certified support for vLLM, SGLang, or a popular LLM inference runtime for Crescent Island specifically. The positive point is that vLLM already officially maintains an XPU backend for Intel GPUs, with supported hardware listed as “Intel Data Center GPU, Intel ARC GPU” in its own installation documentation. This is the most likely integration path for Crescent Island once samples are distributed, but nothing guarantees day-one support for this specific chip before its actual release. Intel also maintains an OpenVINO backend for llama.cpp on its current GPUs, reinforcing this coherent software trajectory without guaranteeing it for Crescent Island.
#Buy now or wait?
- Need for enterprise LLM inference today
- Current HBM GPUs (H100, MI300X) or existing multi-GPU solutions remain the only measurable and deliverable options; Crescent Island is neither available for purchase nor has its price or performance been quantified to date.
- Infrastructure project looking ahead to 2027
- Follow the first independent benchmarks of Crescent Island once the samples are released, especially actual tokens/s and cost per GB.
- Personal or small-team use
- This GPU is not relevant to this use case: prioritize a consumer GPU configuration or a unified-memory APU already documented on this site.
- Decision based on the announced 480 GB
- This figure corresponds to a maximum configuration reported by the press for certain partners, not the 160 GB reference configuration confirmed by Intel.
- Available alternative: multi-GPU LLM with llama.cpp
- Choosing a GPU for a local LLM
- LLM GPU server costs: what to budget for
- VRAM: what it is and how much is needed for AI
- Source: Intel announcement about Crescent Island
- Source: Tom's Hardware on Crescent Island
- Source: vLLM documentation, Intel XPU backend
Is Crescent Island available for purchase?+
Can Crescent Island replace a consumer GPU for local AI?+
Does 160 GB of LPDDR5X mean faster models?+
Does Crescent Island go up to 480 GB?+
Will Crescent Island be compatible with vLLM or llama.cpp?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.