Best LLM for retail and e-commerce customer service

Setting up a local-commerce AI for customer service means running an LLM on your own machines to answer customers of an online store without sending the catalog, orders, or mailing addresses to a third-party service. The right model for this local-commerce AI isn't the largest or highest-ranked on a general-purpose benchmark: it's the one that correctly reads a product page, cites the right order status, applies the return policy to the letter, and hands off to a human when it doesn't know. This article details the four tasks to cover, proposes a tiered hardware-based selection from the quelllm.fr catalog, describes a data-grounding evaluation protocol, and then reviews licensing and deployment before an FAQ.

The four tasks of an e-commerce customer service LLM

The scope is intentionally narrow. Multilingual support is covered in the dedicated guide to multilingual local customer support and internal IT support is not covered here. Four functions remain that define the requirements specification:

The capability that therefore differentiates models isancrage : answer from the provided data and gracefully refuse when information is missing. The recommended architecture (document retrieval, tool calls, refusal rules) is described in the local agent architecture guide.

Selection for 128 GB Mac and dual-GPU workstation (Q4 between 68 and 75 GB)

This tier corresponds to a Mac Studio with 128 GB of unified memory or a workstation with a 96 GB professional card. The models below use mixture-of-experts (MoE) architectures: few active parameters per token, so throughput is compatible with real-time chat despite the weight size. All throughput values are estimates derived from the number of active parameters and memory bandwidth, to be confirmed on your machine.

Qwen 3.5 122B-A10B (Alibaba, Apache 2.0) - Settings : 122B total, approximately 10B active per token - Q4 VRAM : ~73 GB; Q5 ~88 GB (estimated); Q8 ~130 GB (estimated); FP16 ~245 GB (estimated) - Context : 262,000 tokens, more than enough for a complete return policy, a conversation history, and a dozen product sheets - Throughput : estimated 30 to 45 tokens/s on a Mac Studio M4 Max 128 GB in Q4; estimated 80 tokens/s or more on a 96 GB card under vLLM - Why for e-commerce : the Qwen family follows formatting instructions and tool calls reliably, which matters for order statuses. Weights published by Alibaba on Hugging Face.

Mistral Small 4 (Mistral, Apache 2.0) - Settings : 119B - Q4 VRAM : ~72 GB; Q8 ~127 GB (estimated); FP16 ~240 GB (estimated) - Context : 256,000 tokens - Throughput : to be confirmed depending on the architecture's number of active parameters; expect an order of magnitude comparable to the previous one if the model is MoE - Why for e-commerce : French writing quality for client-facing content, Apache 2.0 license with no commercial clause. Weights published by Mistral on Hugging Face.

Qwen3.8 Flash Next 125B-A6B (Qwen, open-weights license, terms to verify before commercial use) - Settings : 125B total, approximately 6B active— Q4 VRAM : ~72 GB; Q8 ~133 GB (estimated) - Context : 256,000 tokens - Throughput : the fastest in its tier, estimated at 50 to 70 tokens/s on Mac Studio 128 GB - Why for e-commerce : short latency for high-traffic chat. Tradeoff: fewer active parameters, so test carefully with ambiguous questions.

Mistral Medium 3.5 128B (Mistral, Modified MIT) - Q4 VRAM : ~74 GB; Q8 ~136 GB (estimated) - Context : 256,000 tokens - Why for e-commerce : an alternative to Mistral Small 4 when you want an extra level of reasoning for dispute cases, at the cost of lower throughput (to be confirmed).

To choose the machine, the page dedicated to 128 GB Macs lists the models that actually fit with context.

Server tier: 140 GB and up

If you have a server with multiple cards or a 192 GB machine and above, three models offer extra quality headroom for complex cases (multi-step claims, partially shipped orders).

Models above 400 GB in the catalog (DeepSeek V4 Pro, Kimi K3, GLM 5.2) are beyond the reach of a store deployment and provide no measurable benefit on such narrowly defined tasks.

Evaluating grounding and refusal: the protocol that matters

Public benchmarks (MMLU, HumanEval, AIME) measure general knowledge, coding, or math. None measures “did the model invent a return time.” The Open LLM Leaderboard is used to rule out a weak model, not to choose one for this use case. Build your own test set; half a day is enough:

For the document retrieval layer, the Firecrawl and local RAG guide shows how to build the catalog index; the GraphRAG guide handles catalogs with many relationships between products.

Licenses and customer data

A customer service handles personal data: name, address, purchase history. Local deployment eliminates the transfer to a subcontractor, but two points still need to be addressed.

Deployment: engine, interface, target latency

Le deployment guide for an intranet team chatbot covers setup step by step.

FAQ

Q: Which model should I choose to get started with a 128 GB Mac Studio?

Qwen 3.5 122B-A10B in Q4 (~73 GB) is the most balanced starting point: Apache 2.0 license, good tool calling, estimated throughput of 30 to 45 tokens/s. If latency takes priority over nuance, Qwen3.8 Flash Next 125B-A6B is faster but should be tested more thoroughly on ambiguous questions. Run the 50-question protocol on both before deciding.

Q: Would a model smaller than 100B be sufficient for this use case?

Often yes for catalog searches and order statuses, where the model mainly reformulates provided data. The gains from models in this tier show up in ambiguous cases and refusals. The catalog filter by VRAM to compare with smaller models; the evaluation protocol remains the same.

Q: How can I prevent the model from making up a delivery time?

Three layers: the status always comes from a tool call, never from the model's memory; the system prompt requires “no date without a tool result”; an application rule blocks any response containing a date if no call was made. Then measure the hallucination rate on your 20 unanswered questions: it must stay below 2%.

Q: Do you need a 1,000,000-token context to load the entire catalog?

Rarely. Loading an entire catalog for every message multiplies prefill latency and cache memory costs. A search that returns 5 to 10 relevant records in a context of 8,000 to 16,000 tokens produces better results and a faster response. The very long context of DeepSeek V4 Flash is better suited to occasional analyses.

Q: How does this differ from the guide on multilingual support?

This comparison is limited to transactional tasks in French: catalog, orders, returns, escalation. The multilingual customer support guide handles language detection, translation, and per-language test sets. The two work together: choose the model here, then add the multilingual layer if your store serves customers outside the French-speaking world.

Q: How can I measure the actual throughput on my machine?

Load the model in Q4 with llama.cpp or vLLM, send the same typical conversation with 8 000 context tokens ten times, and record tokens/s during generation and time to the first token. The figures in this article are estimates; your results depend on memory bandwidth, quantization, and the number of simultaneous sessions. The local benchmark page describes a reproducible method.

Conclusion

A local-commerce AI for customer service is judged on grounding and refusal, not on an MMLU score. On a 128 GB Mac or 96 GB workstation, Qwen 3.5 122B-A10B and Mistral Small 4 cover catalogs, orders, returns, and escalation under an Apache 2.0 license. Beyond 140 GB, Qwen 3 235B-A22B and MiniMax-M2.7 provide more headroom for complex cases. Check compatibility with your hardware in the configurator, then validate the selected model with your own 50 questions.

The hardware for running an LLM locally

To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:

Amazon GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.