Best LLM for retail and e-commerce customer service
Setting up a local-commerce AI for customer service means running an LLM on your own machines to answer customers of an online store without sending the catalog, orders, or mailing addresses to a third-party service. The right model for this local-commerce AI isn't the largest or highest-ranked on a general-purpose benchmark: it's the one that correctly reads a product page, cites the right order status, applies the return policy to the letter, and hands off to a human when it doesn't know. This article details the four tasks to cover, proposes a tiered hardware-based selection from the quelllm.fr catalog, describes a data-grounding evaluation protocol, and then reviews licensing and deployment before an FAQ.
The four tasks of an e-commerce customer service LLM
The scope is intentionally narrow. Multilingual support is covered in the dedicated guide to multilingual local customer support and internal IT support is not covered here. Four functions remain that define the requirements specification:
- Search the product catalog : the model receives a catalog excerpt (product sheets, inventory, variants, prices) through a retrieval system and must answer solely from that excerpt. An invented clothing size or guessed accessory compatibility costs you a return.
- Order status : the model calls a tool (function) that queries your order database and reformulates the result. It must never produce a tracking number or delivery date that did not come from that call.
- Returns and Refund Policy : the policy text is provided in the context. The model must cite the exact deadlines and conditions, including exclusions, and must not round “14 days” to “about two weeks.”
- Human escalation : as soon as a request falls outside the scope (dispute, goodwill gesture, missing data), the model must produce a transfer response and a structured summary for the human agent.
The capability that therefore differentiates models isancrage : answer from the provided data and gracefully refuse when information is missing. The recommended architecture (document retrieval, tool calls, refusal rules) is described in the local agent architecture guide.
Selection for 128 GB Mac and dual-GPU workstation (Q4 between 68 and 75 GB)
This tier corresponds to a Mac Studio with 128 GB of unified memory or a workstation with a 96 GB professional card. The models below use mixture-of-experts (MoE) architectures: few active parameters per token, so throughput is compatible with real-time chat despite the weight size. All throughput values are estimates derived from the number of active parameters and memory bandwidth, to be confirmed on your machine.
Qwen 3.5 122B-A10B (Alibaba, Apache 2.0) - Settings : 122B total, approximately 10B active per token - Q4 VRAM : ~73 GB; Q5 ~88 GB (estimated); Q8 ~130 GB (estimated); FP16 ~245 GB (estimated) - Context : 262,000 tokens, more than enough for a complete return policy, a conversation history, and a dozen product sheets - Throughput : estimated 30 to 45 tokens/s on a Mac Studio M4 Max 128 GB in Q4; estimated 80 tokens/s or more on a 96 GB card under vLLM - Why for e-commerce : the Qwen family follows formatting instructions and tool calls reliably, which matters for order statuses. Weights published by Alibaba on Hugging Face.
Mistral Small 4 (Mistral, Apache 2.0) - Settings : 119B - Q4 VRAM : ~72 GB; Q8 ~127 GB (estimated); FP16 ~240 GB (estimated) - Context : 256,000 tokens - Throughput : to be confirmed depending on the architecture's number of active parameters; expect an order of magnitude comparable to the previous one if the model is MoE - Why for e-commerce : French writing quality for client-facing content, Apache 2.0 license with no commercial clause. Weights published by Mistral on Hugging Face.
Qwen3.8 Flash Next 125B-A6B (Qwen, open-weights license, terms to verify before commercial use) - Settings : 125B total, approximately 6B active— Q4 VRAM : ~72 GB; Q8 ~133 GB (estimated) - Context : 256,000 tokens - Throughput : the fastest in its tier, estimated at 50 to 70 tokens/s on Mac Studio 128 GB - Why for e-commerce : short latency for high-traffic chat. Tradeoff: fewer active parameters, so test carefully with ambiguous questions.
Mistral Medium 3.5 128B (Mistral, Modified MIT) - Q4 VRAM : ~74 GB; Q8 ~136 GB (estimated) - Context : 256,000 tokens - Why for e-commerce : an alternative to Mistral Small 4 when you want an extra level of reasoning for dispute cases, at the cost of lower throughput (to be confirmed).
To choose the machine, the page dedicated to 128 GB Macs lists the models that actually fit with context.
Server tier: 140 GB and up
If you have a server with multiple cards or a 192 GB machine and above, three models offer extra quality headroom for complex cases (multi-step claims, partially shipped orders).
- Qwen 3 235B-A22B (Alibaba, Apache 2.0): ~142 GB in Q4, ~250 GB in Q8 (estimated), 131,072-token context, approximately 22B active. A reliable choice for sequential tool calls.
- MiniMax-M2.7 (MiniMax, Apache 2.0): ~138 GB in Q4, 205,000 tokens of context. Agent-oriented and useful when customer service chains catalog searches with ticket updates.
- DeepSeek V4 Flash 284B (DeepSeek, MIT): ~170 GB in Q4, 1,000,000-token context. The very long context makes it possible to load an entire catalog containing several thousand references without prior searching, but prefill latency increases proportionally. Weights published by DeepSeek on Hugging Face. The choice between Flash and Pro is detailed in the DeepSeek V4 Pro vs. Flash comparison.
- GLM 5.3 Flash 320B-A18B (Zhipu, MIT): ~186 GB in Q4, 128,000 tokens, 18B active. Weights published by Zhipu on Hugging Face.
Models above 400 GB in the catalog (DeepSeek V4 Pro, Kimi K3, GLM 5.2) are beyond the reach of a store deployment and provide no measurable benefit on such narrowly defined tasks.
Evaluating grounding and refusal: the protocol that matters
Public benchmarks (MMLU, HumanEval, AIME) measure general knowledge, coding, or math. None measures “did the model invent a return time.” The Open LLM Leaderboard is used to rule out a weak model, not to choose one for this use case. Build your own test set; half a day is enough:
- 50 questions in three categories : 20 where the answer is in the provided data, 20 where the answer is absent, and 10 ambiguous (an existing product in two variants, an order with two packages).
- Identical context for all models : same product sheets, same return policy, same simulated output from the ordering tool.
- Three metrics : accuracy on questions present; correct refusal rate on questions absent (“I don’t have that information; I’ll transfer you to an advisor”); hallucination rate, meaning a confident answer unsupported by the context.
- Acceptance threshold : an invention rate above 2% on unanswerable questions disqualifies the model, regardless of its quality elsewhere. A customer who receives a false delivery date creates a ticket, a complaint, and sometimes a negative review.
- System prompt test : add the explicit instruction “if the information is not in the context, say you don't know and offer a transfer.” Compare before and after. The models listed above generally respond well to this instruction, but the difference between them is most apparent on ambiguous questions.
For the document retrieval layer, the Firecrawl and local RAG guide shows how to build the catalog index; the GraphRAG guide handles catalogs with many relationships between products.
Licenses and customer data
A customer service handles personal data: name, address, purchase history. Local deployment eliminates the transfer to a subcontractor, but two points still need to be addressed.
- Model license : Apache 2.0 (Qwen 3.5, Mistral Small 4, Qwen 3 235B, MiniMax-M2.7) and MIT (DeepSeek V4 Flash, GLM 5.3 Flash) allow commercial use without volume restrictions. The Modified MIT license for Mistral Medium 3.5 and the “other” license for Qwen3.8 Flash Next require reviewing the text before production deployment. The NVIDIA Open Model License of Nemotron 3 Super 120B has its own terms.
- Logging : retain conversations for evaluation, but with a defined retention period and pseudonymized order identifiers. The local LLM guide for businesses and GDPR details the processing register.
Deployment: engine, interface, target latency
- Inference engine : llama.cpp for a Mac Studio or a single-user workstation in GGUF; vLLM as soon as multiple clients chat in parallel, because batching multiplies overall throughput on the same card.
- Interface and tools : Open WebUI provides a test frontend and tool (function) management for connecting your command API. In production, integration is generally handled in your existing chat widget through the OpenAI-compatible API exposed by the engine.
- Context to reserve : expect 8,000 to 16,000 tokens per conversation (return policy, 5 to 10 product sheets, history). On a 128 GB Mac with a 73 GB Q4 model, there is still room for several simultaneous sessions.
- Target latency : first token in under 2 seconds, complete response in under 10 seconds. MoE models in the 72 GB tier meet this target; dense models of comparable size do not.
Le deployment guide for an intranet team chatbot covers setup step by step.
FAQ
Q: Which model should I choose to get started with a 128 GB Mac Studio?
Qwen 3.5 122B-A10B in Q4 (~73 GB) is the most balanced starting point: Apache 2.0 license, good tool calling, estimated throughput of 30 to 45 tokens/s. If latency takes priority over nuance, Qwen3.8 Flash Next 125B-A6B is faster but should be tested more thoroughly on ambiguous questions. Run the 50-question protocol on both before deciding.
Q: Would a model smaller than 100B be sufficient for this use case?
Often yes for catalog searches and order statuses, where the model mainly reformulates provided data. The gains from models in this tier show up in ambiguous cases and refusals. The catalog filter by VRAM to compare with smaller models; the evaluation protocol remains the same.
Q: How can I prevent the model from making up a delivery time?
Three layers: the status always comes from a tool call, never from the model's memory; the system prompt requires “no date without a tool result”; an application rule blocks any response containing a date if no call was made. Then measure the hallucination rate on your 20 unanswered questions: it must stay below 2%.
Q: Do you need a 1,000,000-token context to load the entire catalog?
Rarely. Loading an entire catalog for every message multiplies prefill latency and cache memory costs. A search that returns 5 to 10 relevant records in a context of 8,000 to 16,000 tokens produces better results and a faster response. The very long context of DeepSeek V4 Flash is better suited to occasional analyses.
Q: How does this differ from the guide on multilingual support?
This comparison is limited to transactional tasks in French: catalog, orders, returns, escalation. The multilingual customer support guide handles language detection, translation, and per-language test sets. The two work together: choose the model here, then add the multilingual layer if your store serves customers outside the French-speaking world.
Q: How can I measure the actual throughput on my machine?
Load the model in Q4 with llama.cpp or vLLM, send the same typical conversation with 8 000 context tokens ten times, and record tokens/s during generation and time to the first token. The figures in this article are estimates; your results depend on memory bandwidth, quantization, and the number of simultaneous sessions. The local benchmark page describes a reproducible method.
Conclusion
A local-commerce AI for customer service is judged on grounding and refusal, not on an MMLU score. On a 128 GB Mac or 96 GB workstation, Qwen 3.5 122B-A10B and Mistral Small 4 cover catalogs, orders, returns, and escalation under an Apache 2.0 license. Beyond 140 GB, Qwen 3 235B-A22B and MiniMax-M2.7 provide more headroom for complex cases. Check compatibility with your hardware in the configurator, then validate the selected model with your own 50 questions.
The hardware for running an LLM locally
To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:
Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.