Multilingual customer support with a local LLM: 6 languages without cloud
You manage customer support that receives tickets in French, English, Spanish, German, Italian, and Arabic. Sending every message to a cloud API exposes emails, order numbers, and sometimes personal data to a third party—and you pay by the token. This guide builds a 100% local multilingual customer-support chatbot with Qwen 3.8 27B, automatic language detection, a tone adapted to each culture, and direct Zendesk or Freshdesk integration via webhook.
#Why use a local LLM for multilingual support
Cloud APIs (GPT-4o, Claude, Gemini) charge per token and require client messages to be transferred outside the EU. For B2C support handling 5,000 tickets per day, the monthly bill quickly exceeds €1,500, and GDPR compliance becomes difficult as soon as a customer sends an IBAN or social security number in the ticket body.
A local LLM solves both problems at once. Qwen 3.8 27B (Alibaba, released August 14, 2026) natively supports more than 100 languages, handles the brief's six European languages with quality that rivals entry-level cloud models, and runs on a single RTX 4090 or a Mac M4 Max. Marginal cost per ticket: zero.
#Prerequisites
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
- 24 GB VRAM GPU minimum
- RTX 3090, 4090, 5090, or a Mac M3/M4 Max with 32 GB of unified memory. Qwen 3.8 27B in Q4_K_M uses approximately 18 GB.
- Ollama installed
- Recent version for native support of Qwen 3.8 (vision and long context).
- A Zendesk or Freshdesk account
- With admin rights to create a webhook integration and a trigger.
- An accessible HTTPS endpoint
- Either a VPS that reverse-proxies to your Ollama, or ngrok / Cloudflare Tunnel to expose the local server.
- Python 3.11+
- For the language-detection layer and webhook router. FastAPI + langdetect are enough.
#1. Install Qwen 3.8 27B with Ollama
The model is available directly in the Ollama library. Start the download and a quick test:
The pull downloads about 18 GB. On a RTX 4090, inference runs at 40–60 tokens/second in Q4—enough to generate a support response in 3 to 5 seconds. With reasoning effort set to “low,” latency drops further.
Expose Ollama on the network so your webhook can reach it:
OLLAMA_KEEP_ALIVE=30m keeps the model in VRAM for 30 minutes after the last request, avoiding a cold start for every ticket during off-peak hours.
#2. Automatic language detection
Before sending a message to Qwen3, we identify its language to choose the right system prompt. The fast-langdetect library (based on Meta's fastText) detects 176 languages in under 5 ms per request, with over 99% accuracy on messages longer than 50 characters.
#3. System prompts by language and adapted tone
Good multilingual support does not translate a French prompt into English—it adapts the register. Professional French uses formal address, business English is more direct, German expects pronounced formal politeness, Latin American Spanish allows more warmth, Italian is more expressive, and Arabic requires polite formulas at the beginning and end of the message.
#4. Integrate a Zendesk or Freshdesk webhook
Zendesk and Freshdesk trigger an HTTP POST webhook for every new ticket. We expose a FastAPI endpoint that detects the language, selects the prompt, calls Ollama, and returns the response to the support API so it can be added as an internal comment (the human agent approves it before sending).
The post_internal_note function pushes the draft into a private comment through the Zendesk API (PUT /api/v2/tickets/{id}.json) or Freshdesk (POST /api/v2/tickets/{id}/notes). A human agent reviews it, adjusts it, and clicks "Send." You save about 60% of drafting time without ever letting the bot reply on its own.
- 01Expose the endpointCloudflare Tunnel or ngrok points to http://localhost:8000. Note the public HTTPS URL.
- 02Create the webhook targetIn Zendesk Admin → Apps & Integrations → Webhooks, create a target with your URL and the POST method.
- 03Connect a triggerIn Triggers, condition "Ticket Created", action "Notify webhook" with a JSON payload containing {ticket: {id, description, requester}}.
- 04Test with a dummy ticketCreate a ticket in German. The endpoint must receive the event, detect "de", and publish a draft in German in the internal notes.
- 05Enable it gradually in productionStart with a single queue (for example, the Spanish queue). Measure quality for one week. Expand one queue at a time.
#5. Measure quality by language
Qwen 3.8 does not have the same quality in all six languages. French and English are excellent, Spanish and Italian are very good, German is acceptable, and Arabic varies by dialect (standard MSA works well, while Maghrebi dialects are less reliable). You need to measure to manage it.
Three metrics to log per language, starting on day one:
- Validation rate without editing
- Percentage of drafts the agent sends as is. This is the simplest and most meaningful indicator. Target: 40–60% at 3 months.
- Edit rate (words changed / words generated)
- Measure with a diff between the final answer and the draft. If editing exceeds 30% in a language, the prompt needs reworking.
- CSAT by language
- The post-resolution satisfaction survey, cross-referenced with the ticket language. If German drops to 3.5/5 while French stays at 4.5, you have a tone problem.
#Common pitfalls
- The bot responds in the wrong language
- It's almost always a mixed-language message (an English signature beneath a French ticket). Detect it in the first paragraph, or force the draft's language with an explicit instruction such as "Reply in {lang}" in addition to the system prompt.
- Latency > 10 seconds
- Make sure OLLAMA_KEEP_ALIVE is enabled and that the model stays in VRAM. ollama ps should show 100% GPU. If not, lower num_ctx to 4096 for short tickets.
- Poorly formatted Arabic responses (RTL)
- Qwen3 generates Arabic well, but some support interfaces don't automatically display RTL. Add dir="rtl" lang="ar" to the response block in Zendesk/Freshdesk.
- Hallucinated nonexistent promotions
- Reduce the temperature to 0.2, and specify in the prompt "do not mention any promotion or discount code unless it is present in the ticket." The LLM likes to offer gifts that do not exist.
- Webhook replaying multiple times
- Zendesk may retry on timeout. Store ticket_id in Redis with a 10-minute TTL for idempotency—otherwise you’ll generate 3 drafts for the same ticket.
#Go further
Once the baseline is stable, two extensions greatly increase the validation rate: connect a local RAG to your knowledge base (FAQ, return policy, warranty terms) to ground responses in facts, and add a LoRA fine-tuning layer on a few thousand resolved tickets to match your house style. For multi-workstation production deployment, deploying on an intranet behind Nginx handles networking and authentication.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.