Intermediate 16 minSupport

Multilingual customer support with a local LLM: 6 languages without cloud

You manage customer support that receives tickets in French, English, Spanish, German, Italian, and Arabic. Sending every message to a cloud API exposes emails, order numbers, and sometimes personal data to a third party—and you pay by the token. This guide builds a 100% local multilingual customer-support chatbot with Qwen 3.8 27B, automatic language detection, a tone adapted to each culture, and direct Zendesk or Freshdesk integration via webhook.

By Marie L.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why use a local LLM for multilingual support

Cloud APIs (GPT-4o, Claude, Gemini) charge per token and require client messages to be transferred outside the EU. For B2C support handling 5,000 tickets per day, the monthly bill quickly exceeds €1,500, and GDPR compliance becomes difficult as soon as a customer sends an IBAN or social security number in the ticket body.

A local LLM solves both problems at once. Qwen 3.8 27B (Alibaba, released August 14, 2026) natively supports more than 100 languages, handles the brief's six European languages with quality that rivals entry-level cloud models, and runs on a single RTX 4090 or a Mac M4 Max. Marginal cost per ticket: zero.

i
Why Qwen 3.8 27B specifically
This is Alibaba's flagship general-purpose model for 2026: a 262,000-token context window, vision, and an Apache 2.0 license (free commercial use). For support, this huge context lets you inject the entire ticket history without truncation, and its multilingual quality remains the best in its tier (18 GB in Q4, fits on a 24 GB card). One setting to know: set its reasoning effort to “low.” By default, it overthinks, unnecessarily increasing latency for support responses.

#Prerequisites

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
24 GB VRAM GPU minimum
RTX 3090, 4090, 5090, or a Mac M3/M4 Max with 32 GB of unified memory. Qwen 3.8 27B in Q4_K_M uses approximately 18 GB.
Ollama installed
Recent version for native support of Qwen 3.8 (vision and long context).
A Zendesk or Freshdesk account
With admin rights to create a webhook integration and a trigger.
An accessible HTTPS endpoint
Either a VPS that reverse-proxies to your Ollama, or ngrok / Cloudflare Tunnel to expose the local server.
Python 3.11+
For the language-detection layer and webhook router. FastAPI + langdetect are enough.

#1. Install Qwen 3.8 27B with Ollama

The model is available directly in the Ollama library. Start the download and a quick test:

Terminal
ollama pull qwen3.8:27b
ollama run qwen3.8:27b "Réponds en une phrase : qu'est-ce qu'un LLM open-weight ?"

The pull downloads about 18 GB. On a RTX 4090, inference runs at 40–60 tokens/second in Q4—enough to generate a support response in 3 to 5 seconds. With reasoning effort set to “low,” latency drops further.

→
Test multilingual support right away
Ask the same question in 6 languages in succession with ollama run. Qwen 3.8 doesn't make a latent transition between languages, unlike older Llama generations that sometimes drifted into English. That's exactly what you want for support.

Expose Ollama on the network so your webhook can reach it:

systemd override
sudo systemctl edit ollama
# Ajoutez :
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_KEEP_ALIVE=30m"

sudo systemctl restart ollama

OLLAMA_KEEP_ALIVE=30m keeps the model in VRAM for 30 minutes after the last request, avoiding a cold start for every ticket during off-peak hours.

#2. Automatic language detection

Before sending a message to Qwen3, we identify its language to choose the right system prompt. The fast-langdetect library (based on Meta's fastText) detects 176 languages in under 5 ms per request, with over 99% accuracy on messages longer than 50 characters.

Installation
pip install fast-langdetect fastapi uvicorn httpx
detector.py
from fast_langdetect import detect

SUPPORTED = {"fr", "en", "es", "de", "it", "ar"}

def detect_language(text: str) -> str:
    text = text.replace("\n", " ").strip()
    if len(text) < 10:
        return "en"  # message trop court, fallback raisonnable
    result = detect(text, low_memory=False)
    lang = result["lang"]
    return lang if lang in SUPPORTED else "en"
!
The mixed-message trap
A French customer who pastes an error message in English can throw off the detector. Detect based only on the first paragraph (before the first blank line), not the entire message. Otherwise, you will reply in English to a French-speaking customer.

#3. System prompts by language and adapted tone

Good multilingual support does not translate a French prompt into English—it adapts the register. Professional French uses formal address, business English is more direct, German expects pronounced formal politeness, Latin American Spanish allows more warmth, Italian is more expressive, and Arabic requires polite formulas at the beginning and end of the message.

prompts.py
SYSTEM_PROMPTS = {
    "fr": (
        "Tu es un agent de support client professionnel. "
        "Réponds en français, avec vouvoiement systématique. "
        "Sois concis, empathique, factuel. Ne promets jamais de remboursement "
        "sans validation. Si tu ne sais pas, propose une escalade humaine."
    ),
    "en": (
        "You are a professional customer support agent. "
        "Reply in English, business-friendly tone, direct and concise. "
        "Never commit to a refund without confirmation. "
        "Escalate to a human if uncertain."
    ),
    "es": (
        "Eres un agente de soporte profesional. Responde en español, "
        "trato de usted, tono cálido pero conciso. Nunca prometas reembolsos "
        "sin confirmación. Escala a un humano en caso de duda."
    ),
    "de": (
        "Du bist ein professioneller Kundensupport-Agent. Antworte auf Deutsch "
        "mit Sie-Form, höflich-formell, präzise und sachlich. Versprich nie eine "
        "Rückerstattung ohne Bestätigung. Eskaliere im Zweifel an einen Menschen."
    ),
    "it": (
        "Sei un agente di assistenza clienti professionale. Rispondi in italiano, "
        "forma di cortesia (Lei), tono cordiale e conciso. Mai promettere rimborsi "
        "senza conferma. Scala a un umano in caso di dubbio."
    ),
    "ar": (
        "أنت موظف دعم عملاء محترف. أجب باللغة العربية الفصحى، "
        "بأسلوب رسمي ومهذب يبدأ بتحية وينتهي بعبارة لطيفة. "
        "لا تعد بأي استرداد دون تأكيد. إذا لم تكن متأكدًا، اطلب تدخل موظف بشري."
    ),
}
→
Business guardrails in the prompt
The three lines “never promise a refund,” “escalate if you do not know,” and “do not offer a discount” are the only ones that prevent an LLM from causing real damage in support. Put them in every language. The tone changes; the business rules do not.

#4. Integrate a Zendesk or Freshdesk webhook

Zendesk and Freshdesk trigger an HTTP POST webhook for every new ticket. We expose a FastAPI endpoint that detects the language, selects the prompt, calls Ollama, and returns the response to the support API so it can be added as an internal comment (the human agent approves it before sending).

webhook_server.py
from fastapi import FastAPI, Request
import httpx
from detector import detect_language
from prompts import SYSTEM_PROMPTS

app = FastAPI()
OLLAMA_URL = "http://localhost:11434/api/chat"

@app.post("/webhook/zendesk")
async def handle_ticket(req: Request):
    payload = await req.json()
    ticket_id = payload["ticket"]["id"]
    message = payload["ticket"]["description"]

    lang = detect_language(message)
    system = SYSTEM_PROMPTS[lang]

    async with httpx.AsyncClient(timeout=60) as client:
        r = await client.post(OLLAMA_URL, json={
            "model": "qwen3.8:27b",
            "messages": [
                {"role": "system", "content": system},
                {"role": "user", "content": message},
            ],
            "stream": False,
            "options": {"temperature": 0.3, "num_ctx": 8192},
        })
    draft = r.json()["message"]["content"]

    # Ajoute la réponse en note interne sur le ticket
    await post_internal_note(ticket_id, lang, draft)
    return {"status": "ok", "lang": lang}

The post_internal_note function pushes the draft into a private comment through the Zendesk API (PUT /api/v2/tickets/{id}.json) or Freshdesk (POST /api/v2/tickets/{id}/notes). A human agent reviews it, adjusts it, and clicks "Send." You save about 60% of drafting time without ever letting the bot reply on its own.

  1. 01
    Expose the endpoint
    Cloudflare Tunnel or ngrok points to http://localhost:8000. Note the public HTTPS URL.
  2. 02
    Create the webhook target
    In Zendesk Admin → Apps & Integrations → Webhooks, create a target with your URL and the POST method.
  3. 03
    Connect a trigger
    In Triggers, condition "Ticket Created", action "Notify webhook" with a JSON payload containing {ticket: {id, description, requester}}.
  4. 04
    Test with a dummy ticket
    Create a ticket in German. The endpoint must receive the event, detect "de", and publish a draft in German in the internal notes.
  5. 05
    Enable it gradually in production
    Start with a single queue (for example, the Spanish queue). Measure quality for one week. Expand one queue at a time.
!
Never respond to the client directly
The LLM drafts; a human validates. That's the rule. A Qwen 3.8 that hallucinates an order number or invents a refund policy ends up as a 1-star Trustpilot review. Until you have 6 months of quality measurements per language, stay in draft mode.

#5. Measure quality by language

Qwen 3.8 does not have the same quality in all six languages. French and English are excellent, Spanish and Italian are very good, German is acceptable, and Arabic varies by dialect (standard MSA works well, while Maghrebi dialects are less reliable). You need to measure to manage it.

Three metrics to log per language, starting on day one:

Validation rate without editing
Percentage of drafts the agent sends as is. This is the simplest and most meaningful indicator. Target: 40–60% at 3 months.
Edit rate (words changed / words generated)
Measure with a diff between the final answer and the draft. If editing exceeds 30% in a language, the prompt needs reworking.
CSAT by language
The post-resolution satisfaction survey, cross-referenced with the ticket language. If German drops to 3.5/5 while French stays at 4.5, you have a tone problem.
i
Practical adjustment example
In a real deployment at a French e-commerce company expanded into Germany, the German approval rate had stalled at 18%. The identified cause: the literally translated prompt said "be empathetic." In business German, "empathisch" sounds therapeutic. Replaced with "verbindlich und lösungsorientiert" (engaging and solution-oriented) → 45% in two weeks.

#Common pitfalls

The bot responds in the wrong language
It's almost always a mixed-language message (an English signature beneath a French ticket). Detect it in the first paragraph, or force the draft's language with an explicit instruction such as "Reply in {lang}" in addition to the system prompt.
Latency > 10 seconds
Make sure OLLAMA_KEEP_ALIVE is enabled and that the model stays in VRAM. ollama ps should show 100% GPU. If not, lower num_ctx to 4096 for short tickets.
Poorly formatted Arabic responses (RTL)
Qwen3 generates Arabic well, but some support interfaces don't automatically display RTL. Add dir="rtl" lang="ar" to the response block in Zendesk/Freshdesk.
Hallucinated nonexistent promotions
Reduce the temperature to 0.2, and specify in the prompt "do not mention any promotion or discount code unless it is present in the ticket." The LLM likes to offer gifts that do not exist.
Webhook replaying multiple times
Zendesk may retry on timeout. Store ticket_id in Redis with a 10-minute TTL for idempotency—otherwise you’ll generate 3 drafts for the same ticket.

#Go further

Once the baseline is stable, two extensions greatly increase the validation rate: connect a local RAG to your knowledge base (FAQ, return policy, warranty terms) to ground responses in facts, and add a LoRA fine-tuning layer on a few thousand resolved tickets to match your house style. For multi-workstation production deployment, deploying on an intranet behind Nginx handles networking and authentication.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.