Best LLM for technical support at an IT services company (IT)
Setting up a local IT support AI addresses a concrete constraint for IT services companies: tickets, logs, and runbooks contain customer data covered by confidentiality clauses. With a local IT support AI, these elements stay on a machine you control, and each response can be traced to a source.
This page covers only internal IT support: technical ticket qualification, reading redacted logs, searching runbooks, traceability, and escalation. Commerce, refunds, and customer service translation are covered by the guide multilingual customer support with a local LLM.
The plan: selection criteria, model selection, memory by quantization, throughput and benchmarks, licenses, then the deployment pipeline.
What an IT support provider requires from a model
Internal support requires different qualities than a general-purpose assistant. Four criteria matter:
- Context window : a log excerpt, a runbook, and a ticket history quickly exceed 30,000 tokens. Models limited to 4,096 tokens (Snowflake Arctic Instruct) or 8,192 tokens (Grok-1 (base)) are excluded for this use case.
- Structured output : the model must produce stable JSON (category, severity, target team) that the ticketing tool can consume.
- Source fidelity : the answer must cite the runbook or log line used, or traceability is lost.
- License : the model often runs as part of a billed service, and therefore in commercial use.
The selection: seven models between 68 and 85 GB in Q4
The models below are the lightest in this page's reference catalog. All require a high-memory server or workstation, not a technician's desktop.
- Mistral Small 4 : 119B, Apache 2.0, ~72 GB in Q4, 256,000-token context. Default first choice thanks to its permissive license and large context. Weights on the Hugging Face page for Mistral.
- Qwen 3.5 122B-A10B : 122B, Apache 2.0, ~73 GB in Q4, 262,000-token context. The A10B suffix indicates approximately 10 billion active parameters per token, which favors throughput. See Alibaba on Hugging Face.
- Nemotron 3 Super 120B : 120B, NVIDIA Open Model License, ~72 GB in Q4, 128,000-token context. Relevant on infrastructure already equipped with NVIDIA (NVIDIA on Hugging Face).
- Laguna S 2.1 : 118B, OpenMDW 1.1, ~68 GB in Q4, 262,144 context. The lightest in the selection, geared toward code, useful for tickets involving scripts and pipelines.
- Qwen3.8 Flash Next 125B-A6B : 125B, “Other (open weights)” license, ~72 GB in Q4, 256,000 context. Read the license before any use with a client.
- Mistral Medium 3.5 128B : 128B, Modified MIT, ~74 GB in Q4, 256,000 context.
- Mixtral 8x22B Instruct : 141B, Apache 2.0, ~82 GB in Q4, 64,000-token context. Shorter context, sufficient for a ticket and a runbook, but only just enough for long log excerpts.
Two models are limited by their 32,768-token context: dots.llm1 Instruct (142B, MIT, ~85 GB) and DBRX Instruct (132B, Databricks Open Model License, ~76 GB). They are suitable for ticket triage, less so for log analysis.
For a team with more memory, Step 3.5 Flash (196B, Apache 2.0, ~118 GB) and MiniMax-M2.7 (229B, Apache 2.0, ~138 GB) make up the next tier.
Memory by quantization and hardware
Only the Q4 values come from the catalog. The other levels are rough estimates based on proportional scaling and must be confirmed on each product page.
For the 118B–128B class:
- Q4 : 68 to 74 GB (catalog)
- Q5 : ~82 to 90 GB (estimated)
- Q8 : ~120 to 135 GB (estimated)
- FP16 : ~236 to 256 GB (estimated)
For Mixtral 8x22B Instruct (141B):
- Q4 : ~82 GB (catalog)
- Q5 : ~97 GB (estimated)
- Q8 : ~150 GB (estimated)
- FP16 : ~282 GB (estimated)
These figures do not include context: loading 100,000 log tokens adds several GB of cache. Plan for a margin of at least 15 to 20% (estimated).
On the hardware side, a 72 GB Q4 fits on a 96 GB unified-memory machine with little headroom, and more comfortably on 128 GB. The pages Mac 96 GB et Mac 128 GB detail these configurations. On a PC, you need to combine multiple graphics cards: see the guide choose a GPU for a local LLM.
Throughput and benchmarks: what remains to be confirmed
No measured throughput is published here for these seven models: tokens/sec must be confirmed on your hardware. Two qualitative reference points:
- MoE architecture : at the same size, a model with a low number of active parameters (A6B, A10B) reads less data per token and generates faster than a dense model.
- Competition : a support team sends several simultaneous requests. A server such as vLLM processes requests in batches and increases total throughput, whereas llama.cpp remains the simplest option for a single GGUF pipeline.
For MMLU or HumanEval scores, see theOpen LLM Leaderboard and the profiles from the catalog. These scores must be confirmed model by model and do a poor job of measuring support work.
An internal test set is more reliable: 50 closed and anonymized tickets, with the actual category, severity, and resolution. Measure the correct categorization rate, the rate of exact runbook citations, and the number of erroneous escalations. The page local benchmark describes the throughput measurement method.
Licenses: verify before installing for a client
- Apache 2.0 (Mistral Small 4, Qwen 3.5 122B-A10B, Mixtral 8x22B Instruct): commercial use permitted, provided the license notices are retained.
- MIT (dots.llm1 Instruct): commercial use permitted, minimal restrictions.
- Modified MIT (Mistral Medium 3.5 128B): read the added clauses and confirm the conditions.
- NVIDIA Open Model License, OpenMDW 1.1, Databricks Open Model License : licenses specific to the vendors, which should be reviewed before deployment as a service.
- Other (open weights) (Qwen3.8 Flash Next 125B-A6B): conditions to be confirmed on a case-by-case basis.
The guide Local LLMs in business and GDPR completes this point for personal data contained in tickets.
Deployment pipeline: tickets, redacted logs, runbooks, escalation
A local support chain has five steps:
- Deterministic redaction : a script replaces IP addresses, hostnames, tokens, and email addresses with neutral identifiers before anything is sent to the model. The mapping table remains outside the model.
- Searching runbooks : procedures are split and indexed, then the model receives only the relevant passages with their references.
- Qualification : the model returns JSON with a category, severity, suspected cause, and cited sources.
- Escalation rule : any response without a source, any high-severity incident, and any action that modifies production are escalated to a technician.
- Journal : each exchange is recorded with the model version, quantization, prompt, and passages used.
For the team interface, Open WebUI connects to a local server. The guide deploy a team chatbot on the intranet details the installation, and architecture of a local agent covers tool calling.
FAQ
Q: Do these models run on a technician’s workstation?
No. The lightest model in the selection, Laguna S 2.1, requires about 68 GB in Q4, excluding context. The realistic setup is a shared server or a workstation with 96-128 GB of memory, accessed by the team over the internal network. For more modest machines, the site's configurator suggests models suited to the available memory.
Q: Should we fine-tune the model on our tickets?
Not initially. Searching the runbooks, with citations, already provides the ESN’s vocabulary and procedures without retraining. Fine-tuning becomes useful if the output format remains unstable or categorization plateaus on your test set. It requires anonymized data and verification of the model license.
Q: Can the model redact the logs itself?
This is not recommended. A model may forget an IP address or token in a long excerpt. Redaction should be handled by a deterministic, tested, version-controlled script placed before the model. The model then works with neutral identifiers, while the mapping to the real values remains in a separate system.
Q: Can the model close a ticket on its own?
It’s best to avoid this at startup. The model suggests a classification and a resolution path, then a technician validates it. After a few weeks of measurement, some low-risk categories can be automated, such as documentation requests. Production incidents and privileged actions remain subject to human approval.
Q: What minimum context should I target for log analysis?
Aim for at least 64,000 tokens, which Mixtral 8x22B Instruct provides, and preferably 128,000 or more to correlate multiple files. Models with 32,768 tokens require heavily splitting excerpts. A large context uses more memory and slows generation: filter logs before sending them.
Q: Does this selection cover customer service?
No. It targets a managed service provider's internal IT support: incidents, logs, runbooks, and escalation. Sales conversations, refunds, and translating conversations with end customers fall under other criteria, especially multilingual quality. They are covered in the guide dedicated to multilingual customer support with a local LLM.
Conclusion
For a local IT support AI solution at an ESN, Mistral Small 4 and Qwen 3.5 122B-A10B are the safest starting point: Apache 2.0 license, approximately 256,000-token context, 72 to 73 GB in Q4. Throughput and scores still need to be confirmed on your hardware and with your own tickets. Enter your available memory in the configurator to verify compatibility, or browse the catalog to compare licenses and VRAM requirements.
The hardware for running an LLM locally
To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:
Affiliate links — QuelLLM may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.