Deploy an internal AI chatbot for SMBs (from 0 to 50 collaborateurs)
You run an SMB with fewer than 50 employees and want to offer your teams an AI assistant without sending their data to OpenAI. This guide provides a complete, costed stack for deploying an internal SMB AI chatbot: concrete hardware, open-source software, Microsoft SSO, a 4-week plan, change management, and ROI calculations. No hand-waving, no pitch—the shopping list and schedule.
#Why an internal chatbot instead of ChatGPT Business
ChatGPT Team costs €25 per user per month. For 30 employees, that comes to €9,000 per year, indefinitely. For the same budget, you can buy a workstation that runs for 4 to 5 years, and your conversations never leave the office. That is the central trade-off.
But the real issue isn't the price. It's what your teams dare to paste into the chat window. A salesperson asking an external LLM to rewrite a proposal includes a customer's name, an amount, and sometimes a margin. An HR employee summarizing an interview reveals someone's career path. Multiply that by 30 people over 12 months—the leak is guaranteed, even with an acceptable-use policy.
- Privacy
- Conversations stay on your LAN. No Cloud Act transit, no exposure to a US subcontractor, and no policy settings to review every quarter.
- Fixed cost
- The hardware is CAPEX amortized over 4-5 years. The SaaS subscription is OPEX that increases with headcount.
- Mastery
- You choose the model, adjust the system prompt, and connect your internal documents through RAG. No one changes your tool behind your back.
- Compliance
- It's easier to justify compliance with the AI Act and GDPR using an on-premises system than relying on a provider from outside the EU.
#1. Hardware budget: the €5,000–10,000 workstation
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
To serve 30 to 50 simultaneous users on a 20B to 35B model at Q4_K_M quality, you don't need a DGX server. A well-chosen workstation can handle it. Here are three configurations based on the use case.
#Profile A — 10 to 20 users (€5,000)
- GPU
- 1× RTX 4080 16 GB used (variable price) or RTX 5070 Ti 16 GB new (≈ €1,400 in late September 2026). Runs a 20-24B model in Q4_K_M with a 16k-token context.
- CPU
- AMD Ryzen 9 7900X or Intel Core i7-14700K. Not very critical, but 12+ cores help with embedding preparation.
- System RAM
- 64 GB DDR5. Enough for the system, Qdrant, and KV cache spilling over from the GPU.
- Storage
- 2 TB NVMe Gen4. Models + vector database + local backups.
- Served model
- Mistral Small 24B (general-purpose, good in French) or gpt-oss 20B in Q4_K_M (~14 GB VRAM). ~30 tok/s latency.
#Profile B — 20 to 35 users (€7,500)
- GPU
- 1× RTX 4090 24 GB — no longer sold new; look for it used (variable price). The ideal compromise for serving a 27–30B model in Q4_K_M.
- CPU
- Ryzen 9 7950X or Threadripper 7960X. Useful cores for serving several simultaneous conversations.
- System RAM
- 128 GB DDR5 ECC if possible. Indexed RAG grows quickly.
- Storage
- 2 TB NVMe Gen4 + 4 TB SATA for archiving.
- Served model
- Qwen 3.8 27B (the 2026 “Copilot-like” model, 262k ctx) or Granite 4.2 30B in Q4_K_M (~18 GB VRAM). Latency ~20-25 tok/s.
#Profile C — 35 to 50 users (€10,000)
- Option NVIDIA
- 2× RTX 4090 24 GB in tensor parallel. Serves a 27-35B model in full-quality Q8 (~30-40 GB of distributed VRAM), or runs two models in parallel with acceptable latency.
- Option Apple
- Mac Studio M4 Max or Ultra with 64 to 128 GB of unified memory. Excellent noise-to-performance ratio, ideal for an open office.
- CPU
- Threadripper 7970X / 7980X on the NVIDIA side. Critical for handling 50 simultaneous WebSocket connections.
- System RAM
- 128 to 256 GB. Headroom for the OS cache, Qdrant in RAM, and a possible auxiliary embedding model.
- Served model
- Qwen 3.8 27B in Q8 or Qwen 3.6 35B-A3B (fast MoE) in Q5_K_M. Latency: ~15–20 tok/s.
#2. Recommended technical stack
Four open-source building blocks are enough. No vendor has a say, no license needs renewal, and everything runs in Docker on the workstation.
- Ollama
- The daemon that loads and serves the model. Listens by default on http://localhost:11434. Automatically detects the NVIDIA GPU and handles quantization. MIT License.
- Open WebUI
- The ChatGPT-style web interface. Native multi-user support, RBAC, OIDC integration for SSO, per-user history. BSD-3-Clause license.
- Qdrant
- Vector database for RAG. Indexes your internal documents (procedures, template contracts, product sheets). Faster and simpler than pgvector for this volume. Apache 2 license.
- Traefik or Nginx
- Reverse proxy to expose Open WebUI over HTTPS on the LAN with an internal certificate, terminate TLS, and route to the backends.
This skeleton is intentionally short. You will add Traefik on top for HTTPS, plus a bind-mounted volume for nightly backups to the NAS. Nothing more.
#3. SSO with Microsoft Entra ID
Most SMBs have Microsoft 365. Connecting the chatbot's authentication to Entra ID (formerly Azure AD) eliminates orphaned accounts and ensures that an employee who leaves loses access the same day. Open WebUI natively supports OIDC with Microsoft.
- 01Create an App registrationIn the Entra ID portal, App registrations → New registration. Name: “Internal AI Chatbot.” Redirect URI: https://chatbot.votreboite.local/oauth/microsoft/callback (Web).
- 02Retrieve the credentialsNote the Application (client) ID and the Directory (tenant) ID. Under Certificates & secrets, generate a Client secret with a 24-month expiration and note it immediately (it will never be displayed again).
- 03Define permissionsAPI permissions → Microsoft Graph → Delegated: openid, profile, email, User.Read. Application permissions aren't needed; the flow is delegated.
- 04Restrict accessEnterprise applications → your app → Properties → Assignment required = Yes. Then, under Users and groups, add the “Tous les collaborateurs” group or a pilot group. Without this step, any account in the tenant can sign in.
- 05Inject into Open WebUICopy CLIENT_ID, CLIENT_SECRET, and TENANT_ID into the docker-compose .env file. Restart. The “Sign in with Microsoft” button appears on the login page.
#4. 4-week deployment plan
This timeline assumes an internal technical lead (IT department, IT manager, or service provider) who dedicates ~50% of their time to the project. An SMB does not need more.
#Week 1 — Hardware and installation
- J1-J2
- Hardware order (workstation, 1500 VA UPS, managed switch if not already in place).
- J3-J4
- Receiving, assembly if needed, installation of Ubuntu Server LTS 24.04, drivers NVIDIA, Docker + NVIDIA Container Toolkit.
- J5
- Pull Docker images, first ollama pull mistral-small, raw chat test in the CLI. You should see tokens streaming out.
#Week 2—Stack and integrations
- J6-J7
- docker-compose up, configuration Open WebUI (admin, paramètres généraux, branding interne).
- J8-J9
- Entra ID SSO configuration, testing with 3 pilot accounts, and validation of the login and logout flow.
- J10
- Traefik reverse proxy with an internal certificate through your PKI, or Let’s Encrypt DNS-01 if the domain is public. Test from 2–3 LAN machines.
#Week 3 — RAG and pilot
- J11-J12
- Select 50 to 200 relevant internal documents (procedures, FAQs, sales sheets). Index them in Qdrant via Open WebUI or a Python script.
- J13-J14
- SMB-colored system prompt: tone, permitted topics, topics to refuse (sensitive HR, salaries, etc.). Test with the 3 pilot users.
- J15
- Open the pilot to a representative group of 5–8 employees (1 sales, 1 HR, 1 accounting, 1 operations…). Collect structured feedback.
#Week 4 — Training and cutover
- J16-J17
- Adjustments based on pilot feedback (system prompt, RAG documents, temperature settings).
- J18
- 1h30 group training session: demo, role-specific use cases, usage rules, and what NOT to put in the chat (health data, identifiers, etc.).
- J19
- Gradual rollout to all employees through the Entra ID group. Internal announcement.
- J20
- Set up monitoring (Glances or Netdata for the workstation, plus an aggregated conversation log to identify requested topics).
#5. Training and change management
An internal chatbot without training means 70% of the investment is wasted. Employees don't know what to ask it, mentally compare it with ChatGPT for the general public, and conclude that “it's not as good” after two attempts.
- Format
- A group session of 1h30 in batches of 10-15 people. No slides for 1h. 20 min of demos + 1h of hands-on practice with their real topics.
- Use cases by profession
- Prepare 3 concrete prompts for each function (HR, sales, accounting, ops, management). People copy and adapt them—that’s exactly what we want.
- Usage rules
- Clearly indicate what may be pasted (internal texts, drafts, already-public data) and what must not be pasted (health information, convictions, named banking data, technical identifiers).
- Reference
- Appoint 1 AI lead per department. This person shares best practices locally and reports gaps in the RAG's knowledge.
- 30-day follow-up
- Measure usage. If some teams never touch the tool, that is a signal—either they are missing an obvious use case, or the training did not stick.
#6. GDPR and AI Act compliance
The fact that the system is on-premises greatly simplifies compliance, but does not eliminate it. GDPR obligations apply to the processing, not its location.
- Record of processing activities
- Add a record for the internal chatbot. Purpose: drafting assistance and document search. Legal basis: the employer's legitimate interest. Data processed: conversation content, SSO identifier.
- Retention period
- Define a clear policy: 90 days of conversation history per user, with automatic purging afterward. Document it in the charter.
- Employee information
- Update the IT policy and the CSE/CSE information notice. Explicitly state that conversations are stored and accessible to the technical administrator in the event of an incident.
- AI Act
- An internal productivity-assistance chatbot falls into the « limited risk » category (Article 50). Main obligation: inform the user that they are interacting with AI — the Open WebUI interface does this by default.
- Technical security
- Encrypted backups of the Open WebUI database, audited admin access, MFA required for Entra ID accounts, monthly Docker image updates.
#7. ROI: the honest calculation
Take a small business with 30 employees and a Profile A deployment costing €5,000. Over a 3-year period.
- Alternative SaaS cost
- 30 × 25 € × 12 months × 3 years = 27,000 € (ChatGPT Team).
- Internal cost
- Workstation 5 000 € + electricity ~200 €/year + 5 person-days of installation in the 1st year (~3 500 €) + 2 person-days/year of maintenance (~1 400 €/year × 2 years) = ~11 500 € over 3 years.
- Direct net savings
- ~15 500 € over 3 years, or about 5 200 € per year starting in year 2.
- Indirect gains
- Reduced data-leakage risk (you can make the case that avoiding a single leak more than covers the project), easier GDPR compliance, and independence from SaaS price increases.
- Productivity gain
- Conservative estimate: 15 minutes saved per user per working day. At €35/hour fully loaded and 220 days, that comes to ~€38,500 per year for 30 people. This figure depends heavily on actual adoption—be cautious in your business cases.
#Go further
This guide lays the foundation. Three natural next steps: add serious RAG over your internal documents so the chatbot knows your procedures, harden the multi-user deployment on your intranet, and formalize GDPR compliance with a proper file.
Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.