Intermediate 20 minUse Case

Deploy an internal AI chatbot for SMBs (from 0 to 50 collaborateurs)

You run an SMB with fewer than 50 employees and want to offer your teams an AI assistant without sending their data to OpenAI. This guide provides a complete, costed stack for deploying an internal SMB AI chatbot: concrete hardware, open-source software, Microsoft SSO, a 4-week plan, change management, and ROI calculations. No hand-waving, no pitch—the shopping list and schedule.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why an internal chatbot instead of ChatGPT Business

ChatGPT Team costs €25 per user per month. For 30 employees, that comes to €9,000 per year, indefinitely. For the same budget, you can buy a workstation that runs for 4 to 5 years, and your conversations never leave the office. That is the central trade-off.

But the real issue isn't the price. It's what your teams dare to paste into the chat window. A salesperson asking an external LLM to rewrite a proposal includes a customer's name, an amount, and sometimes a margin. An HR employee summarizing an interview reveals someone's career path. Multiply that by 30 people over 12 months—the leak is guaranteed, even with an acceptable-use policy.

Privacy
Conversations stay on your LAN. No Cloud Act transit, no exposure to a US subcontractor, and no policy settings to review every quarter.
Fixed cost
The hardware is CAPEX amortized over 4-5 years. The SaaS subscription is OPEX that increases with headcount.
Mastery
You choose the model, adjust the system prompt, and connect your internal documents through RAG. No one changes your tool behind your back.
Compliance
It's easier to justify compliance with the AI Act and GDPR using an on-premises system than relying on a provider from outside the EU.
i
Who this guide is for
Small and medium-sized businesses with 5 to 50 employees, with or without an internal IT department. If you have 200 people or more, some trade-offs change (switching to vLLM, high availability, structured ticketing)—this guide remains valid as a foundation.

#1. Hardware budget: the €5,000–10,000 workstation

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

To serve 30 to 50 simultaneous users on a 20B to 35B model at Q4_K_M quality, you don't need a DGX server. A well-chosen workstation can handle it. Here are three configurations based on the use case.

#Profile A — 10 to 20 users (€5,000)

GPU
1× RTX 4080 16 GB used (variable price) or RTX 5070 Ti 16 GB new (≈ €1,400 in late September 2026). Runs a 20-24B model in Q4_K_M with a 16k-token context.
CPU
AMD Ryzen 9 7900X or Intel Core i7-14700K. Not very critical, but 12+ cores help with embedding preparation.
System RAM
64 GB DDR5. Enough for the system, Qdrant, and KV cache spilling over from the GPU.
Storage
2 TB NVMe Gen4. Models + vector database + local backups.
Served model
Mistral Small 24B (general-purpose, good in French) or gpt-oss 20B in Q4_K_M (~14 GB VRAM). ~30 tok/s latency.

#Profile B — 20 to 35 users (€7,500)

GPU
1× RTX 4090 24 GB — no longer sold new; look for it used (variable price). The ideal compromise for serving a 27–30B model in Q4_K_M.
CPU
Ryzen 9 7950X or Threadripper 7960X. Useful cores for serving several simultaneous conversations.
System RAM
128 GB DDR5 ECC if possible. Indexed RAG grows quickly.
Storage
2 TB NVMe Gen4 + 4 TB SATA for archiving.
Served model
Qwen 3.8 27B (the 2026 “Copilot-like” model, 262k ctx) or Granite 4.2 30B in Q4_K_M (~18 GB VRAM). Latency ~20-25 tok/s.

#Profile C — 35 to 50 users (€10,000)

Option NVIDIA
2× RTX 4090 24 GB in tensor parallel. Serves a 27-35B model in full-quality Q8 (~30-40 GB of distributed VRAM), or runs two models in parallel with acceptable latency.
Option Apple
Mac Studio M4 Max or Ultra with 64 to 128 GB of unified memory. Excellent noise-to-performance ratio, ideal for an open office.
CPU
Threadripper 7970X / 7980X on the NVIDIA side. Critical for handling 50 simultaneous WebSocket connections.
System RAM
128 to 256 GB. Headroom for the OS cache, Qdrant in RAM, and a possible auxiliary embedding model.
Served model
Qwen 3.8 27B in Q8 or Qwen 3.6 35B-A3B (fast MoE) in Q5_K_M. Latency: ~15–20 tok/s.
→
Don't over-specify on day 1
Start with profile A even if there are 40 of you. Most users don't ask questions in parallel. Measure the actual wait time after 3 weeks and upgrade the GPU if necessary. The motherboard and power supply must, however, support a larger GPU from the start.

#2. Recommended technical stack

Four open-source building blocks are enough. No vendor has a say, no license needs renewal, and everything runs in Docker on the workstation.

Ollama
The daemon that loads and serves the model. Listens by default on http://localhost:11434. Automatically detects the NVIDIA GPU and handles quantization. MIT License.
Open WebUI
The ChatGPT-style web interface. Native multi-user support, RBAC, OIDC integration for SSO, per-user history. BSD-3-Clause license.
Qdrant
Vector database for RAG. Indexes your internal documents (procedures, template contracts, product sheets). Faster and simpler than pgvector for this volume. Apache 2 license.
Traefik or Nginx
Reverse proxy to expose Open WebUI over HTTPS on the LAN with an internal certificate, terminate TLS, and route to the backends.
docker-compose.yml — minimal skeleton
services:
  ollama:
    image: ollama/ollama:latest
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    volumes:
      - ollama_data:/root/.ollama
    restart: unless-stopped

  qdrant:
    image: qdrant/qdrant:latest
    volumes:
      - qdrant_data:/qdrant/storage
    restart: unless-stopped

  openwebui:
    image: ghcr.io/open-webui/open-webui:main
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
      - QDRANT_URL=http://qdrant:6333
      - ENABLE_OAUTH_SIGNUP=true
      - OAUTH_PROVIDER_NAME=Microsoft
      - MICROSOFT_CLIENT_ID=${ENTRA_CLIENT_ID}
      - MICROSOFT_CLIENT_SECRET=${ENTRA_CLIENT_SECRET}
      - MICROSOFT_CLIENT_TENANT_ID=${ENTRA_TENANT_ID}
    volumes:
      - openwebui_data:/app/backend/data
    depends_on:
      - ollama
      - qdrant
    restart: unless-stopped

volumes:
  ollama_data:
  qdrant_data:
  openwebui_data:

This skeleton is intentionally short. You will add Traefik on top for HTTPS, plus a bind-mounted volume for nightly backups to the NAS. Nothing more.

#3. SSO with Microsoft Entra ID

Most SMBs have Microsoft 365. Connecting the chatbot's authentication to Entra ID (formerly Azure AD) eliminates orphaned accounts and ensures that an employee who leaves loses access the same day. Open WebUI natively supports OIDC with Microsoft.

  1. 01
    Create an App registration
    In the Entra ID portal, App registrations → New registration. Name: “Internal AI Chatbot.” Redirect URI: https://chatbot.votreboite.local/oauth/microsoft/callback (Web).
  2. 02
    Retrieve the credentials
    Note the Application (client) ID and the Directory (tenant) ID. Under Certificates & secrets, generate a Client secret with a 24-month expiration and note it immediately (it will never be displayed again).
  3. 03
    Define permissions
    API permissions → Microsoft Graph → Delegated: openid, profile, email, User.Read. Application permissions aren't needed; the flow is delegated.
  4. 04
    Restrict access
    Enterprise applications → your app → Properties → Assignment required = Yes. Then, under Users and groups, add the “Tous les collaborateurs” group or a pilot group. Without this step, any account in the tenant can sign in.
  5. 05
    Inject into Open WebUI
    Copy CLIENT_ID, CLIENT_SECRET, and TENANT_ID into the docker-compose .env file. Restart. The “Sign in with Microsoft” button appears on the login page.
!
The classic pitfall: Assignment required
If you miss step 4, your entire Microsoft tenant could potentially connect — including all external guest accounts (customers, contractors). Double-check this setting before sharing the URL with the teams.

#4. 4-week deployment plan

This timeline assumes an internal technical lead (IT department, IT manager, or service provider) who dedicates ~50% of their time to the project. An SMB does not need more.

#Week 1 — Hardware and installation

J1-J2
Hardware order (workstation, 1500 VA UPS, managed switch if not already in place).
J3-J4
Receiving, assembly if needed, installation of Ubuntu Server LTS 24.04, drivers NVIDIA, Docker + NVIDIA Container Toolkit.
J5
Pull Docker images, first ollama pull mistral-small, raw chat test in the CLI. You should see tokens streaming out.

#Week 2—Stack and integrations

J6-J7
docker-compose up, configuration Open WebUI (admin, paramètres généraux, branding interne).
J8-J9
Entra ID SSO configuration, testing with 3 pilot accounts, and validation of the login and logout flow.
J10
Traefik reverse proxy with an internal certificate through your PKI, or Let’s Encrypt DNS-01 if the domain is public. Test from 2–3 LAN machines.

#Week 3 — RAG and pilot

J11-J12
Select 50 to 200 relevant internal documents (procedures, FAQs, sales sheets). Index them in Qdrant via Open WebUI or a Python script.
J13-J14
SMB-colored system prompt: tone, permitted topics, topics to refuse (sensitive HR, salaries, etc.). Test with the 3 pilot users.
J15
Open the pilot to a representative group of 5–8 employees (1 sales, 1 HR, 1 accounting, 1 operations…). Collect structured feedback.

#Week 4 — Training and cutover

J16-J17
Adjustments based on pilot feedback (system prompt, RAG documents, temperature settings).
J18
1h30 group training session: demo, role-specific use cases, usage rules, and what NOT to put in the chat (health data, identifiers, etc.).
J19
Gradual rollout to all employees through the Entra ID group. Internal announcement.
J20
Set up monitoring (Glances or Netdata for the workstation, plus an aggregated conversation log to identify requested topics).

#5. Training and change management

An internal chatbot without training means 70% of the investment is wasted. Employees don't know what to ask it, mentally compare it with ChatGPT for the general public, and conclude that “it's not as good” after two attempts.

Format
A group session of 1h30 in batches of 10-15 people. No slides for 1h. 20 min of demos + 1h of hands-on practice with their real topics.
Use cases by profession
Prepare 3 concrete prompts for each function (HR, sales, accounting, ops, management). People copy and adapt them—that’s exactly what we want.
Usage rules
Clearly indicate what may be pasted (internal texts, drafts, already-public data) and what must not be pasted (health information, convictions, named banking data, technical identifiers).
Reference
Appoint 1 AI lead per department. This person shares best practices locally and reports gaps in the RAG's knowledge.
30-day follow-up
Measure usage. If some teams never touch the tool, that is a signal—either they are missing an obvious use case, or the training did not stick.
→
What works best in a demo
Three cases to show first: rewriting a difficult email, summarizing a long meeting report, and translating a technical document. These are the use cases whose value is obvious in 30 seconds.

#6. GDPR and AI Act compliance

The fact that the system is on-premises greatly simplifies compliance, but does not eliminate it. GDPR obligations apply to the processing, not its location.

Record of processing activities
Add a record for the internal chatbot. Purpose: drafting assistance and document search. Legal basis: the employer's legitimate interest. Data processed: conversation content, SSO identifier.
Retention period
Define a clear policy: 90 days of conversation history per user, with automatic purging afterward. Document it in the charter.
Employee information
Update the IT policy and the CSE/CSE information notice. Explicitly state that conversations are stored and accessible to the technical administrator in the event of an incident.
AI Act
An internal productivity-assistance chatbot falls into the « limited risk » category (Article 50). Main obligation: inform the user that they are interacting with AI — the Open WebUI interface does this by default.
Technical security
Encrypted backups of the Open WebUI database, audited admin access, MFA required for Entra ID accounts, monthly Docker image updates.
i
Is a DPIA necessary?
A data protection impact assessment (DPIA) is not mandatory for a standard productivity use case, but it is recommended if you plan to connect RAG to sensitive data (HR records, identifiable customer data). If in doubt, your DPO decides.

#7. ROI: the honest calculation

Take a small business with 30 employees and a Profile A deployment costing €5,000. Over a 3-year period.

Alternative SaaS cost
30 × 25 € × 12 months × 3 years = 27,000 € (ChatGPT Team).
Internal cost
Workstation 5 000 € + electricity ~200 €/year + 5 person-days of installation in the 1st year (~3 500 €) + 2 person-days/year of maintenance (~1 400 €/year × 2 years) = ~11 500 € over 3 years.
Direct net savings
~15 500 € over 3 years, or about 5 200 € per year starting in year 2.
Indirect gains
Reduced data-leakage risk (you can make the case that avoiding a single leak more than covers the project), easier GDPR compliance, and independence from SaaS price increases.
Productivity gain
Conservative estimate: 15 minutes saved per user per working day. At €35/hour fully loaded and 220 days, that comes to ~€38,500 per year for 30 people. This figure depends heavily on actual adoption—be cautious in your business cases.
!
The classic productivity ROI mistake
Counting 1 hour/day saved and multiplying it out is wrong. Most regular users save 10 to 20 minutes per day, and 40% of accounts become inactive after 2 months. Use a generous adoption estimate (50–60%) and a conservative per-user gain (15 min). The ROI remains strongly positive.

#Go further

This guide lays the foundation. Three natural next steps: add serious RAG over your internal documents so the chatbot knows your procedures, harden the multi-user deployment on your intranet, and formalize GDPR compliance with a proper file.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.