Advanced 16 minDeployment

Enterprise coding AI: protect proprietary code, NDAs, and secrets industriel

A developer pastes a business function into a cloud assistant to refactor it. In one second, a code fragment covered by an assignment clause and an NDA has left your perimeter, passing through the servers of a U.S. third party. For a software publisher, engineering firm, or IT services company operating under trade-secret protection, this innocuous action is a legal and contractual vulnerability. This guide is for the technical lead or CIO who wants to equip developers with a high-performance copilot without a single line of proprietary code ever leaving the company network. We will examine why GitHub Copilot and Cursor are problematic despite their “enterprise” options, what the GDPR and AI Act actually require, and how to deploy a 100% local stack (Ollama, Cline, Aider, Tabby) on a workstation or shared GPU server, with an auditable no-exfiltration policy.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Ubuntu 24.04

#The Real Risk: Your Code Is Sent to a Third Party

Source code is not data like any other. It contains your algorithms, trade secrets, hard-coded API keys (it happens), your security architecture, and, legally, it is often covered by a rights-assignment clause benefiting a client. A cloud code assistant reads the context around the cursor, sometimes the entire repository for indexing, and sends these fragments to a remote model. The risk is not theoretical: it combines trade-secret leakage, NDA violations, and GDPR noncompliance as soon as a comment contains personal data.

Trade secret
A proprietary algorithm disclosed to a third party loses its status as a trade secret (Article L151-1 of the French Commercial Code): legal protection is lost.
Assignment clause
For code delivered to a client, the contract often prohibits any communication with an unapproved subcontractor. A cloud assistant is an undeclared subcontractor.
Vendor NDA
You are working on a partner’s code under a confidentiality agreement: sending it to OpenAI or Anthropic is a direct violation.
Personal data
A test dataset, an inline log, or a comment containing an email address brings the transmission under the GDPR.
!
The trap in “we don't train on your data”
The no-training promise (zero data retention) solves only part of the problem. The code still travels over the network, is processed in memory on servers outside the EU, and remains subject to the U.S. CLOUD Act. No training does not mean no transfer.

#Why Copilot and Cursor violate your NDAs

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

GitHub Copilot and Cursor are excellent productivity tools, but their business model relies on hosted LLMs (Copilot on Azure/OpenAI infrastructure, Cursor on OpenAI and Anthropic). Even with Business or Enterprise plans, your code leaves your machine to be completed on the server side. “Content exclusion” or “privacy mode” options reduce retention, but the network flow to a third party remains. For an NDA clause that prohibits any disclosure to an unnamed third party, this flow is a violation, regardless of the retention policy.

Copilot Business
Code sent to GitHub/Azure for completion. “No training” enabled by default, but data is transferred outside the EU and submitted to the CLOUD Act.
Cursor (Privacy Mode)
Private mode prevents retention on Cursor’s side, but requests still go through OpenAI and Anthropic APIs.
Cursor without Privacy Mode
The code may be retained and used to improve the product. Unacceptable under a strict NDA.
Common contractual limit
None of these offerings signs a DPA that covers a clause assigning client code to a specific third-party subcontractor.
i
The decisive test
Ask your legal team: 'May I send client X's code to a third-party server Y located in the United States?'. If the answer is no for even one of your contracts, no cloud assistant can be deployed consistently across the team. Local deployment then becomes the only coherent policy.

#What the GDPR and AI Act require from source code

The GDPR does not refer to “source code” but to personal data. Code contains more of it than you might think: emails in comments, test identifiers, seed data, and log traces. As soon as any of these elements is sent to a cloud assistant, you are carrying out a data transfer, which requires a legal basis, a processor under a DPA (Art. 28), and, outside the EU, a valid transfer mechanism (standard contractual clauses). The AI Act, meanwhile, classifies most coding assistants as limited risk, with primarily transparency obligations; but the real issue for you comes earlier, in governing the model’s input data.

GDPR Art. 28
Any cloud assistant processing your data is a processor and requires a signed DPA. Many teams use one without a contract.
GDPR transfer outside the EU
Without EU hosting, you need safeguards (SCCs) and a transfer impact assessment. Local hosting eliminates the issue entirely.
AI Act (transparency)
Limited risk for the coding assistant: disclose that content is AI-generated. Low burden, but document it.
Minimization by design
The simplest path to compliance: don’t transfer anything at all. A local stack is minimized by design.
→
Local: compliance by subtraction
Rather than piling on DPAs, SCCs, impact assessments, and content exclusions for a cloud tool, 100% local deployment eliminates data transfer. No transfer, no processor, no GDPR question about model input. This is the most defensible approach before an auditor.

#The 100% local stack: Ollama, Cline, Aider, Tabby

A local copilot stack has two layers: an inference server that runs the model on your hardware, and clients that connect to it from the IDE or terminal. The reference server is Ollama, which exposes a local HTTP API and handles downloading GGUF models. On top of that, three complementary clients cover the main use cases: Cline for the agent in VS Code, Aider for commit-oriented terminal pair programming, and Tabby for Copilot-style autocompletion in shared-server mode for the entire team.

Ollama
Local inference server (MIT). Serves models via http://localhost:11434. No code telemetry and no outbound calls for inference.
Cline
VS Code extension (agent). Reads and writes files, runs commands, and plans multi-file tasks. Connects to Ollama as a local provider.
Aider
Pair-programming CLI (Apache 2.0). Edits code and makes Git commits; excellent for guided refactoring. Points to the Ollama API.
Tabby
Self-hosted autocomplete server (Apache 2.0). Replaces Copilot for inline suggestions; ideal for shared GPU server mode.
Install the Ollama foundation + a coding model
# Serveur d'inference local
curl -fsSL https://ollama.com/install.sh | sh

# Modele de code recommande pour 16 Go VRAM (specialiste agent de code)
ollama pull devstral:24b

# Verifier que l'API locale repond (aucun appel sortant)
curl http://localhost:11434/api/tags
Connect Aider to Ollama (terminal)
pip install aider-install && aider-install

# Pointer Aider vers le serveur local, jamais vers une API cloud
export OLLAMA_API_BASE=http://127.0.0.1:11434
aider --model ollama_chat/devstral:24b

For Cline in VS Code, select the 'Ollama' provider in the extension settings and enter the server URL (local or that of your internal GPU server). No cloud API key is entered: this is the technical guarantee that no fragment is sent to OpenAI or Anthropic. Tabby, meanwhile, runs in a container on the GPU server, and workstations point to its internal endpoint.

#Architecture: isolated workstation vs. shared GPU server

Two topologies dominate. The isolated workstation runs Ollama directly on the developer's machine (M-series Mac or RTX PC): the code never leaves the workstation, which is ideal for the strictest NDAs, but it's limited by individual VRAM and expensive to deploy at scale. The shared GPU server centralizes one or more GPUs on the internal network; workstations connect to it through the API. You share a larger model, rationalize the hardware, and keep the data flow strictly within the network (LAN or corporate VPN).

Comparison of the two deployment topologies
CriterionIsolated workstation (Ollama local)Shared GPU server (Tabby/Ollama)
Code scopeNever leave the machineStays within the internal network (LAN/VPN)
Model sizeLimited by the machine's VRAMLarger shared model (24–80 GB)
Hardware costHigh (1 GPU per developer)Streamlined (1 server for N developers)
Real-time autocompleteGood if you have a local GPUExcellent with Tabby + batching
Strict NDA complianceMaximum (zero network traffic)Strong (internal flow only, traceable)
MaintenanceDecentralized, heterogeneousCentralized, managed updates
→
The recommended hybrid architecture
In practice, you combine Tabby on a GPU server for team-wide shared autocompletion (high throughput, batching) with Ollama locally on workstations handling the most sensitive repositories or for offline work. All under a network policy that blocks known cloud AI endpoints.

#Which coding models for which VRAM

The quality of a local copilot depends on the model. Three open families dominate coding in 2026: Qwen3-Coder (Alibaba, MoE 30B-A3B, 256k context, fast thanks to its 3B active parameters, Apache 2.0), Devstral (Mistral AI, a 24B model designed for agentic use with Cline and OpenHands, Apache 2.0), and strong general-purpose models for coding such as Qwen 3.8 27B or GLM 4.7 Flash. The choice depends on available VRAM and the use case: a small, fast model for Tabby autocompletion, a larger one for Cline’s agentic reasoning.

Local coding models by VRAM budget (Q4 quantization)
VRAMRecommended modelTypical use
8 GBQwen2.5-Coder 7B base (FIM)Tabby inline autocompletion (the 2026 FIM benchmark), fast completions
16 GBDevstral 24B / gpt-oss 20BCline and Aider agents on medium-sized repositories
24 GBQwen3-Coder 30B-A3B / Qwen 3.8 27BMulti-file refactoring, reasoning
48-80 GBQwen3-Coder 30B-A3B Q8 / Granite 4.2 30BMulti-developer shared server, long context
i
Verify actual VRAM before making promises
The advertised VRAM required by a model depends on quantization and context length (the KV cache grows with context). A Qwen3-Coder 30B-A3B in Q4 (19 GB) fits on 24 GB with a moderate context, but a 256k context may exceed the limit. Test with your actual workload.

#“Local provider only” policy

A local stack protects you only when it is enforced by a policy. “Local provider only” means no development tool can point to a cloud AI API. This has three complementary layers: tool configuration (no cloud key), network blocking (cloud AI endpoints are unreachable from development workstations), and an organizational rule (signed policy). The combination of all three makes the policy genuinely enforceable during an audit.

  1. 01
    Block cloud API keys
    There must be no OPENAI_API_KEY, ANTHROPIC_API_KEY, or equivalent variable on development machines. Cline and Aider are configured exclusively to use the internal Ollama endpoint.
  2. 02
    Block endpoints at the firewall
    Filter out api.openai.com, api.anthropic.com, Copilot domains, and Cursor. Completions then physically cannot leave the network.
  3. 03
    Pin the extension configuration
    Deploy VS Code/Cline settings through GPO or MDM to prevent a developer from pointing back to a cloud provider.
  4. 04
    Charter and training
    A signed policy reiterates the ban on pasting code into a cloud chatbot (the residual risk is human, not technical).
  5. 05
    Log access to the inference server
    The Ollama/Tabby server logs prove that completions are served internally. A key piece of the audit.

#Audit non-exfiltration

The “it’s local” argument only holds if it can be proven. A non-exfiltration audit consists of demonstrating, with evidence, that no code fragment leaves the perimeter while the copilot is in use. The most convincing method is network observation: capture a workstation’s traffic during an intensive coding session and verify that no connection goes to a cloud AI endpoint. Complement this with an inspection of the configuration and server logs.

Capture and verify outbound traffic during a session
# 1. Capturer le trafic du poste pendant une session de codage Cline/Aider
sudo tcpdump -i any -n 'tcp port 443' -w /tmp/session_code.pcap

# 2. Lister les IP/destinations contactees (hors reseau interne)
tcpdump -r /tmp/session_code.pcap -n | awk '{print $5}' | cut -d. -f1-4 | sort -u

# 3. Verifier qu'aucune resolution ne vise un endpoint d'IA cloud
grep -Ei 'openai|anthropic|githubcopilot|cursor' /var/log/dnsmasq.log || echo 'OK : aucune requete IA cloud'
Network evidence
tcpdump/Wireshark capture: only the internal server and internal Git repository IPs appear. Zero cloud AI endpoints.
Configuration proof
Cline/Aider settings export showing the internal Ollama provider and the absence of a cloud key.
Proof of blocking
Negative test: a manual attempt to reach api.openai.com from a developer workstation fails (firewall).
Server proof
Timestamped Ollama/Tabby logs correlating completions with internal workstations.
→
Document once, replay often
Script the audit (capture + analysis + report) and schedule it periodically. A reproducible audit folder is far more valuable than a one-off claim when the client or the CNIL asks questions.

#An honest trade-off: local vs. cloud

To be fair, the cloud still leads in the raw quality of the largest models and in zero-effort infrastructure. A proprietary cloud assistant may outperform a local Qwen3-Coder 30B-A3B on some complex reasoning tasks. Local deployment requires a hardware investment, a team to maintain the inference server, and slightly lower quality on the most demanding tasks. The right tradeoff isn't ideological: it depends on how sensitive your code is.

Choose local if
You have code under an NDA, assignment clauses, trade secrets, or clients who prohibit non-approved third-party subcontractors.
The cloud may be sufficient if
Your code is open source, contains no personal data, and carries no contractual confidentiality obligation to a third party.
The true cost of local deployment
A depreciated GPU server (24–48 GB) shared by a team often costs less than per-seat cloud licenses over 2–3 years.
The hidden risk of the cloud
The cost of a single NDA violation (lost contract, litigation) exceeds years of local licensing.
i
The decision rule in one sentence
If you can't answer 'yes' to 'can I send this code to a third party?' for all your repositories, deploy a consistent local stack: it's easier to govern than a mixed cloud/local environment handled case by case.

#Conclusion and implementation

Equipping a development team with a capable copilot without ever exposing proprietary code is now realistic: Ollama as the inference server, Cline and Aider for the agent and terminal, Tabby for shared autocomplete, all under a 'local provider only' policy and a reproducible non-exfiltration audit. GDPR and AI Act compliance is achieved through subtraction: no transfer, no processor, no question. To save time on setup, the paid guide “Local Code Copilot” provides a turnkey pack: Ollama + Cline + Aider; ready-to-use configurations, model choices by VRAM budget, a network policy blocking cloud endpoints, and the non-exfiltration audit script. Enough to deploy a stack in a few hours that you can defend to your legal team and customers.

Frequently asked questions
Is Copilot Enterprise really incompatible with an NDA?+
It depends on the NDA. If your agreement prohibits any disclosure to a third party that is not specifically authorized by name, yes: even with no-training enabled, the code passes through GitHub/Azure infrastructure, which constitutes disclosure to a third party. Have your legal team review each client contract before any cloud deployment.
Is a local Ollama stack powerful enough to replace Copilot?+
For autocompletion and everyday pair programming, yes: Tabby handles inline suggestions (with Qwen2.5-Coder 7B base, still the FIM reference in 2026), and for agentic chat, Devstral 24B or Qwen3-Coder 30B-A3B are competitive on most tasks. For highly complex reasoning, the largest cloud models still have an edge, but the gap has narrowed and the privacy guarantee changes the equation.
Which GPU should you choose for a team of 10 developers?+
A shared GPU server with 24 to 48 GB of VRAM (for example, a 24 GB card to start) running Tabby and Ollama is generally enough for about ten developers using autocompletion, with a Devstral 24B or a Qwen3-Coder 30B-A3B. Measure your actual workload: long context increases the KV cache and the required VRAM.
Does the GDPR really require local processing for code?+
No, the GDPR does not require local deployment. As soon as personal data is transferred (comments, logs, test games), it requires a legal basis, a DPA, and, outside the EU, transfer safeguards. Local deployment is simply the easiest path to compliance because it eliminates the transfer and therefore most of these obligations.
How can you prove to a client that not a single line leaks?+
Use a reproducible audit package: capture network traffic during a coding session showing zero connections to a cloud AI endpoint, export the tool configuration without a cloud key, perform a negative test confirming that firewall access is blocked, and collect the internal inference server logs. Script everything so it can be replayed on demand.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.