Enterprise coding AI: protect proprietary code, NDAs, and secrets industriel
A developer pastes a business function into a cloud assistant to refactor it. In one second, a code fragment covered by an assignment clause and an NDA has left your perimeter, passing through the servers of a U.S. third party. For a software publisher, engineering firm, or IT services company operating under trade-secret protection, this innocuous action is a legal and contractual vulnerability. This guide is for the technical lead or CIO who wants to equip developers with a high-performance copilot without a single line of proprietary code ever leaving the company network. We will examine why GitHub Copilot and Cursor are problematic despite their “enterprise” options, what the GDPR and AI Act actually require, and how to deploy a 100% local stack (Ollama, Cline, Aider, Tabby) on a workstation or shared GPU server, with an auditable no-exfiltration policy.
#The Real Risk: Your Code Is Sent to a Third Party
Source code is not data like any other. It contains your algorithms, trade secrets, hard-coded API keys (it happens), your security architecture, and, legally, it is often covered by a rights-assignment clause benefiting a client. A cloud code assistant reads the context around the cursor, sometimes the entire repository for indexing, and sends these fragments to a remote model. The risk is not theoretical: it combines trade-secret leakage, NDA violations, and GDPR noncompliance as soon as a comment contains personal data.
- Trade secret
- A proprietary algorithm disclosed to a third party loses its status as a trade secret (Article L151-1 of the French Commercial Code): legal protection is lost.
- Assignment clause
- For code delivered to a client, the contract often prohibits any communication with an unapproved subcontractor. A cloud assistant is an undeclared subcontractor.
- Vendor NDA
- You are working on a partner’s code under a confidentiality agreement: sending it to OpenAI or Anthropic is a direct violation.
- Personal data
- A test dataset, an inline log, or a comment containing an email address brings the transmission under the GDPR.
#Why Copilot and Cursor violate your NDAs
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
GitHub Copilot and Cursor are excellent productivity tools, but their business model relies on hosted LLMs (Copilot on Azure/OpenAI infrastructure, Cursor on OpenAI and Anthropic). Even with Business or Enterprise plans, your code leaves your machine to be completed on the server side. “Content exclusion” or “privacy mode” options reduce retention, but the network flow to a third party remains. For an NDA clause that prohibits any disclosure to an unnamed third party, this flow is a violation, regardless of the retention policy.
- Copilot Business
- Code sent to GitHub/Azure for completion. “No training” enabled by default, but data is transferred outside the EU and submitted to the CLOUD Act.
- Cursor (Privacy Mode)
- Private mode prevents retention on Cursor’s side, but requests still go through OpenAI and Anthropic APIs.
- Cursor without Privacy Mode
- The code may be retained and used to improve the product. Unacceptable under a strict NDA.
- Common contractual limit
- None of these offerings signs a DPA that covers a clause assigning client code to a specific third-party subcontractor.
#What the GDPR and AI Act require from source code
The GDPR does not refer to “source code” but to personal data. Code contains more of it than you might think: emails in comments, test identifiers, seed data, and log traces. As soon as any of these elements is sent to a cloud assistant, you are carrying out a data transfer, which requires a legal basis, a processor under a DPA (Art. 28), and, outside the EU, a valid transfer mechanism (standard contractual clauses). The AI Act, meanwhile, classifies most coding assistants as limited risk, with primarily transparency obligations; but the real issue for you comes earlier, in governing the model’s input data.
- GDPR Art. 28
- Any cloud assistant processing your data is a processor and requires a signed DPA. Many teams use one without a contract.
- GDPR transfer outside the EU
- Without EU hosting, you need safeguards (SCCs) and a transfer impact assessment. Local hosting eliminates the issue entirely.
- AI Act (transparency)
- Limited risk for the coding assistant: disclose that content is AI-generated. Low burden, but document it.
- Minimization by design
- The simplest path to compliance: don’t transfer anything at all. A local stack is minimized by design.
#The 100% local stack: Ollama, Cline, Aider, Tabby
A local copilot stack has two layers: an inference server that runs the model on your hardware, and clients that connect to it from the IDE or terminal. The reference server is Ollama, which exposes a local HTTP API and handles downloading GGUF models. On top of that, three complementary clients cover the main use cases: Cline for the agent in VS Code, Aider for commit-oriented terminal pair programming, and Tabby for Copilot-style autocompletion in shared-server mode for the entire team.
- Ollama
- Local inference server (MIT). Serves models via http://localhost:11434. No code telemetry and no outbound calls for inference.
- Cline
- VS Code extension (agent). Reads and writes files, runs commands, and plans multi-file tasks. Connects to Ollama as a local provider.
- Aider
- Pair-programming CLI (Apache 2.0). Edits code and makes Git commits; excellent for guided refactoring. Points to the Ollama API.
- Tabby
- Self-hosted autocomplete server (Apache 2.0). Replaces Copilot for inline suggestions; ideal for shared GPU server mode.
For Cline in VS Code, select the 'Ollama' provider in the extension settings and enter the server URL (local or that of your internal GPU server). No cloud API key is entered: this is the technical guarantee that no fragment is sent to OpenAI or Anthropic. Tabby, meanwhile, runs in a container on the GPU server, and workstations point to its internal endpoint.
#Architecture: isolated workstation vs. shared GPU server
Two topologies dominate. The isolated workstation runs Ollama directly on the developer's machine (M-series Mac or RTX PC): the code never leaves the workstation, which is ideal for the strictest NDAs, but it's limited by individual VRAM and expensive to deploy at scale. The shared GPU server centralizes one or more GPUs on the internal network; workstations connect to it through the API. You share a larger model, rationalize the hardware, and keep the data flow strictly within the network (LAN or corporate VPN).
| Criterion | Isolated workstation (Ollama local) | Shared GPU server (Tabby/Ollama) |
|---|---|---|
| Code scope | Never leave the machine | Stays within the internal network (LAN/VPN) |
| Model size | Limited by the machine's VRAM | Larger shared model (24–80 GB) |
| Hardware cost | High (1 GPU per developer) | Streamlined (1 server for N developers) |
| Real-time autocomplete | Good if you have a local GPU | Excellent with Tabby + batching |
| Strict NDA compliance | Maximum (zero network traffic) | Strong (internal flow only, traceable) |
| Maintenance | Decentralized, heterogeneous | Centralized, managed updates |
#Which coding models for which VRAM
The quality of a local copilot depends on the model. Three open families dominate coding in 2026: Qwen3-Coder (Alibaba, MoE 30B-A3B, 256k context, fast thanks to its 3B active parameters, Apache 2.0), Devstral (Mistral AI, a 24B model designed for agentic use with Cline and OpenHands, Apache 2.0), and strong general-purpose models for coding such as Qwen 3.8 27B or GLM 4.7 Flash. The choice depends on available VRAM and the use case: a small, fast model for Tabby autocompletion, a larger one for Cline’s agentic reasoning.
| VRAM | Recommended model | Typical use |
|---|---|---|
| 8 GB | Qwen2.5-Coder 7B base (FIM) | Tabby inline autocompletion (the 2026 FIM benchmark), fast completions |
| 16 GB | Devstral 24B / gpt-oss 20B | Cline and Aider agents on medium-sized repositories |
| 24 GB | Qwen3-Coder 30B-A3B / Qwen 3.8 27B | Multi-file refactoring, reasoning |
| 48-80 GB | Qwen3-Coder 30B-A3B Q8 / Granite 4.2 30B | Multi-developer shared server, long context |
#“Local provider only” policy
A local stack protects you only when it is enforced by a policy. “Local provider only” means no development tool can point to a cloud AI API. This has three complementary layers: tool configuration (no cloud key), network blocking (cloud AI endpoints are unreachable from development workstations), and an organizational rule (signed policy). The combination of all three makes the policy genuinely enforceable during an audit.
- 01Block cloud API keysThere must be no OPENAI_API_KEY, ANTHROPIC_API_KEY, or equivalent variable on development machines. Cline and Aider are configured exclusively to use the internal Ollama endpoint.
- 02Block endpoints at the firewallFilter out api.openai.com, api.anthropic.com, Copilot domains, and Cursor. Completions then physically cannot leave the network.
- 03Pin the extension configurationDeploy VS Code/Cline settings through GPO or MDM to prevent a developer from pointing back to a cloud provider.
- 04Charter and trainingA signed policy reiterates the ban on pasting code into a cloud chatbot (the residual risk is human, not technical).
- 05Log access to the inference serverThe Ollama/Tabby server logs prove that completions are served internally. A key piece of the audit.
#Audit non-exfiltration
The “it’s local” argument only holds if it can be proven. A non-exfiltration audit consists of demonstrating, with evidence, that no code fragment leaves the perimeter while the copilot is in use. The most convincing method is network observation: capture a workstation’s traffic during an intensive coding session and verify that no connection goes to a cloud AI endpoint. Complement this with an inspection of the configuration and server logs.
- Network evidence
- tcpdump/Wireshark capture: only the internal server and internal Git repository IPs appear. Zero cloud AI endpoints.
- Configuration proof
- Cline/Aider settings export showing the internal Ollama provider and the absence of a cloud key.
- Proof of blocking
- Negative test: a manual attempt to reach api.openai.com from a developer workstation fails (firewall).
- Server proof
- Timestamped Ollama/Tabby logs correlating completions with internal workstations.
#An honest trade-off: local vs. cloud
To be fair, the cloud still leads in the raw quality of the largest models and in zero-effort infrastructure. A proprietary cloud assistant may outperform a local Qwen3-Coder 30B-A3B on some complex reasoning tasks. Local deployment requires a hardware investment, a team to maintain the inference server, and slightly lower quality on the most demanding tasks. The right tradeoff isn't ideological: it depends on how sensitive your code is.
- Choose local if
- You have code under an NDA, assignment clauses, trade secrets, or clients who prohibit non-approved third-party subcontractors.
- The cloud may be sufficient if
- Your code is open source, contains no personal data, and carries no contractual confidentiality obligation to a third party.
- The true cost of local deployment
- A depreciated GPU server (24–48 GB) shared by a team often costs less than per-seat cloud licenses over 2–3 years.
- The hidden risk of the cloud
- The cost of a single NDA violation (lost contract, litigation) exceeds years of local licensing.
#Conclusion and implementation
Equipping a development team with a capable copilot without ever exposing proprietary code is now realistic: Ollama as the inference server, Cline and Aider for the agent and terminal, Tabby for shared autocomplete, all under a 'local provider only' policy and a reproducible non-exfiltration audit. GDPR and AI Act compliance is achieved through subtraction: no transfer, no processor, no question. To save time on setup, the paid guide “Local Code Copilot” provides a turnkey pack: Ollama + Cline + Aider; ready-to-use configurations, model choices by VRAM budget, a network policy blocking cloud endpoints, and the non-exfiltration audit script. Enough to deploy a stack in a few hours that you can defend to your legal team and customers.
Is Copilot Enterprise really incompatible with an NDA?+
Is a local Ollama stack powerful enough to replace Copilot?+
Which GPU should you choose for a team of 10 developers?+
Does the GDPR really require local processing for code?+
How can you prove to a client that not a single line leaks?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.