Qwen Code: the terminal coding agent with Ollama
Qwen Code is the command-line coding agent published by Qwen's team at Alibaba. It reads your repository, modifies files, runs commands, and chains steps until the task is complete, like Claude Code or OpenCode. The point that matters here is that it speaks the OpenAI protocol, so it can connect to a model running locally through Ollama or LM Studio. This guide covers installation, local connectivity, context tuning—which makes the difference between a useful agent and one that spins in circles—and the limitations to know before adopting it.
#What Qwen Code is and why to run it locally
Qwen Code is a fork of Gemini CLI, Google's open-source terminal agent, which the Qwen team adapted for its Qwen3-Coder models. The project is released under the Apache 2.0 license on GitHub (QwenLM/qwen-code), installed via npm, and used with the qwen command. It follows the mechanics of modern coding agents: a model receives your request, has access to tools (reading and writing files, searching the repository, shell execution, web requests, and MCP servers), and loops over those tools until it produces a verifiable result.
By default, Qwen Code routes you to a “Qwen OAuth” connection: you sign in with an Qwen account, and requests are sent to Alibaba Cloud's servers. A free offering is available through this route, but its quotas may change and depend on the region. We therefore provide no figure here: the authentication page in the official documentation is the only up-to-date source. What does not change is the other mode, called “OpenAI-compatible”: Qwen Code accepts any server that exposes the OpenAI API, including Ollama and LM Studio on your machine.
- Privacy
- In local mode, source code, executed commands, and their output never leave the machine. That's the decisive argument for client code or code covered by a confidentiality agreement.
- Cost
- No quota, no per-token billing. The only cost is electricity and hardware you have already purchased.
- Availability
- No service outage, no queue during peak hours. The agent responds as long as the GPU is running.
- Trade-off
- A 7- to 30-billion-parameter model in Q4 quantization is not on the same level as a cloud model with several hundred billion parameters. You need to break tasks down more finely and review more carefully.
#Prerequisites
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
- Node.js 20 or later
- Qwen Code is an npm package. Check with node --version. On Linux and macOS, nvm or fnm prevent permission problems during global installation.
- Ollama installed and working
- The daemon is listening on http://localhost:11434. ollama list should respond without an error. If it does not, start with the Ollama installation guide.
- A model that supports tool calling
- This is non-negotiable: a coding agent chains together structured tool calls. The Qwen3-Coder and Qwen2.5-Coder model families, as well as Devstral, support them under Ollama. A model without tool support will produce text instead of actions, and the agent will remain stuck.
- GPU memory
- Q4_K_M guidelines: a 7B uses about 5 GB of VRAM, a 14B about 9 GB, and a 32B about 19 GB, excluding context. The long context an agent requires adds several gigabytes: leave plenty of headroom.
- A Git repository
- Not required, but strongly recommended. The agent modifies files; git diff and git checkout are your safety net.
#1. Install Qwen Code
The installation recommended by the repository uses npm globally. On macOS, a Homebrew package is also published. The binary is called qwen.
When qwen is launched without configuration for the first time, the tool prompts you to choose an authentication method. Do not choose Qwen OAuth if your goal is local use: select the OpenAI option, or better yet, exit and first prepare the configuration described in the next step. You can always change the method later with the /auth command in a session.
#2. Connect Qwen Code to Ollama
Ollama exposes an OpenAI-compatible API at the /v1 path on port 11434. Qwen Code reads three environment variables for this mode: the base URL, an API key, and the model name. Ollama does not require a key, but Qwen Code rejects an empty value, so use any arbitrary string.
- 01Download a tools-compatible coding modelTake Qwen3-Coder 30B-A3B, a mixture-of-experts (MoE) model with 30 billion parameters, 3 billion of which are active for each token, making it fast for its size. It weighs around 19 GB in Q4: you need 24 GB of VRAM, or an Apple Silicon machine with at least 32 GB of unified memory, to keep it entirely on the GPU. On a 12 GB card, choose qwen2.5-coder:7b or a 14B model instead.
- 02Verify that the OpenAI API from Ollama respondsA request to /v1/models should list your models. If it fails, Ollama is not running or is listening on another address.
- 03Create the .env file at the project rootQwen Code automatically loads a .env file present in the current directory, in a .qwen subdirectory of the project, or in ~/.qwen for global configuration. The file closest to the working directory takes precedence.
- 04Launch qwen in the projectThe bottom of the window displays the active model. If you see your model's name, Ollama, the connection is working. Enter a first simple request, such as summarizing the repository structure, to verify that the reading tools work.
The same parameters can be passed as command-line options for a one-off session without touching the .env file. This is useful for testing a second model without breaking the configuration that works.
#3. Alternative: LM Studio as a server
If you prefer LM Studio, the principle is the same. Load a coding model in the application, open the Developer tab, and start the local server: it listens on port 1234 by default and exposes the same OpenAI-compatible API. Be sure to enable tool calling in the server options if it is not already enabled, and set the model's context length in the interface (see the next step).
The model name to enter is the identifier displayed by LM Studio in the list of loaded models, or returned by a request to http://localhost:1234/v1/models. It differs from the Ollama names.
#4. Configure the context window: the step everyone skips
This is the number-one cause of local Qwen Code failures. A coding agent sends a long system prompt on every turn (tool descriptions, behavior rules, and the contents of the QWEN.md file), followed by the session history and then the files it has read. Within the first few exchanges, you exceed 10,000 tokens. However, Ollama opens with a short context window by default (4,096 tokens in recent versions): everything beyond that is silently truncated, the model “forgets” the tool instructions and starts responding in prose instead of taking action, or loops on the same action.
You therefore need to enforce a context of at least 32,000 tokens. Two methods under Ollama: a global environment variable on the daemon, or a Modelfile that sets num_ctx for a specific model.
The second method is cleaner: it does not affect other models, and the derived model name reminds you of its setting. The cost is memory: the key-value cache grows with the context. For a 7B model in Q4, 32,000 context tokens add roughly 2 to 4 GB, depending on the architecture and cache quantization. If the model no longer fits on the GPU, Ollama offloads some layers to the CPU and performance collapses: monitor the PROCESSOR column in ollama ps, which should show 100% GPU.
On the Qwen Code side, there is also a session limit. The sessionTokenLimit setting in the settings file caps the cumulative number of tokens in a conversation; once it is reached, the tool prompts you to compress the history with /compress or start over with /clear. Set this value to match what your model actually supports: a limit of 32,000 for a model served with num_ctx 32768 prevents silent truncation on the Ollama side.
#5. Settings file and QWEN.md
Qwen Code reads a settings.json file at two levels: ~/.qwen/settings.json for the user, and .qwen/settings.json in the project, which takes precedence. The most useful keys for local use are the session limit, action approval mode, and MCP servers. The exact names have changed between versions; the example below follows the public documentation and should be checked against the Settings page for your version.
The QWEN.md file serves the same purpose as CLAUDE.md for Claude Code or AGENTS.md for other agents: it is the permanent memo injected into every session. Describe the stack, build and test commands, naming conventions, and what the agent must never touch. The /init command generates an initial version from the repository; /memory show displays what the agent actually loaded.
The approval mode controls what the agent can do without asking you. By default, every file write and shell command waits for your approval. The --approval-mode auto-edit option allows file changes but not commands; --yolo removes all confirmation prompts. With a local model that makes mistakes more often than a cloud model, keep the default mode until you trust it, and reserve --yolo for a clean, committed repository.
#6. First working session
A Qwen Code session is controlled in natural language, with a few shortcuts. The @ prefix inserts a file or folder into the request (@src/api/routes.py), the ! prefix runs a shell command without going through the model, and commands beginning with / control the tool itself.
- /help
- List of commands available in your version.
- /auth
- Changes the authentication method, useful for switching between local and cloud use.
- /model
- Displays or changes the current session's model.
- /stats
- Tokens consumed and session duration: the first thing to check when responses deteriorate.
- /compress
- Summarize the history to free up context without losing the thread.
- /clear
- Starts over with an empty conversation; QWEN.md remains loaded.
- /init et /memory
- Generates and then inspects the project context file.
- /mcp
- Status of configured MCP servers and the tools they expose.
- /quit
- Exits the session.
A workflow that works well with a local model: first ask for an explanation ("explain how authentication is handled in @src/auth/"), then a bounded change ("add a token expiration check in verify_token and a corresponding test"), and finally verification ("run make test and fix whatever breaks"). Each step fits within a few thousand tokens, and the model keeps track of the context. Requests like "refactor the entire module" exceed what 7B to 30B models can handle reliably.
For automation, non-interactive mode accepts a request as an argument and returns control when finished. It integrates into a script or Git hook.
#Local limitations of Qwen Code
Qwen Code is built around Qwen3-Coder models served by Alibaba Cloud, and that becomes apparent as soon as you run it on a smaller local model. Here is what you need to accept.
- Heavy system prompt
- The tool sends a long tool description on every turn. On a 7B model, this instruction alone takes up part of the model's context and attention, so it follows the tool-calling format less reliably than larger models. Loops and prose responses instead of actions occur more often than with OpenCode or Aider, which use more compact prompts.
- Vision limited to the cloud
- Image support (screenshots, mockups) relies on vision models served online. Locally, it works only if your server exposes a compatible multimodal model, which is not the case for most coding models.
- No native support for local models
- Unlike OpenCode, which lists the Ollama models in a menu, Qwen Code requires you to enter the model name and URL in a file or as an option. Changing models means editing the .env or relaunching with --model.
- Evolving configuration format
- The project is young, and its settings.json file has changed structure over successive versions. An example found on a forum may no longer be valid. The official documentation, as of the date you read it, is authoritative.
- Rewrite-based editing
- Like Gemini CLI, from which it is derived, Qwen Code modifies files by replacing blocks. Aider, on the other hand, applies unified diffs and automatically commits each change, making the history easier to read. If you want one commit per modification, Aider remains better suited.
In return, Qwen Code provides full MCP support, mature session-management commands inherited from Gemini CLI, a clean non-interactive mode, and an integration that extends to IDEs through an extension. It is a good fit if you already use Qwen models and want a single tool for switching between the Alibaba cloud and your GPU. If the goal is local use only, OpenCode or Aider require less configuration to achieve the same result. We do not publish a numerical comparison: quality depends primarily on the model you choose, not the agent.
#Troubleshooting
- The agent responds in text instead of executing actions
- Either the model does not support tool calling (check its Ollama entry), or the context is too short and the tool descriptions were truncated. Apply step 4 and verify with ollama ps that the model is loaded with the correct num_ctx.
- 404 error or “model not found”
- The name in OPENAI_MODEL does not exactly match ollama list. Copy and paste the name with its tag.
- Connection error on localhost:11434
- Ollama is not running, or is listening on another interface (OLLAMA_HOST). Test it with curl http://localhost:11434/v1/models.
- Very slow responses after a few exchanges
- The context has grown, and the model no longer fits on the GPU. ollama ps shows some CPU usage. Reduce num_ctx, switch to a smaller model, or run /compress earlier in the session.
- Qwen Code prompts for OAuth authentication again
- The environment variables are not being read: .env is not in the current directory or in ~/.qwen. Run /auth in the session and choose the OpenAI option, or pass the parameters on the command line to isolate the problem.
- The model ignores QWEN.md
- Use /memory show to verify that the file is loaded. If contextFileName was changed in settings.json, the name must match.
- npm installation fails with EACCES
- Insufficient permissions for the global npm directory. Install Node through nvm or fnm instead of the system package, then run the installation again.
#Go further
Qwen Code is just one of the terminal agents that accept a local server. The following guides cover the alternatives and model selection, which this article deliberately leaves aside.
- OpenCode + Ollama: a coding agent in your terminal
- Aider + Ollama: coding in the terminal with an agent that is 100% local
- Goose (Block): the local AI agent in your terminal
- Best local LLM for coding: Devstral, Qwen3-Coder
Guide written on October 11, 2026; the references are the GitHub repository and the Qwen Code documentation, along with the Ollama documentation. No speed or quality measurements were performed for this article; the memory figures are rough estimates. Commands and key names change with versions: verify them on the pages below before copying them.
- QwenLM/qwen-code GitHub repository (README, installation, license)
- Official Qwen Code documentation (authentication, settings, commands)
- Ollama documentation (OpenAI-compatible API, environment variables)
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.