Llamafile: a complete LLM in a single file executable
Llamafile (the mozilla-ai/llamafile project, Apache 2.0 license) packages a GGUF model and its llama.cpp engine into a single executable using Cosmopolitan Libc. You download one file, make it executable (or rename it to .exe on Windows), and a chat interface opens at http://localhost:8080, with no installation or daemon. The same file runs on Windows, macOS, Linux, and BSD, with a 4 GB limit on Windows.
Llamafile is a Mozilla project that compresses a complete language model and its inference engine into a single executable file. No installation, no dependencies, no daemon: download a file, launch it, and an LLM runs on your machine with a web interface. The same file works identically on Windows, macOS, and Linux. This guide shows how to launch a llamafile, what's inside it, and when it's a better choice than tools such as Ollama.
#Why llamafile
Most local LLM tools require an installation, a background service, and then a model download. Llamafile removes all these steps: the inference engine and model weights live in a single file. Make it executable, launch it, and your browser opens a chat interface. It is the shortest path from a download link to a conversation with a model running locally.
This format shines in three situations. For quickly testing a model without cluttering your machine. For distributing an LLM to colleagues or nontechnical users: one file to send, nothing to configure. And for archiving: a llamafile will still work years from now, without depending on a runtime to install or an online repository that remains accessible.
- Zero installation
- No packages, no daemons, no environment variables to configure.
- One file
- Engine and model combined, easy to copy, back up, or share.
- Multi-OS
- The same binary runs on Windows, macOS, Linux, and BSD (x86-64 and ARM64).
- 100% local
- No data leaves the machine, and no network connection is required after the download.
#How it works: the Mozilla format
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Llamafile combines two open-source components. The first is llama.cpp, the C/C++ inference engine that runs GGUF models on both CPU and GPU. The second is Cosmopolitan Libc, a library created by Justine Tunney that lets you compile a “polyglot” executable: a single binary recognized and launched natively by multiple operating systems.
In practice, a .llamafile is both a program and an archive. It contains the inference server code and, packaged inside it, the GGUF weights file. At launch, it detects the OS and architecture, loads the model from itself, then starts a local HTTP server with a web chat interface.
The llamafile project itself is licensed under Apache 2.0; the modifications made to llama.cpp and whisper.cpp for the project remain under the MIT license, like the original projects, to remain compatible and suitable for upstream integration. Since version 0.10.0, llamafile has used a new build system aligned with recent versions of llama.cpp, expanding support for newer models and features; older versions remain available for those who prefer the “classic” experience.
The project also includes whisperfile, a single-file companion built on whisper.cpp and the same Cosmopolitan packaging, for audio transcription and translation without installation, on the same platforms as a standard llamafile.
#Prerequisites
The requirements are deliberately minimal. The key is having enough memory for the selected model; the file size and required RAM depend directly on the quantization.
- OS
- Windows 10+, macOS 11+, recent Linux or BSD. x86-64 or ARM64 (Apple Silicon included).
- RAM
- Roughly count the file size plus some headroom. A 7B model in Q4 (~5 GB) runs comfortably on 8 to 16 GB.
- GPU (optional)
- Acceleration is possible via NVIDIA (CUDA) or Apple Metal. Without a GPU, inference runs on the CPU—slower but functional.
- Disk space
- From ~500 MB for a small model to several dozen GB for a large, unquantized model.
#Download and launch in a few seconds
The official mozilla-ai/llamafile repository on GitHub lists ready-to-use llamafiles, and Hugging Face hosts many more. The project, initially launched by Mozilla Builders, is now taken over and maintained by Mozilla.ai. Once you have the file, the procedure differs only slightly by OS.
- 01Download the fileDownload a .llamafile from the project’s GitHub page or from Hugging Face (search for “llamafile” in the model search bar).
- 02macOS and Linux: make executableOpen a terminal in the download folder and allow execution with chmod, then run the file.
- 03Windows: Rename to .exeAdd the .exe extension to the filename, then double-click it (or launch it from PowerShell).
- 04Open the interfaceThe program starts a local server and usually opens your browser automatically. If not, go to http://localhost:8080.
#The same file on Windows, Mac, and Linux
That is llamafile’s most surprising promise: the file you download on Linux is exactly the same one that will run on a Mac or Windows PC. There is no “Windows,” “Mac,” or “Linux” version to choose. This portability comes from Cosmopolitan Libc, which produces an executable understood by several operating systems at once.
In practice, this means the same llamafile placed on a USB drive or network share works for an entire team, regardless of their workstations. You distribute a model without worrying about everyone's OS or asking anyone to install anything.
#Useful command-line options
The web interface is enough for chatting, but a few options let you adjust its behavior. Add them after the filename when launching.
- -ngl N
- Offloads N model layers to the GPU (for example, -ngl 999 to put everything in VRAM if there’s enough).
- -c N
- Sets the context window size in tokens (e.g. -c 4096).
- --host / --port
- Changes the server's listening address and port (localhost:8080 by default).
- -m file.gguf
- Uses an external weights file instead of the bundled one — useful for bypassing Windows' 4 GB limit.
- --server --nobrowser
- Starts in server mode without opening a browser, which is convenient for headless use.
The server also exposes an OpenAI-compatible API at /v1/chat/completions. You can therefore connect a llamafile behind any client that speaks the OpenAI protocol by simply pointing the base URL to http://localhost:8080/v1.
#Create your own llamafile
You are not limited to files prepared by others: any GGUF model can be packaged. The project provides a zipalign binary and an “empty” executable (the engine alone), to which you add the weight file and a default arguments file.
- 01Get the toolsDownload the latest release from GitHub: it contains the llamafile binary (engine only) and the zipalign utility.
- 02Have a GGUFGet the .gguf file for the model you want from Hugging Face, using the quantization of your choice (Q4_K_M is a good compromise).
- 03Write the default argumentsCreate a .args file listing the options to apply at launch (model, web interface, etc.).
- 04Package with zipalignCopy the engine under a new name, then inject the GGUF and .args file into it with zipalign.
- 05TestMake the result executable and run it like any other llamafile.
#Llamafile vs. Ollama: two philosophies
Llamafile and Ollama address the same need — running an LLM locally — but use opposite approaches. Ollama installs a daemon that listens on http://localhost:11434 and manages a library of models downloaded on demand. Llamafile manages nothing: each model is a standalone file that you launch directly.
| Criterion | Llamafile | Ollama |
|---|---|---|
| Mental model | One file = one model, no service | Model manager with a central daemon |
| Installation | None, the file runs as is | Install once, then “pull” the models |
| Distribution | Handed over as-is, runs anywhere | Assumes the recipient installs Ollama |
| Several models | Quickly becomes unwieldy, one file per model | Excels at instant swapping via a command |
| Audio (STT) | Whisperfile as a separate companion | Non-native |
| OpenAI-compatible API | Yes | Yes |
| License | Apache 2.0 (MIT for reused components) | MIT |
| A good choice if | Test or distribute without installing | Juggle multiple models every day |
The split is simple: llamafile wins for occasional testing, distribution to nontechnical users, or archiving a fixed model; Ollama wins as soon as you juggle multiple models daily or integrate an LLM into a durable stack with an interface such as Open WebUI or LM Studio.
#Troubleshooting
- “run-detectors” or launch refusal (Linux)
- An overzealous binutils/binfmt can block the polyglot format. Solution: install the APE loader package, or run it via sh -c "./fichier.llamafile".
- File too large on Windows
- Beyond 4 GB, keep the GGUF separate and pass it with -m to a small “engine-only” llamafile.
- No GPU acceleration
- Verify that the CUDA drivers (NVIDIA) are present and add -ngl 999. On Mac, Metal is used automatically.
- Port already in use
- Another service is running on 8080: change it with --port 8090, for example.
- Slow responses
- On CPU alone, this is expected for large models. Choose a lighter quantization or a smaller model (3B rather than 7B).
#Go further
Llamafile is an ideal entry point. For more complete and sustainable local use, these guides take the topic further.
- Install an LLM locally: the step-by-step guide
- Ollama vs llama.cpp: which one to choose
- Alternatives to Ollama: the overview
- Install Ollama in 5 minutes
- Source: llamafile GitHub repository
- Source: official llamafile documentation
- Source: llamafile's original announcement post
#FAQ
Is Llamafile free?+
Who develops llamafile?+
Can Llamafile also transcribe audio?+
Why won't my llamafile launch on Windows?+
Do you need to install llama.cpp to use llamafile?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.