Beginner 11 minOther tools

Llamafile: a complete LLM in a single file executable

Direct response

Llamafile (the mozilla-ai/llamafile project, Apache 2.0 license) packages a GGUF model and its llama.cpp engine into a single executable using Cosmopolitan Libc. You download one file, make it executable (or rename it to .exe on Windows), and a chat interface opens at http://localhost:8080, with no installation or daemon. The same file runs on Windows, macOS, Linux, and BSD, with a 4 GB limit on Windows.

Llamafile is a Mozilla project that compresses a complete language model and its inference engine into a single executable file. No installation, no dependencies, no daemon: download a file, launch it, and an LLM runs on your machine with a web interface. The same file works identically on Windows, macOS, and Linux. This guide shows how to launch a llamafile, what's inside it, and when it's a better choice than tools such as Ollama.

By Clara M.·Update 2026-09-28·Tested on Windows, macOS, and Linux

#Why llamafile

Most local LLM tools require an installation, a background service, and then a model download. Llamafile removes all these steps: the inference engine and model weights live in a single file. Make it executable, launch it, and your browser opens a chat interface. It is the shortest path from a download link to a conversation with a model running locally.

This format shines in three situations. For quickly testing a model without cluttering your machine. For distributing an LLM to colleagues or nontechnical users: one file to send, nothing to configure. And for archiving: a llamafile will still work years from now, without depending on a runtime to install or an online repository that remains accessible.

Zero installation
No packages, no daemons, no environment variables to configure.
One file
Engine and model combined, easy to copy, back up, or share.
Multi-OS
The same binary runs on Windows, macOS, Linux, and BSD (x86-64 and ARM64).
100% local
No data leaves the machine, and no network connection is required after the download.

#How it works: the Mozilla format

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Llamafile combines two open-source components. The first is llama.cpp, the C/C++ inference engine that runs GGUF models on both CPU and GPU. The second is Cosmopolitan Libc, a library created by Justine Tunney that lets you compile a “polyglot” executable: a single binary recognized and launched natively by multiple operating systems.

In practice, a .llamafile is both a program and an archive. It contains the inference server code and, packaged inside it, the GGUF weights file. At launch, it detects the OS and architecture, loads the model from itself, then starts a local HTTP server with a web chat interface.

i
GGUF, the common foundation
The weights embedded in a llamafile use the GGUF format, the same one used by llama.cpp, Ollama or LM Studio. A llamafile does not invent a new model format: it packages an existing GGUF with its runtime.

The llamafile project itself is licensed under Apache 2.0; the modifications made to llama.cpp and whisper.cpp for the project remain under the MIT license, like the original projects, to remain compatible and suitable for upstream integration. Since version 0.10.0, llamafile has used a new build system aligned with recent versions of llama.cpp, expanding support for newer models and features; older versions remain available for those who prefer the “classic” experience.

The project also includes whisperfile, a single-file companion built on whisper.cpp and the same Cosmopolitan packaging, for audio transcription and translation without installation, on the same platforms as a standard llamafile.

#Prerequisites

The requirements are deliberately minimal. The key is having enough memory for the selected model; the file size and required RAM depend directly on the quantization.

OS
Windows 10+, macOS 11+, recent Linux or BSD. x86-64 or ARM64 (Apple Silicon included).
RAM
Roughly count the file size plus some headroom. A 7B model in Q4 (~5 GB) runs comfortably on 8 to 16 GB.
GPU (optional)
Acceleration is possible via NVIDIA (CUDA) or Apple Metal. Without a GPU, inference runs on the CPU—slower but functional.
Disk space
From ~500 MB for a small model to several dozen GB for a large, unquantized model.
!
The 4 GB limit on Windows
Windows caps executable size at 4 GB. Beyond that, you must either choose a smaller llamafile or keep the .gguf weight file separate and pass it to llamafile with the -m option. Other systems are not affected.

#Download and launch in a few seconds

The official mozilla-ai/llamafile repository on GitHub lists ready-to-use llamafiles, and Hugging Face hosts many more. The project, initially launched by Mozilla Builders, is now taken over and maintained by Mozilla.ai. Once you have the file, the procedure differs only slightly by OS.

  1. 01
    Download the file
    Download a .llamafile from the project’s GitHub page or from Hugging Face (search for “llamafile” in the model search bar).
  2. 02
    macOS and Linux: make executable
    Open a terminal in the download folder and allow execution with chmod, then run the file.
  3. 03
    Windows: Rename to .exe
    Add the .exe extension to the filename, then double-click it (or launch it from PowerShell).
  4. 04
    Open the interface
    The program starts a local server and usually opens your browser automatically. If not, go to http://localhost:8080.
Terminal — macOS / Linux
# Rendre le fichier exécutable
chmod +x Llama-3.2-3B-Instruct.Q4_K_M.llamafile

# Lancer : ouvre l'interface web sur http://localhost:8080
./Llama-3.2-3B-Instruct.Q4_K_M.llamafile
PowerShell — Windows
# Renommer pour ajouter l'extension .exe
Rename-Item .\Llama-3.2-3B-Instruct.Q4_K_M.llamafile Llama-3.2-3B.exe

# Lancer
.\Llama-3.2-3B.exe
→
Double-clicking
On Windows after renaming it to .exe, and on macOS/Linux once the executable bit is set, a simple double-click in the file explorer is enough to start llamafile. No terminal is needed for normal use.

#The same file on Windows, Mac, and Linux

That is llamafile’s most surprising promise: the file you download on Linux is exactly the same one that will run on a Mac or Windows PC. There is no “Windows,” “Mac,” or “Linux” version to choose. This portability comes from Cosmopolitan Libc, which produces an executable understood by several operating systems at once.

In practice, this means the same llamafile placed on a USB drive or network share works for an entire team, regardless of their workstations. You distribute a model without worrying about everyone's OS or asking anyone to install anything.

i
ARM and Apple Silicon
Recent llamafiles also include ARM64 code. On an M1/M2/M3/M4 Mac, inference uses Metal and runs natively, without Rosetta emulation.

#Useful command-line options

The web interface is enough for chatting, but a few options let you adjust its behavior. Add them after the filename when launching.

-ngl N
Offloads N model layers to the GPU (for example, -ngl 999 to put everything in VRAM if there’s enough).
-c N
Sets the context window size in tokens (e.g. -c 4096).
--host / --port
Changes the server's listening address and port (localhost:8080 by default).
-m file.gguf
Uses an external weights file instead of the bundled one — useful for bypassing Windows' 4 GB limit.
--server --nobrowser
Starts in server mode without opening a browser, which is convenient for headless use.
Terminal — options
# Tout charger sur le GPU, contexte 8k, port personnalisé
./Llama-3.2-3B-Instruct.Q4_K_M.llamafile -ngl 999 -c 8192 --port 8090

# Serveur headless avec un GGUF externe (contourne la limite 4 Go)
./llava-v1.5-7b-q4.llamafile --server --nobrowser -m modele.gguf

The server also exposes an OpenAI-compatible API at /v1/chat/completions. You can therefore connect a llamafile behind any client that speaks the OpenAI protocol by simply pointing the base URL to http://localhost:8080/v1.

#Create your own llamafile

You are not limited to files prepared by others: any GGUF model can be packaged. The project provides a zipalign binary and an “empty” executable (the engine alone), to which you add the weight file and a default arguments file.

  1. 01
    Get the tools
    Download the latest release from GitHub: it contains the llamafile binary (engine only) and the zipalign utility.
  2. 02
    Have a GGUF
    Get the .gguf file for the model you want from Hugging Face, using the quantization of your choice (Q4_K_M is a good compromise).
  3. 03
    Write the default arguments
    Create a .args file listing the options to apply at launch (model, web interface, etc.).
  4. 04
    Package with zipalign
    Copy the engine under a new name, then inject the GGUF and .args file into it with zipalign.
  5. 05
    Test
    Make the result executable and run it like any other llamafile.
Terminal — packaging
# Fichier d'arguments par défaut
cat > .args <<'EOF'
-m
modele.Q4_K_M.gguf
--host
0.0.0.0
...
EOF

# Copier le moteur, puis y injecter poids + args
cp llamafile mon-modele.llamafile
zipalign -j0 mon-modele.llamafile modele.Q4_K_M.gguf .args

# Tester
chmod +x mon-modele.llamafile
./mon-modele.llamafile
→
Choose the right quantization
For a llamafile intended for sharing, Q4_K_M offers the best size/quality balance. Reserve Q5_K_M or Q8_0 for cases where quality matters more than file size, and use FP16 only for lossless archiving.

#Llamafile vs. Ollama: two philosophies

Llamafile and Ollama address the same need — running an LLM locally — but use opposite approaches. Ollama installs a daemon that listens on http://localhost:11434 and manages a library of models downloaded on demand. Llamafile manages nothing: each model is a standalone file that you launch directly.

Llamafile or Ollama, a quick comparison
CriterionLlamafileOllama
Mental modelOne file = one model, no serviceModel manager with a central daemon
InstallationNone, the file runs as isInstall once, then “pull” the models
DistributionHanded over as-is, runs anywhereAssumes the recipient installs Ollama
Several modelsQuickly becomes unwieldy, one file per modelExcels at instant swapping via a command
Audio (STT)Whisperfile as a separate companionNon-native
OpenAI-compatible APIYesYes
LicenseApache 2.0 (MIT for reused components)MIT
A good choice ifTest or distribute without installingJuggle multiple models every day

The split is simple: llamafile wins for occasional testing, distribution to nontechnical users, or archiving a fixed model; Ollama wins as soon as you juggle multiple models daily or integrate an LLM into a durable stack with an interface such as Open WebUI or LM Studio.

i
They are not rivals
Nothing prevents using both: a llamafile to show a model to a client on their laptop, Ollama as the permanent engine on their workstation. Both are also based on llama.cpp.

#Troubleshooting

“run-detectors” or launch refusal (Linux)
An overzealous binutils/binfmt can block the polyglot format. Solution: install the APE loader package, or run it via sh -c "./fichier.llamafile".
File too large on Windows
Beyond 4 GB, keep the GGUF separate and pass it with -m to a small “engine-only” llamafile.
No GPU acceleration
Verify that the CUDA drivers (NVIDIA) are present and add -ngl 999. On Mac, Metal is used automatically.
Port already in use
Another service is running on 8080: change it with --port 8090, for example.
Slow responses
On CPU alone, this is expected for large models. Choose a lighter quantization or a smaller model (3B rather than 7B).

#Go further

Llamafile is an ideal entry point. For more complete and sustainable local use, these guides take the topic further.


#FAQ

Is Llamafile free?+
Yes. The llamafile project is open source under the Apache 2.0 license, and modifications to llama.cpp and whisper.cpp remain under the MIT license. You only pay for the hardware and electricity that run the model, as with any local inference tool you run yourself.
Who develops llamafile?+
The project was launched by Mozilla Builders in 2023, then taken over and modernized by Mozilla.ai. The official repository consequently changed its GitHub address: it now lives under the mozilla-ai organization rather than the former Mozilla-Ocho, with a new build system since version 0.10.0.
Can Llamafile also transcribe audio?+
Yes, via whisperfile, a single-file companion based on whisper.cpp and the same Cosmopolitan packaging. It handles audio transcription and translation without installation, on the same platforms as a standard chat llamafile, using exactly the same principle of a single, self-contained file to download.
Why won't my llamafile launch on Windows?+
The most common cause is the 4 GB limit imposed on Windows executables: a larger llamafile cannot run as-is. The solution is to keep the GGUF weights file separate and pass it with the -m option to a lighter “engine-only” llamafile.
Do you need to install llama.cpp to use llamafile?+
No. Llamafile already embeds the llama.cpp engine compiled into the file using Cosmopolitan Libc. That is precisely what makes the format self-contained: there are no external dependencies to install before you can run the model, and you do not even need to know that llama.cpp exists beforehand.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.