Using Open source AI in my local machine
Open source AI is the gold standard for public AI thought as public good. Apertus is an open source AI developed by the Swiss AI Initiative. The question is: how does one use it in your own machine?
The LLM Apertus is published in Swiss AI Initiative hugging face profile. Until today, its latest version is 1.5 and comes in two sizes (8B and 70B). 70B is hard to fit in consumer hardware (does not fit in mine), so we will go for Apertus v1.5 8B.
As it stands, the recommended setup to perform inference is using vLLM. Unfortunately, vLLM is designed for large scale deployment, not single usage on a laptop, so it is not the best deployment tool for us. Instead, the best tool for personal local usage is llama.cpp.
Moreover, Apertus v1.5 is multimodal, accepting text, audio, and visual input. For now, we simply want to chat with it, so text only. Therefore, we will use a stripped down version Apertus v1.5 8B for text-only. Lastly, to perform inference with llama.cpp instead of vLLM, we need the model in the GGUF format (instead of the published safetensors format). Therefore, we will use andreasmartin/apertus-v1.5-8b-text-Q8_0-GGUF.
Here is the simplest interface that works for me:
- Download apertus-v1.5-8b-text-q8_0.gguf
- Install llama.cpp. For example by running
brew install llama.cpp. - Run a server performing inference on the model by running
llama-server --alias apertus-1.5-8b --jinja --ctx-size 16384 --parallel 1 --threads 8 --threads-batch 16 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --cache-reuse 256 --temp 0.8 --top-p 0.9 --host 0.0.0.0 --port 8080 --model path/to/apertus-v1.5-8b-text-q8_0.ggufThis is using a few options explained below.
| Option | What it does | Default | Your value |
|---|---|---|---|
--model |
Path to the GGUF file | none (required) | the Apertus Q8_0 file |
--alias |
Model name shown in the UI and returned by /v1/models |
derived from the model path | apertus-1.5-8b |
--jinja |
Formats chat messages with the Jinja chat template embedded in the GGUF | on in recent builds, off in older ones | on (safe to keep either way) |
--ctx-size |
Context window in tokens; sets how much KV-cache memory is reserved | older builds: 4096; newer builds: 0, meaning the model’s full training context (262K here) | 16384 |
--parallel |
Number of simultaneous request slots, which share the context | auto (−1) in recent builds, often 4; 1 in older builds | 1, so one conversation gets all 16K |
--threads |
CPU threads for token generation | auto, roughly the physical core count | 8 |
--threads-batch |
CPU threads for prompt processing | same as --threads |
16 |
--flash-attn |
Flash attention: faster, uses less memory, and is required for a quantized V cache | auto |
on |
--cache-type-k |
Precision of the K half of the KV cache | f16 |
q8_0 (half the memory) |
--cache-type-v |
Precision of the V half of the KV cache | f16 |
q8_0 |
--cache-reuse |
Minimum matching chunk size (in tokens) for reusing cached prompt parts via KV shifting; speeds up edits and regenerations | 0 (off) | 256 |
--temp |
Sampling temperature | 0.8 | 0.8 (already the default) |
--top-p |
Nucleus sampling cutoff | 0.95 | 0.9 (Swiss AI’s recommendation) |
--host |
Address the server listens on | 127.0.0.1 (localhost only) |
0.0.0.0 (all interfaces, needed for Tailscale) |
--port |
HTTP port | 8080 | 8080 (already the default) |
Warning: If, for some reason, your llama.cpp installation does not include an UI, then you have to do the following.
- Run this to download the llama UI
mkdir path/to/llama-ui curl -L https://huggingface.co/buckets/ggml-org/llama-ui/resolve/latest/dist.tar.gz | tar -xz -C path/to/llama-ui - Add
--path ~/.local/share/llama-uito the llama-server call