Instructions to use eins78/Kolibri-1-mlx-mixed-4-8-bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use eins78/Kolibri-1-mlx-mixed-4-8-bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("eins78/Kolibri-1-mlx-mixed-4-8-bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use eins78/Kolibri-1-mlx-mixed-4-8-bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "eins78/Kolibri-1-mlx-mixed-4-8-bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "eins78/Kolibri-1-mlx-mixed-4-8-bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use eins78/Kolibri-1-mlx-mixed-4-8-bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "eins78/Kolibri-1-mlx-mixed-4-8-bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "eins78/Kolibri-1-mlx-mixed-4-8-bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "eins78/Kolibri-1-mlx-mixed-4-8-bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use eins78/Kolibri-1-mlx-mixed-4-8-bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "eins78/Kolibri-1-mlx-mixed-4-8-bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default eins78/Kolibri-1-mlx-mixed-4-8-bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use eins78/Kolibri-1-mlx-mixed-4-8-bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "eins78/Kolibri-1-mlx-mixed-4-8-bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "eins78/Kolibri-1-mlx-mixed-4-8-bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
eins78/Kolibri-1-mlx-mixed-4-8-bit
This model was converted to MLX format from Aleph-Alpha/Kolibri-1
(FP8 checkpoint, 128x128 block scales) with the converter in kolibri-mlx, in mlx-lm 0.32.0 format:
mixed affine quantisation, group size 64: routed experts (switch_mlp.*) 4-bit; attention projections, shared expert, embeddings and lm_head 8-bit; MoE router gates (mlp.gate) kept unquantised in bfloat16.
Refer to the original model card for details on the model.
Kolibri 1 is a mixture-of-experts model (50 layers, 384 experts, 6 routed plus 1 shared expert per token, sliding-window attention of 513 tokens in a 4:1 sliding:full pattern). The model is 42 GiB on disk and needs about 45.4 GB of unified memory, so it needs a Mac with 64 GB.
Use with mlx
Stock mlx-lm cannot load this model yet. mlx-lm does not include the kolibri1 architecture, so mlx_lm.load, mlx_lm.generate and mlx_lm.server fail on it with an unsupported model type. Until mlx-lm gains kolibri1 support, install kolibri-mlx and run import kolibri_mlx.register before loading; this registers the out-of-tree model file with mlx-lm in the current process.
git clone https://github.com/eins78/kolibri-mlx && cd kolibri-mlx
uv sync
import kolibri_mlx.register # makes mlx-lm aware of model_type "kolibri1"
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load("eins78/Kolibri-1-mlx-mixed-4-8-bit")
messages = [{"role": "user", "content": "Erkl盲re in zwei S盲tzen, warum der Himmel blau ist."}]
prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, reasoning_effort="low"
)
sampler = make_sampler(temp=1.0, top_p=0.97, top_k=128)
print(generate(model, tokenizer, prompt=prompt, sampler=sampler, max_tokens=400, verbose=True))
Command line (from the repo root):
uv run python generate.py --model eins78/Kolibri-1-mlx-mixed-4-8-bit -p "Hallo!" -m 200 --temp 1.0 --top-p 0.97 --top-k 128
uv run python serve.py --model eins78/Kolibri-1-mlx-mixed-4-8-bit --port 8080 # OpenAI-compatible server
Recommended sampling
temperature 1.0, top_p 0.97, top_k 128 (as recommended by Aleph Alpha). The mlx-lm server accepts all three in the request body.
Reasoning and tool calls
The chat template switches reasoning with reasoning_effort ("none", "low", "medium", "high") or enable_thinking=false.
Pass them as chat_template_kwargs to the server, or as keyword arguments to apply_chat_template:
curl http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
"messages": [{"role": "user", "content": "Give me three tips for clear commit messages."}],
"chat_template_kwargs": {"enable_thinking": false},
"temperature": 1.0, "top_p": 0.97, "top_k": 128, "max_tokens": 400
}'
With reasoning on, the model writes <think>...</think> first; the mlx-lm server returns it in message.reasoning.
Tool calls use Hermes-style <tool_call>{"name": ..., "arguments": ...}</tool_call> and are returned as tool_calls (mlx-lm json_tools parser, inferred from the chat template).
Verification
Compared with a layer-by-layer fp32 PyTorch CPU reference written from Aleph Alpha's vLLM plugin
(aleph_alpha_inference/kolibri1.py), on five German and English prompts, including two longer than the 513-token window.
Caveat: the reference is a torch rebuild of the plugin's math, checked line by line against the plugin source; the vLLM plugin itself was not run, because it needs CUDA and no CUDA machine was available. Dequantisation of the FP8 weights was validated separately against Aleph Alpha's BF16 release (below).
| Check | Result |
|---|---|
| Top-1 token agreement (this model vs reference), per prompt | 0.385 (13 tok), 0.867 (15), 0.958 (24), 0.966 (618), 0.892 (612); unquantised bf16 gives 0.769, 0.933, 1.000, 0.968, 0.931 |
| Top-5 overlap, per prompt | 0.754, 0.880, 0.925, 0.935, 0.875 |
| KL(reference, this model), mean per prompt | 0.526, 0.0146, 0.00996, 0.0174, 0.0690 (bf16 floor: 0.490, 0.0057, 0.00067, 0.0082, 0.036) |
| MLX implementation in fp32 vs reference, chained end to end | top-1 1.000 on all prompts, KL about 1e-11, max logit difference 2.3e-4 |
| Decode with cache vs full prefill | identical to 1e-6 in fp32 (tiny model, across the window boundary); bf16-level noise on the real model |
| Server smoke test (German + English, reasoning on/off, tool call) | passes |
| Speed / memory on Apple M4 Pro, 64 GB | 50 tokens/s generation, about 240 tokens/s prompt, 45.4 GB peak |
Reproduction: scripts/verify_layerwise.py, scripts/verify_e2e.py, scripts/smoke_server.py in kolibri-mlx; the full verification write-up is in its README.
The FP8 dequantisation was validated against the BF16 release (1362 FP8 tensors of shard 1, worst error 3.6 % of the tensor maximum, consistent with e4m3 rounding; bf16 tensors bit-identical). Quantisation costs accuracy; the numbers above are for this quantised model, not for the original weights. The 13-token prompt is noisy even in bf16.
Evaluation
Two of the converter author's own evaluation suites were run against this model through the mlx-lm server (temperature 0, reasoning off); the per-case results, baselines and timings are in the Evaluation section of the kolibri-mlx README.
| Suite | Result |
|---|---|
| MeteoSwiss forecast suite (German weather summaries from MCP tool output, 91 deterministic cases + 4 LLM-judged) | 78/91 deterministic (85.7 %), 4/4 judged |
| Hermes B1 tool-calling eval (8 agentic cases) | 3 of 8 cases run (two at 100/100, one hit the iteration cap); the rest were not run because the machine ran out of memory headroom during long prefills, see the README's Evaluation section |
Known gaps
- Requires the out-of-tree model file from kolibri-mlx; stock mlx-lm does not support
kolibri1yet. - KV-cache quantisation (
--kv-bits) does not work with the rotating sliding-window cache in mlx-lm. - Early positions of a context (the first 16 tokens) are the most affected by quantisation and bf16 noise.
AI assistance
The MLX port, the converter, the verification scripts, this model card and the conversion itself were produced by an AI coding agent (Claude Code, Claude Fable 5.1, with Opus and Sonnet subagents) working from written briefs, directed and reviewed by the repository owner (eins78). The verification numbers above were produced by the scripts in kolibri-mlx, not written by hand. The upload to this Hub repository was done by the agent with the owner's approval.
License
Apache License 2.0, inherited from Aleph-Alpha/Kolibri-1. The LICENSE file is included. Model weights and tokenizer are by Aleph Alpha; this repository only changes the format and precision. The converted weights stay under Apache License 2.0 from Aleph Alpha. The conversion code is in the kolibri-mlx repository under its own licences (CC0 1.0 by default, MIT and Apache 2.0 for derived files; see its NOTICE).
- Downloads last month
- 1,341
4-bit