eins78/Kolibri-1-mlx-mixed-4-8-bit

This model was converted to MLX format from Aleph-Alpha/Kolibri-1 (FP8 checkpoint, 128x128 block scales) with the converter in kolibri-mlx, in mlx-lm 0.32.0 format: mixed affine quantisation, group size 64: routed experts (switch_mlp.*) 4-bit; attention projections, shared expert, embeddings and lm_head 8-bit; MoE router gates (mlp.gate) kept unquantised in bfloat16. Refer to the original model card for details on the model.

Kolibri 1 is a mixture-of-experts model (50 layers, 384 experts, 6 routed plus 1 shared expert per token, sliding-window attention of 513 tokens in a 4:1 sliding:full pattern). The model is 42 GiB on disk and needs about 45.4 GB of unified memory, so it needs a Mac with 64 GB.

Use with mlx

Stock mlx-lm cannot load this model yet. mlx-lm does not include the kolibri1 architecture, so mlx_lm.load, mlx_lm.generate and mlx_lm.server fail on it with an unsupported model type. Until mlx-lm gains kolibri1 support, install kolibri-mlx and run import kolibri_mlx.register before loading; this registers the out-of-tree model file with mlx-lm in the current process.

git clone https://github.com/eins78/kolibri-mlx && cd kolibri-mlx
uv sync
import kolibri_mlx.register  # makes mlx-lm aware of model_type "kolibri1"
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load("eins78/Kolibri-1-mlx-mixed-4-8-bit")

messages = [{"role": "user", "content": "Erkl盲re in zwei S盲tzen, warum der Himmel blau ist."}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, reasoning_effort="low"
)

sampler = make_sampler(temp=1.0, top_p=0.97, top_k=128)
print(generate(model, tokenizer, prompt=prompt, sampler=sampler, max_tokens=400, verbose=True))

Command line (from the repo root):

uv run python generate.py --model eins78/Kolibri-1-mlx-mixed-4-8-bit -p "Hallo!" -m 200 --temp 1.0 --top-p 0.97 --top-k 128
uv run python serve.py --model eins78/Kolibri-1-mlx-mixed-4-8-bit --port 8080   # OpenAI-compatible server

Recommended sampling

temperature 1.0, top_p 0.97, top_k 128 (as recommended by Aleph Alpha). The mlx-lm server accepts all three in the request body.

Reasoning and tool calls

The chat template switches reasoning with reasoning_effort ("none", "low", "medium", "high") or enable_thinking=false. Pass them as chat_template_kwargs to the server, or as keyword arguments to apply_chat_template:

curl http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "messages": [{"role": "user", "content": "Give me three tips for clear commit messages."}],
  "chat_template_kwargs": {"enable_thinking": false},
  "temperature": 1.0, "top_p": 0.97, "top_k": 128, "max_tokens": 400
}'

With reasoning on, the model writes <think>...</think> first; the mlx-lm server returns it in message.reasoning. Tool calls use Hermes-style <tool_call>{"name": ..., "arguments": ...}</tool_call> and are returned as tool_calls (mlx-lm json_tools parser, inferred from the chat template).

Verification

Compared with a layer-by-layer fp32 PyTorch CPU reference written from Aleph Alpha's vLLM plugin (aleph_alpha_inference/kolibri1.py), on five German and English prompts, including two longer than the 513-token window. Caveat: the reference is a torch rebuild of the plugin's math, checked line by line against the plugin source; the vLLM plugin itself was not run, because it needs CUDA and no CUDA machine was available. Dequantisation of the FP8 weights was validated separately against Aleph Alpha's BF16 release (below).

Check Result
Top-1 token agreement (this model vs reference), per prompt 0.385 (13 tok), 0.867 (15), 0.958 (24), 0.966 (618), 0.892 (612); unquantised bf16 gives 0.769, 0.933, 1.000, 0.968, 0.931
Top-5 overlap, per prompt 0.754, 0.880, 0.925, 0.935, 0.875
KL(reference, this model), mean per prompt 0.526, 0.0146, 0.00996, 0.0174, 0.0690 (bf16 floor: 0.490, 0.0057, 0.00067, 0.0082, 0.036)
MLX implementation in fp32 vs reference, chained end to end top-1 1.000 on all prompts, KL about 1e-11, max logit difference 2.3e-4
Decode with cache vs full prefill identical to 1e-6 in fp32 (tiny model, across the window boundary); bf16-level noise on the real model
Server smoke test (German + English, reasoning on/off, tool call) passes
Speed / memory on Apple M4 Pro, 64 GB 50 tokens/s generation, about 240 tokens/s prompt, 45.4 GB peak

Reproduction: scripts/verify_layerwise.py, scripts/verify_e2e.py, scripts/smoke_server.py in kolibri-mlx; the full verification write-up is in its README. The FP8 dequantisation was validated against the BF16 release (1362 FP8 tensors of shard 1, worst error 3.6 % of the tensor maximum, consistent with e4m3 rounding; bf16 tensors bit-identical). Quantisation costs accuracy; the numbers above are for this quantised model, not for the original weights. The 13-token prompt is noisy even in bf16.

Evaluation

Two of the converter author's own evaluation suites were run against this model through the mlx-lm server (temperature 0, reasoning off); the per-case results, baselines and timings are in the Evaluation section of the kolibri-mlx README.

Suite Result
MeteoSwiss forecast suite (German weather summaries from MCP tool output, 91 deterministic cases + 4 LLM-judged) 78/91 deterministic (85.7 %), 4/4 judged
Hermes B1 tool-calling eval (8 agentic cases) 3 of 8 cases run (two at 100/100, one hit the iteration cap); the rest were not run because the machine ran out of memory headroom during long prefills, see the README's Evaluation section

Known gaps

  • Requires the out-of-tree model file from kolibri-mlx; stock mlx-lm does not support kolibri1 yet.
  • KV-cache quantisation (--kv-bits) does not work with the rotating sliding-window cache in mlx-lm.
  • Early positions of a context (the first 16 tokens) are the most affected by quantisation and bf16 noise.

AI assistance

The MLX port, the converter, the verification scripts, this model card and the conversion itself were produced by an AI coding agent (Claude Code, Claude Fable 5.1, with Opus and Sonnet subagents) working from written briefs, directed and reviewed by the repository owner (eins78). The verification numbers above were produced by the scripts in kolibri-mlx, not written by hand. The upload to this Hub repository was done by the agent with the owner's approval.

License

Apache License 2.0, inherited from Aleph-Alpha/Kolibri-1. The LICENSE file is included. Model weights and tokenizer are by Aleph Alpha; this repository only changes the format and precision. The converted weights stay under Apache License 2.0 from Aleph Alpha. The conversion code is in the kolibri-mlx repository under its own licences (CC0 1.0 by default, MIT and Apache 2.0 for derived files; see its NOTICE).

Downloads last month
1,341
Safetensors
Model size
78B params
Tensor type
U32
路
BF16
路
F32
路
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for eins78/Kolibri-1-mlx-mixed-4-8-bit

Quantized
(18)
this model