KOLIBRI-1-DERISKED-FP8
Block-FP8 release of the behaviorally modified Kolibri 1 derivative for direct, high-agency reasoning, long-context, and tool-enabled workflows.
Release status: public and ungated. The standalone FP8 checkpoint, configuration, tokenizer, model card, account banner, and deployment kit are included.
Artifact contents
| Component | Status |
|---|---|
| Model card | Included |
| Blackfrost-AI account banner | Included |
| Apache-2.0 license | Included |
| TP4 deployment kit | Included under DEPLOYMENT/ |
| Standalone FP8 weights | Included: 32 SafeTensors shards |
| Model configuration and tokenizer assets | Included |
This is a standalone FP8 checkpoint. It is not an adapter and does not require the parent checkpoint at inference time.
Why this release exists
Kolibri 1 is an efficient bilingual mixture-of-experts model with explicit reasoning, long-context support, and structured tool calling. This derivative provides an alternative behavioral operating point intended for users who want a more direct assistant while preserving the architecture's reasoning, retrieval, multi-turn, and tool-use strengths.
The weights are behaviorally modified. The proprietary transformation process and reconstruction materials are not distributed.
Model specifications
| Property | Value |
|---|---|
| Architecture | Kolibri1ForCausalLM, sparse mixture of experts |
| Precision | FP8 E4M3 weights with 128×128 FP32 block scales and dynamic activations; router gates remain BF16 |
| Total parameters | 78,103,074,560 |
| Active parameters per token | 3,457,573,120 |
| Transformer blocks | 50 |
| Experts per MoE block | 384 routed experts plus one shared expert |
| Routed experts per token | 6 |
| Attention pattern | 4:1 sliding-window to full-attention pattern |
| Languages | German and English |
| Reasoning | Explicit reasoning mode with configurable effort |
| Tool use | Structured tool calling through the Kolibri parser |
| Native model context | 262,144 tokens |
| Blackfrost validated context | 262,144 tokens |
| Upstream extended-context claim | Up to 1,048,576 tokens; not independently qualified on this derivative |
| Checkpoint tensor footprint | 78.83 GB |
| Recommended runtime | Aleph Alpha's official vLLM inference plugin/container |
Upstream architecture and training details are documented in the Kolibri model card and technical report.
Lineage
- Foundational and immediate parent:
Aleph-Alpha/Kolibri-1-BF16 - Immutable parent revision:
8c8b34899c7bc77b349f7916b07158e3ba29456a - Tokenizer and chat-template lineage: parent Kolibri 1 release
- Blackfrost stage: behaviorally modified BF16 derivative; proprietary transformation process not distributed
- Quantization stage: official-layout 128×128 block-FP8 export from the completed derivative
- Public artifact: standalone FP8 SafeTensors checkpoint
No upstream benchmark number is presented here as a measurement of this derivative. Consult the parent card for upstream evaluations and training-data documentation.
Artifact verification
The FP8 checkpoint completed structural and numerical verification before publication:
- all 32 expected checkpoint shards completed atomically;
- 57,950 weights follow Aleph Alpha's official Kolibri FP8 quantization surface;
- attention, routed experts, and shared experts use FP8 E4M3 with 128×128 FP32 block scales;
- 403 router-gate and other excluded tensors remain exactly equal to the derivative BF16 source;
- 185 distributed reconstruction samples passed, with minimum cosine similarity
0.9995573and maximum relative L2 error0.0265193; - the artifact contains 116,303 tensors totaling 78,827,029,120 tensor bytes;
- the source checkpoint remains immutable throughout conversion.
Parent fingerprints used during construction:
| Item | SHA-256 |
|---|---|
Parent config.json |
2577f567b37487b1ee24ccc9f7e9d9895b01ce6d8ae248f481a41130b04be256 |
Parent model.safetensors.index.json |
370e3721203856717b806a3c1c7f07e5b149dc3e534096f047f6de9e34977c0d |
The complete public checkpoint checksum inventory will be added with the weight artifact. Deployment-kit checksums are included in DEPLOYMENT/CHECKSUMS.sha256.
Behavioral and capability evaluation
The behavioral predecessor to this quant completed two independent long-form and tool-integrity harness passes, each satisfying 74 of 75 assertions. These figures describe the selected BF16 derivative before FP8 export; a full post-quant behavioral rerun remains pending.
| Test area | Measured result |
|---|---|
| Independent full harness passes | 2 |
| Assertions passed | 74/75 in each pass |
| Long-form section completion | 9/9 sections plus exact end marker in both passes |
| Distant-context retrieval | Exact requested records from a 1,000-record context in both passes |
| Multi-turn memory | Exact requested state across five turns in both passes |
| Heterogeneous two-stage tool workflows | 8/8 passed across both passes |
| Repeated and parallel weather workflows | 8/8 passed across both passes |
| Malformed tool names or arguments | 0 observed |
| Textual tool-call leakage | 0 observed |
| Output or tool-call loops | 0 observed |
The only failed assertion in each pass was a strict request for at least 1,000 visible words. Both responses were coherent, completed every requested section, and emitted the required end marker, but were shorter than the threshold.
These are focused integrity checks, not a comprehensive capability benchmark. A quantitative refusal-behavior evaluation, broad academic benchmark suite, multilingual regression suite, and extended-context qualification beyond 262,144 tokens remain pending.
Operator review observed the intended softer, more direct behavior when the checkpoint was used with an operator-supplied system prompt. That observation is qualitative: the automated harness did not measure refusal rate or isolate the contribution of the system prompt from the checkpoint.
Prompt, reasoning, and sampling notes
Use the packaged Kolibri chat template. It supports:
reasoning_effort:low,medium,high, ornone;enable_thinking:trueorfalse;- structured tools through the standard OpenAI-compatible
toolsfield; - multi-step tool-result reinjection.
The upstream recommended sampling defaults are:
| Parameter | Value |
|---|---|
temperature |
1.0 |
top_p |
0.97 |
top_k |
128 |
The Blackfrost integrity harness used sampling rather than greedy decoding. Applications should validate their own system prompts, sampling settings, structured-output schemas, and tool policies. This release does not prescribe a default persona or application policy.
Deployment
The included DEPLOYMENT/ kit provides the intended serving shape with:
- the official Aleph Alpha Kolibri inference container;
- vLLM's OpenAI-compatible API;
- FP8 weights with BF16 activations;
- tensor parallelism across four GPUs;
- a 262,144-token context ceiling;
- Kolibri reasoning and tool-call parsers;
- automatic tool choice and prefix caching.
Start with DEPLOYMENT/README.md. The launcher publishes the clean API model identifier KOLIBRI-1-DERISKED-FP8.
Limitations
- This FP8 export has not yet received a comprehensive post-quant academic, coding, German-language, safety, or refusal-behavior evaluation.
- The 1,048,576-token upstream serving mode has not been independently qualified on this derivative; the published deployment kit stays at the validated 262,144-token profile.
- The strict 1,000-visible-word requirement failed in both focused harness runs despite complete section coverage.
- Tool calls are model-generated data. Validate function names, schemas, arguments, authorization, and tool outputs before execution.
- Mixture-of-experts inference has a large memory footprint even though only a fraction of parameters are active per token.
- Outputs may be inaccurate, incomplete, biased, or fabricated. High-impact decisions require independent review.
Security responsibility
Deployers are responsible for authentication, authorization, rate limits, network isolation, logging, sandboxing, secrets handling, and tool-side validation. Keep a human in the loop for consequential actions. Treat retrieved content and tool output as untrusted input, and never give the model broader privileges than the surrounding application requires.
License
The parent model and this derivative are distributed under the Apache License 2.0. Preserve required copyright, patent, trademark, and attribution notices when redistributing the model or derivative works.
The official Aleph Alpha inference plugin is also Apache-2.0. Container layers and transitive dependencies may carry their own licenses; review them for your deployment.
Disclaimer
This research release is provided as-is, without warranty. Users are responsible for determining whether the model, generated output, and deployment configuration are appropriate for their application and jurisdiction.
Reporting
Report reproducible model, card, tokenizer, or deployment-kit issues through this repository's Hugging Face Discussions tab. Include the repository revision, runtime image digest, hardware class, serving flags, prompt/template settings, sampling parameters, and a minimal redacted reproduction.
- Downloads last month
- 28
Model tree for Blackfrost-AI/KOLIBRI-1-DERISKED-FP8
Base model
Aleph-Alpha/Kolibri-1-BF16