BLACKFROST-AI

KOLIBRI-1-DERISKED-FP8

Block-FP8 release of the behaviorally modified Kolibri 1 derivative for direct, high-agency reasoning, long-context, and tool-enabled workflows.

Release status: public and ungated. The standalone FP8 checkpoint, configuration, tokenizer, model card, account banner, and deployment kit are included.

Artifact contents

Component Status
Model card Included
Blackfrost-AI account banner Included
Apache-2.0 license Included
TP4 deployment kit Included under DEPLOYMENT/
Standalone FP8 weights Included: 32 SafeTensors shards
Model configuration and tokenizer assets Included

This is a standalone FP8 checkpoint. It is not an adapter and does not require the parent checkpoint at inference time.

Why this release exists

Kolibri 1 is an efficient bilingual mixture-of-experts model with explicit reasoning, long-context support, and structured tool calling. This derivative provides an alternative behavioral operating point intended for users who want a more direct assistant while preserving the architecture's reasoning, retrieval, multi-turn, and tool-use strengths.

The weights are behaviorally modified. The proprietary transformation process and reconstruction materials are not distributed.

Model specifications

Property Value
Architecture Kolibri1ForCausalLM, sparse mixture of experts
Precision FP8 E4M3 weights with 128×128 FP32 block scales and dynamic activations; router gates remain BF16
Total parameters 78,103,074,560
Active parameters per token 3,457,573,120
Transformer blocks 50
Experts per MoE block 384 routed experts plus one shared expert
Routed experts per token 6
Attention pattern 4:1 sliding-window to full-attention pattern
Languages German and English
Reasoning Explicit reasoning mode with configurable effort
Tool use Structured tool calling through the Kolibri parser
Native model context 262,144 tokens
Blackfrost validated context 262,144 tokens
Upstream extended-context claim Up to 1,048,576 tokens; not independently qualified on this derivative
Checkpoint tensor footprint 78.83 GB
Recommended runtime Aleph Alpha's official vLLM inference plugin/container

Upstream architecture and training details are documented in the Kolibri model card and technical report.

Lineage

  1. Foundational and immediate parent: Aleph-Alpha/Kolibri-1-BF16
  2. Immutable parent revision: 8c8b34899c7bc77b349f7916b07158e3ba29456a
  3. Tokenizer and chat-template lineage: parent Kolibri 1 release
  4. Blackfrost stage: behaviorally modified BF16 derivative; proprietary transformation process not distributed
  5. Quantization stage: official-layout 128×128 block-FP8 export from the completed derivative
  6. Public artifact: standalone FP8 SafeTensors checkpoint

No upstream benchmark number is presented here as a measurement of this derivative. Consult the parent card for upstream evaluations and training-data documentation.

Artifact verification

The FP8 checkpoint completed structural and numerical verification before publication:

  • all 32 expected checkpoint shards completed atomically;
  • 57,950 weights follow Aleph Alpha's official Kolibri FP8 quantization surface;
  • attention, routed experts, and shared experts use FP8 E4M3 with 128×128 FP32 block scales;
  • 403 router-gate and other excluded tensors remain exactly equal to the derivative BF16 source;
  • 185 distributed reconstruction samples passed, with minimum cosine similarity 0.9995573 and maximum relative L2 error 0.0265193;
  • the artifact contains 116,303 tensors totaling 78,827,029,120 tensor bytes;
  • the source checkpoint remains immutable throughout conversion.

Parent fingerprints used during construction:

Item SHA-256
Parent config.json 2577f567b37487b1ee24ccc9f7e9d9895b01ce6d8ae248f481a41130b04be256
Parent model.safetensors.index.json 370e3721203856717b806a3c1c7f07e5b149dc3e534096f047f6de9e34977c0d

The complete public checkpoint checksum inventory will be added with the weight artifact. Deployment-kit checksums are included in DEPLOYMENT/CHECKSUMS.sha256.

Behavioral and capability evaluation

The behavioral predecessor to this quant completed two independent long-form and tool-integrity harness passes, each satisfying 74 of 75 assertions. These figures describe the selected BF16 derivative before FP8 export; a full post-quant behavioral rerun remains pending.

Test area Measured result
Independent full harness passes 2
Assertions passed 74/75 in each pass
Long-form section completion 9/9 sections plus exact end marker in both passes
Distant-context retrieval Exact requested records from a 1,000-record context in both passes
Multi-turn memory Exact requested state across five turns in both passes
Heterogeneous two-stage tool workflows 8/8 passed across both passes
Repeated and parallel weather workflows 8/8 passed across both passes
Malformed tool names or arguments 0 observed
Textual tool-call leakage 0 observed
Output or tool-call loops 0 observed

The only failed assertion in each pass was a strict request for at least 1,000 visible words. Both responses were coherent, completed every requested section, and emitted the required end marker, but were shorter than the threshold.

These are focused integrity checks, not a comprehensive capability benchmark. A quantitative refusal-behavior evaluation, broad academic benchmark suite, multilingual regression suite, and extended-context qualification beyond 262,144 tokens remain pending.

Operator review observed the intended softer, more direct behavior when the checkpoint was used with an operator-supplied system prompt. That observation is qualitative: the automated harness did not measure refusal rate or isolate the contribution of the system prompt from the checkpoint.

Prompt, reasoning, and sampling notes

Use the packaged Kolibri chat template. It supports:

  • reasoning_effort: low, medium, high, or none;
  • enable_thinking: true or false;
  • structured tools through the standard OpenAI-compatible tools field;
  • multi-step tool-result reinjection.

The upstream recommended sampling defaults are:

Parameter Value
temperature 1.0
top_p 0.97
top_k 128

The Blackfrost integrity harness used sampling rather than greedy decoding. Applications should validate their own system prompts, sampling settings, structured-output schemas, and tool policies. This release does not prescribe a default persona or application policy.

Deployment

The included DEPLOYMENT/ kit provides the intended serving shape with:

  • the official Aleph Alpha Kolibri inference container;
  • vLLM's OpenAI-compatible API;
  • FP8 weights with BF16 activations;
  • tensor parallelism across four GPUs;
  • a 262,144-token context ceiling;
  • Kolibri reasoning and tool-call parsers;
  • automatic tool choice and prefix caching.

Start with DEPLOYMENT/README.md. The launcher publishes the clean API model identifier KOLIBRI-1-DERISKED-FP8.

Limitations

  • This FP8 export has not yet received a comprehensive post-quant academic, coding, German-language, safety, or refusal-behavior evaluation.
  • The 1,048,576-token upstream serving mode has not been independently qualified on this derivative; the published deployment kit stays at the validated 262,144-token profile.
  • The strict 1,000-visible-word requirement failed in both focused harness runs despite complete section coverage.
  • Tool calls are model-generated data. Validate function names, schemas, arguments, authorization, and tool outputs before execution.
  • Mixture-of-experts inference has a large memory footprint even though only a fraction of parameters are active per token.
  • Outputs may be inaccurate, incomplete, biased, or fabricated. High-impact decisions require independent review.

Security responsibility

Deployers are responsible for authentication, authorization, rate limits, network isolation, logging, sandboxing, secrets handling, and tool-side validation. Keep a human in the loop for consequential actions. Treat retrieved content and tool output as untrusted input, and never give the model broader privileges than the surrounding application requires.

License

The parent model and this derivative are distributed under the Apache License 2.0. Preserve required copyright, patent, trademark, and attribution notices when redistributing the model or derivative works.

The official Aleph Alpha inference plugin is also Apache-2.0. Container layers and transitive dependencies may carry their own licenses; review them for your deployment.

Disclaimer

This research release is provided as-is, without warranty. Users are responsible for determining whether the model, generated output, and deployment configuration are appropriate for their application and jurisdiction.

Reporting

Report reproducible model, card, tokenizer, or deployment-kit issues through this repository's Hugging Face Discussions tab. Include the repository revision, runtime image digest, hardware class, serving flags, prompt/template settings, sampling parameters, and a minimal redacted reproduction.

Downloads last month
28
Safetensors
Model size
78B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Blackfrost-AI/KOLIBRI-1-DERISKED-FP8

Quantized
(17)
this model