Ornith 1.5 35B-A3B for Mac: 4-bit Splash package with vision

Run Ornith-1.5-35B-A3B on an Apple silicon Mac with Splash: a community 4-bit conversion with vision, 1.47× faster than the same weights on MLX with MTP for short code in a single stream, and 3.46× with four concurrent requests. Independent and unendorsed by Ornith AI or Inco AI.

brew install incoai/tap/splash
splash serve --model ezoushen/ornith-1.5-35b-a3b-splash

Requires Apple M3 or newer, macOS 26.4+, and at least 36 GB unified memory. About 20 GB of Splash-specific packed binaries — not Transformers, GGUF or MLX.

At a glance

Results from an M5 Max (40-core GPU, 128 GB); qualifications below.

This package Comparison or notes
Size 20.9 GB, 56 files, vision included Ornith's MLX 4-bit repo has no vision, and Splash 1.1.0 stops on its config (reported)
Short code, single stream 1.47× faster the same weights on MLX with MTP
Four concurrent requests 3.46× faster the same weights on MLX with MTP
Draft acceptance 2.63 tokens per step 4.45 for Inco's own Qwen3.6 package
Identical greedy output to Ornith on MLX 136 of 200 prompts code 40/40 · vision 40/40 · long context 35/40 · math 21/40 · chat 0/40
GSM8K (32-item subset) 30 / 32 28 / 32 on the MLX production lane
JevBench public decisions (231, /v1/systemone) 83.1% —
Splash versions 1.0.2, with vision also loads on 1.1.0 (checked text-only)

Draft, tokenizer and output fidelity

The draft was trained on Qwen3.6-35B-A3B, not Ornith. Inco's DFlash 2 draft is reused unmodified. Splash decodes only through the draft, with no autoregressive fallback. Acceptance is 2.631 tokens per verify step versus 4.447 for Inco's Qwen3.6 package on the same harness: 59.2%. An Ornith-conditioned draft would be faster; training one remains unfinished.

Combining marks may tokenise differently than on other Ornith engines. The packed tokenizer's pre_tokenizer Split regex omits \p{M} from two character classes where official Ornith's and Inco's shipped tokenizers include it. Vocabulary (248,044 entries) and merges match Ornith's; two pre-tokenisation fields do not. The variant's origin is unknown. Prompt rendering matched on 12 of 12 test conversations, but none contained combining marks, so this does not generalise.

Output is not token-identical to the production MLX lane: 136/200 exact matches, with the breakdown in the table. Speculative decoding was ruled out as the cause (50/50 exact against a zero-draft control), as was prompt rendering (0 differing artifacts of 12). The 32-item GSM8K scores are 93.75% versus the MLX lane's 87.50% — a subset result, not a quality verdict.

What is in it

component origin licence notes
language-model body, 625 sections ornith-ai/Ornith-1.5-35B-A3B MIT repacked bit-exactly; max_abs=0 against the source; 555 direct, 40 fused, 30 derived, no dequantize-requantize
vision tower, 333 bf16 sections ornith-ai/Ornith-1.5-35B-A3B MIT byte-identical to Inco's Qwen3.6 tower, because the two models ship the same vision weights — verified by sampled value comparison against official Ornith
tokenizer ornith-ai/Ornith-1.5-35B-A3B MIT vocabulary and merges verified identical; see the \p{M} caveat above
DFlash 2 draft, 7 files incoai/Qwen3.6-35B-A3B-Splash Apache-2.0 reused unmodified; see below
container layout, manifest, packing scheme Inco's schema 4 Apache-2.0 their format, our writer

Integrity: verify-package.py --full passes over 56 artifacts, 20,949,446,234 bytes, SHA-256 against the manifest.

On redistributing the draft

Inco's incoai/Qwen3.6-35B-A3B-Splash card states that Apache-2.0 covers its DFlash 2 draft, vision encoder and tokenizer. Its perpetual, irrevocable §2 grants permit redistribution; this package complies with §4 by carrying the licence, retaining attribution and stating changes.

The draft's source repository, incoai/Qwen3.6-35B-A3B-DFlash2, is gated and returns HTTP 401; the copy in the Apache-2.0 package is not. The gate signals an intent the licence does not encode. If Inco prefers that this package not carry its draft, it will be removed on request.

Speed

Same weights, same host, one session, arms alternated — Splash against an MLX lane with MTP:

shape speedup band pairs
single stream, short code 1.47× 1.36–1.50 4
four concurrent requests 3.46× 3.38–3.47 3

Absolute tok/s are omitted because the same benchmark cell varied by 1.43× between runs; ratios within each pair cancel that drift. Bands include only cells passing a CPU-quiescence gate that was harsher on MLX, slightly favouring Splash. Excluded cells read 1.34× and 3.18–3.62×.

Limitations

Natural photographs are untested; vision coverage is generated shapes and colours. No safety evaluation has been performed. See above for draft, tokenizer and output differences. Since 2026-09-30 the chat template renders a system message after the first as a system turn, the one-line change incoai's Qwen3.6 Splash package makes to the upstream template. Earlier downloads carry the upstream Ornith template, which raises on it: Splash 1.0.2 and runtimes that use the template as shipped answer messages could not be rendered, so re-download (discussion).

Related

  • Umpire 35B-A3B: a calibrated decision fine-tune of Ornith 1.5, also packaged for Splash.
  • Umpire MLX 4-bit: its MLX checkpoint, for reproducing Umpire served beside Ornith's MLX checkpoint without a second copy of the shared weights, with the shared-expert-weights Splash fork.

Credits and licence

Apache-2.0 for the draft, vision encoder, tokenizer container and packing format, from Inco AI. MIT for the Ornith weights, from Ornith AI. See LICENSE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ezoushen/ornith-1.5-35b-a3b-splash

Quantized
(187)
this model

Collection including ezoushen/ornith-1.5-35b-a3b-splash