ruri-v3-130m-coreai
Model Summary
This is an unofficial Core AI conversion of cl-nagoya/ruri-v3-130m (a general-purpose Japanese text-embedding model), Apple's new on-device AI runtime (iOS/macOS 27+, the successor to Core ML). All credit for the original model goes to its authors (Nagoya University / cl-nagoya).
- Live demo: Ruri Atlas — explore Japanese & English sentence embeddings as a 3-D map in your browser (WebGPU), switch embedding models, toggle the reranker, and see where a prompt-injection lands. Nothing is sent to a server.
Validated on both Neural Engine and GPU
This has been rewritten from scratch for Neural Engine. Simply converting
the existing model as-is does not load on the Neural Engine in standard
PyTorch matmul layout -- it requires rewriting to Apple's BC1S
(Batch, Channel, 1, Sequence) layout, Conv2d-based projections, and
explicit per-head attention. This reimplementation faithfully reproduces
ModernBERT-Ja's specific architecture (19 layers, full/sliding-window
attention alternating every 3rd layer, a different RoPE theta per attention
flavor, and a bidirectional sliding-window mask with radius 65).
The bundled ruri_ane.py is needed to understand or reproduce this
conversion in Python (the .aimodel itself can be embedded directly into a
Swift app without it).
Usage (Python, coreai.runtime)
# /// script
# requires-python = ">=3.11"
# dependencies = ["coreai-core==1.0.0b3", "transformers", "torch", "sentencepiece", "protobuf"]
# ///
import asyncio
import numpy as np
import torch
import transformers
from coreai.runtime import AIModel, ComputeUnitKind, NDArray, SpecializationOptions
async def main():
tokenizer = transformers.AutoTokenizer.from_pretrained("cl-nagoya/ruri-v3-130m")
text = "検索文書: 瑠璃色(るりいろ)は、紫みを帯びた濃い青のことである。"
enc = tokenizer(text, return_tensors="pt", padding="max_length", truncation=True, max_length=128)
# ComputeUnitKind.neural_engine() also available
opts = SpecializationOptions.from_preferred_compute_unit_kind(ComputeUnitKind.gpu())
model = await AIModel.load("ruri-v3-130m_float16_seq128.aimodel", specialization_options=opts)
# For longer text, use ruri-v3-130m-ane_float16_seq256.aimodel (+ max_length=256) or seq512 (+ max_length=512)
fn = model.load_function(model.function_names[0])
out = await fn(inputs={
"input_ids": NDArray(enc["input_ids"].to(torch.int32).numpy()),
"attention_mask": NDArray(enc["attention_mask"].to(torch.int32).numpy()),
})
embedding = out["sentence_embedding"].numpy()[0] # (512,) L2-normalized
print(embedding[:5])
asyncio.run(main())
Run with uv run test.py (uv builds a throwaway environment for this
script automatically, from the inline # /// script metadata above).
Three fixed sequence lengths are bundled: 128 tokens
(ruri-v3-130m_float16_seq128.aimodel), 256 tokens
(ruri-v3-130m-ane_float16_seq256.aimodel) and 512 tokens
(ruri-v3-130m-ane_float16_seq512.aimodel) for longer text. Pass a
different seq_len to RuriAneModel(...) in ruri_ane.py to re-export
other buckets.
Prefixes (important)
ruri-v3 uses a "1+3 prefix scheme" -- prepend the appropriate prefix before tokenizing:
- Empty string: general semantic similarity
トピック:("Topic: "): classification/clustering検索クエリ:("Search query: "): the query side of retrieval検索文書:("Search document: "): the document side of retrieval
Accuracy
Compared against the PyTorch fp32 reference (a real Japanese sentence, padded to 24/128 real tokens):
| Target | seq128 | seq256 | seq512 |
|---|---|---|---|
| GPU specialization | 1.0001926 | 1.0002351 | 1.0001947 |
| Neural Engine specialization | 1.0002038 | 1.0002038 | 1.0002571 |
During conversion, Neural Engine execution specifically showed degraded
accuracy (cosine 0.954) due to a subtle bug: nn.Parameter/nn.Conv2d
default to fp32, and .copy_() during weight loading only overwrites
values, not dtype -- fixed by explicitly casting the whole module to fp16
(model.half()) before loading weights. See the code comments, or Pattern
11 of mlx-coreml-conversion-toolkit
for the full writeup.
Specs
| Item | Value |
|---|---|
| Base model | cl-nagoya/ruri-v3-130m (ModernBERT-Ja, 132M params) |
| Precision | float16 |
| Input | fixed 128 / 256 / 512 tokens |
| Framework | Core AI (coreai-torch, Neural-Engine-targeted BC1S implementation in ruri_ane.py) |
| Compute units | GPU, Neural Engine (both validated) |
Notes
- This is a community conversion, not an official release from the Ruri / cl-nagoya team.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
cl-nagoya/ruri-v3-130m(日本語汎用テキスト埋め込みモデル)を、 Appleの新しいオンデバイスAIランタイム「Core AI」 (iOS/macOS 27以降、Core MLの後継)で動かせるように変換したものです。元モデルの著作権は その作者(名古屋大学 / cl-nagoya)に帰属します。
- ライブデモ: Ruri Atlas — 日本語・英語の文の埋め込みを、ブラウザ上(WebGPU)で3Dの地図として見られます。埋め込みモデルの切り替え、リランカーのON/OFF、プロンプトインジェクションの効き方の可視化もできます。文章はサーバーに送られません。
Neural EngineとGPUの両方で動作確認済み
Neural Engine向けにゼロから書き直しています。単純に既存モデルを変換しただけでは、標準的な
PyTorchの行列演算レイアウトのままNeural Engineにロードできません(Appleが定めるBC1S
(Batch, Channel, 1, Sequence)レイアウト・Conv2dベースの投影・ヘッドごとの明示的attention
計算への書き直しが必要です)。ModernBERT-Ja特有の構成(19層、3層ごとに切り替わる
full/sliding window attention、attentionの種類ごとに異なるRoPE theta、半径65の双方向
sliding windowマスク)を正確に再現した上でNeural Engine向けに再実装しています。
同梱のruri_ane.pyが必要です(.aimodelはSwiftアプリに直接組み込めますが、Pythonで動作確認・
再現する場合はこのソースコードを参照してください)。
使い方(Python, coreai.runtime)
# /// script
# requires-python = ">=3.11"
# dependencies = ["coreai-core==1.0.0b3", "transformers", "torch", "sentencepiece", "protobuf"]
# ///
import asyncio
import numpy as np
import torch
import transformers
from coreai.runtime import AIModel, ComputeUnitKind, NDArray, SpecializationOptions
async def main():
tokenizer = transformers.AutoTokenizer.from_pretrained("cl-nagoya/ruri-v3-130m")
text = "検索文書: 瑠璃色(るりいろ)は、紫みを帯びた濃い青のことである。"
enc = tokenizer(text, return_tensors="pt", padding="max_length", truncation=True, max_length=128)
# ComputeUnitKind.neural_engine() も指定可能
opts = SpecializationOptions.from_preferred_compute_unit_kind(ComputeUnitKind.gpu())
model = await AIModel.load("ruri-v3-130m_float16_seq128.aimodel", specialization_options=opts)
# 長文には ruri-v3-130m-ane_float16_seq256.aimodel(max_length=256)または seq512(max_length=512)を使用
fn = model.load_function(model.function_names[0])
out = await fn(inputs={
"input_ids": NDArray(enc["input_ids"].to(torch.int32).numpy()),
"attention_mask": NDArray(enc["attention_mask"].to(torch.int32).numpy()),
})
embedding = out["sentence_embedding"].numpy()[0] # (512,) L2正規化済み
print(embedding[:5])
asyncio.run(main())
uv run test.pyで実行できます(uvが自動でこのスクリプト専用の環境を作ります)。
固定サイズ128トークン(ruri-v3-130m_float16_seq128.aimodel)、256トークン
(ruri-v3-130m-ane_float16_seq256.aimodel)、512トークン(ruri-v3-130m-ane_float16_seq512.aimodel、長文向け)の
3種類を同梱しています。ruri_ane.pyのRuriAneModel(seq_len=...)を変更すれば、他のバケットも再エクスポートできます。
プレフィックスについて(重要)
ruri-v3は"1+3 prefix scheme"を採用しているため、埋め込み対象のテキストに応じて以下の プレフィックスを付けてからトークナイズしてください。
- 空文字列: 意味的な類似度計算全般
トピック:: 分類・クラスタリング用検索クエリ:: 検索クエリ側検索文書:: 検索対象の文書側
精度検証
PyTorch fp32リファレンスとの比較(実際の日本語文1件、パディング含む24/128トークン):
| ターゲット | seq128 | seq256 | seq512 |
|---|---|---|---|
| GPU specialization | 1.0001926 | 1.0002351 | 1.0001947 |
| Neural Engine specialization | 1.0002038 | 1.0002038 | 1.0002571 |
変換過程で、Neural Engine実行時のみ精度が0.954まで劣化する問題に遭遇しました。原因は
nn.Parameter/nn.Conv2dのデフォルトdtypeがfp32で、重みロード時の.copy_()が値のみを
上書きしdtypeを変えないという罠でした(詳細はリポジトリ内のコメント、および
mlx-coreml-conversion-toolkit
のPattern 11を参照)。model.half()での明示的なfp16化で解消しています。
Specs
| Item | Value |
|---|---|
| ベースモデル | cl-nagoya/ruri-v3-130m(ModernBERT-Ja、132M params) |
| 精度 | float16 |
| 入力 | 固定128/256/512トークン |
| フレームワーク | Core AI(coreai-torch、ruri_ane.pyによるNeural Engine向けBC1S実装) |
| 対応コンピュートユニット | GPU, Neural Engine(両方で動作確認済み) |
備考
- 本変換は非公式のコミュニティ版です。cl-nagoyaチームによる公式リリースではありません。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
Model tree for masahiroid/ruri-v3-130m-coreai
Base model
sbintuitions/modernbert-ja-130m