ruri-v3-310m-coreai
Model Summary
This is an unofficial Core AI conversion of cl-nagoya/ruri-v3-310m (a Japanese text-embedding model), Apple's new on-device AI runtime (iOS/macOS 27+, the successor to Core ML). All credit for the original model goes to its authors (Nagoya University / cl-nagoya).
- Live demo: Ruri Atlas — explore Japanese & English sentence embeddings as a 3-D map in your browser (WebGPU), switch embedding models, toggle the reranker, and see where a prompt-injection lands. Nothing is sent to a server.
Validated on both Neural Engine and GPU
This is rewritten from scratch for the Neural Engine (BC1S
(Batch, Channel, 1, Sequence) layout, Conv2d projections, explicit per-head
attention), following the same approach as
masahiroid/ruri-v3-130m-coreai,
which documents the architecture details (alternating full/sliding-window
attention, per-flavor RoPE theta, radius-65 bidirectional window) and the
fp32-weight dtype pitfall. This model is ModernBERT-Ja with 25 layers and
hidden size 768; ruri_ane.py bundles the implementation (set
ANE_MODEL_NAME=cl-nagoya/ruri-v3-310m to select this size).
Usage (Python, coreai.runtime)
# /// script
# requires-python = ">=3.11"
# dependencies = ["coreai-core==1.0.0b3", "transformers", "torch", "sentencepiece", "protobuf"]
# ///
import asyncio
import torch
import transformers
from coreai.runtime import AIModel, ComputeUnitKind, NDArray, SpecializationOptions
async def main():
tokenizer = transformers.AutoTokenizer.from_pretrained("cl-nagoya/ruri-v3-310m")
text = "検索文書: 瑠璃色(るりいろ)は、紫みを帯びた濃い青のことである。"
enc = tokenizer(text, return_tensors="pt", padding="max_length", truncation=True, max_length=128)
# ComputeUnitKind.neural_engine() also available
opts = SpecializationOptions.from_preferred_compute_unit_kind(ComputeUnitKind.gpu())
model = await AIModel.load("ruri-v3-310m-ane_float16_seq128.aimodel", specialization_options=opts)
fn = model.load_function(model.function_names[0])
out = await fn(inputs={
"input_ids": NDArray(enc["input_ids"].to(torch.int32).numpy()),
"attention_mask": NDArray(enc["attention_mask"].to(torch.int32).numpy()),
})
embedding = out["sentence_embedding"].numpy()[0] # (768,) L2-normalized
print(embedding[:5])
asyncio.run(main())
Three fixed sequence lengths are bundled: ..._seq128.aimodel,
..._seq256.aimodel and ..._seq512.aimodel (set max_length to match).
Prefixes (important)
ruri-v3 uses a "1+3 prefix scheme" -- prepend the appropriate prefix before tokenizing:
- Empty string: general semantic similarity
トピック:("Topic: "): classification/clustering検索クエリ:("Search query: "): the query side of retrieval検索文書:("Search document: "): the document side of retrieval
Accuracy
Cosine similarity against the PyTorch fp32 reference (one real Japanese sentence, padded):
| Target | seq128 | seq256 | seq512 |
|---|---|---|---|
| GPU specialization | 1.0001693 | 1.0002791 | 1.0002584 |
| Neural Engine specialization | 1.0005926 | 1.0005926 | 1.0007246 |
Specs
| Item | Value |
|---|---|
| Base model | cl-nagoya/ruri-v3-310m (ModernBERT-Ja, 25 layers, hidden 768) |
| Precision | float16 |
| Input | fixed 128 / 256 / 512 tokens |
| Framework | Core AI (coreai-torch, BC1S implementation in ruri_ane.py) |
| Compute units | GPU, Neural Engine (both validated) |
Notes
- This is a community conversion, not an official release from the Ruri / cl-nagoya team.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
cl-nagoya/ruri-v3-310m(日本語テキスト埋め込みモデル)を、 Appleの新しいオンデバイスAIランタイム「Core AI」 (iOS/macOS 27以降、Core MLの後継)で動かせるように変換したものです。元モデルの著作権は その作者(名古屋大学 / cl-nagoya)に帰属します。
- ライブデモ: Ruri Atlas — 日本語・英語の文の埋め込みを、ブラウザ上(WebGPU)で3Dの地図として見られます。埋め込みモデルの切り替え、リランカーのON/OFF、プロンプトインジェクションの効き方の可視化もできます。文章はサーバーに送られません。
Neural EngineとGPUの両方で動作確認済み
Neural Engine向けにゼロから書き直しています(BC1Sレイアウト、Conv2dベースの投影、
ヘッドごとの明示的attention計算)。方針は
masahiroid/ruri-v3-130m-coreai
と同じで、アーキテクチャの詳細(full/sliding window attentionの交互配置、attentionの種類ごとの
RoPE theta、半径65の双方向windowマスク)とfp32重みdtypeの罠はそちらに記載しています。
本モデルは25層・hidden 768のModernBERT-Jaで、同梱のruri_ane.pyが実装です
(ANE_MODEL_NAME=cl-nagoya/ruri-v3-310mでこのサイズを選択)。
使い方(Python, coreai.runtime)
上記English節のコードをそのまま使えます(from_pretrained("cl-nagoya/ruri-v3-310m")、
ruri-v3-310m-ane_float16_seq128.aimodel)。出力は768次元のL2正規化済みベクトルです。
固定サイズは128 / 256 / 512トークンの3種類を同梱しています(max_lengthを合わせてください)。
プレフィックスについて(重要)
ruri-v3は"1+3 prefix scheme"を採用しているため、用途に応じて以下を付けてからトークナイズしてください。
- 空文字列: 意味的な類似度計算全般
トピック:: 分類・クラスタリング用検索クエリ:: 検索クエリ側検索文書:: 検索対象の文書側
精度検証
PyTorch fp32リファレンスとのコサイン類似度(実際の日本語文1件、パディング含む):
| ターゲット | seq128 | seq256 | seq512 |
|---|---|---|---|
| GPU specialization | 1.0001693 | 1.0002791 | 1.0002584 |
| Neural Engine specialization | 1.0005926 | 1.0005926 | 1.0007246 |
Specs
| Item | Value |
|---|---|
| ベースモデル | cl-nagoya/ruri-v3-310m(ModernBERT-Ja、25層、hidden 768) |
| 精度 | float16 |
| 入力 | 固定128/256/512トークン |
| フレームワーク | Core AI(coreai-torch、ruri_ane.pyによるBC1S実装) |
| 対応コンピュートユニット | GPU, Neural Engine(両方で動作確認済み) |
備考
- 本変換は非公式のコミュニティ版です。cl-nagoyaチームによる公式リリースではありません。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
Model tree for masahiroid/ruri-v3-310m-coreai
Base model
sbintuitions/modernbert-ja-310m