ruri-v3-310m-coreai

English | 日本語

Model Summary

This is an unofficial Core AI conversion of cl-nagoya/ruri-v3-310m (a Japanese text-embedding model), Apple's new on-device AI runtime (iOS/macOS 27+, the successor to Core ML). All credit for the original model goes to its authors (Nagoya University / cl-nagoya).

  • Live demo: Ruri Atlas — explore Japanese & English sentence embeddings as a 3-D map in your browser (WebGPU), switch embedding models, toggle the reranker, and see where a prompt-injection lands. Nothing is sent to a server.

Validated on both Neural Engine and GPU

This is rewritten from scratch for the Neural Engine (BC1S (Batch, Channel, 1, Sequence) layout, Conv2d projections, explicit per-head attention), following the same approach as masahiroid/ruri-v3-130m-coreai, which documents the architecture details (alternating full/sliding-window attention, per-flavor RoPE theta, radius-65 bidirectional window) and the fp32-weight dtype pitfall. This model is ModernBERT-Ja with 25 layers and hidden size 768; ruri_ane.py bundles the implementation (set ANE_MODEL_NAME=cl-nagoya/ruri-v3-310m to select this size).

Usage (Python, coreai.runtime)

# /// script
# requires-python = ">=3.11"
# dependencies = ["coreai-core==1.0.0b3", "transformers", "torch", "sentencepiece", "protobuf"]
# ///
import asyncio
import torch
import transformers
from coreai.runtime import AIModel, ComputeUnitKind, NDArray, SpecializationOptions

async def main():
    tokenizer = transformers.AutoTokenizer.from_pretrained("cl-nagoya/ruri-v3-310m")
    text = "検索文書: 瑠璃色(るりいろ)は、紫みを帯びた濃い青のことである。"
    enc = tokenizer(text, return_tensors="pt", padding="max_length", truncation=True, max_length=128)

    # ComputeUnitKind.neural_engine() also available
    opts = SpecializationOptions.from_preferred_compute_unit_kind(ComputeUnitKind.gpu())
    model = await AIModel.load("ruri-v3-310m-ane_float16_seq128.aimodel", specialization_options=opts)
    fn = model.load_function(model.function_names[0])

    out = await fn(inputs={
        "input_ids": NDArray(enc["input_ids"].to(torch.int32).numpy()),
        "attention_mask": NDArray(enc["attention_mask"].to(torch.int32).numpy()),
    })
    embedding = out["sentence_embedding"].numpy()[0]  # (768,) L2-normalized
    print(embedding[:5])

asyncio.run(main())

Three fixed sequence lengths are bundled: ..._seq128.aimodel, ..._seq256.aimodel and ..._seq512.aimodel (set max_length to match).

Prefixes (important)

ruri-v3 uses a "1+3 prefix scheme" -- prepend the appropriate prefix before tokenizing:

  • Empty string: general semantic similarity
  • トピック: ("Topic: "): classification/clustering
  • 検索クエリ: ("Search query: "): the query side of retrieval
  • 検索文書: ("Search document: "): the document side of retrieval

Accuracy

Cosine similarity against the PyTorch fp32 reference (one real Japanese sentence, padded):

Target seq128 seq256 seq512
GPU specialization 1.0001693 1.0002791 1.0002584
Neural Engine specialization 1.0005926 1.0005926 1.0007246

Specs

Item Value
Base model cl-nagoya/ruri-v3-310m (ModernBERT-Ja, 25 layers, hidden 768)
Precision float16
Input fixed 128 / 256 / 512 tokens
Framework Core AI (coreai-torch, BC1S implementation in ruri_ane.py)
Compute units GPU, Neural Engine (both validated)

Notes

  • This is a community conversion, not an official release from the Ruri / cl-nagoya team.
  • Security audit uses model-audit-lite (see SECURITY.md for details).

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

cl-nagoya/ruri-v3-310m(日本語テキスト埋め込みモデル)を、 Appleの新しいオンデバイスAIランタイム「Core AI」 (iOS/macOS 27以降、Core MLの後継)で動かせるように変換したものです。元モデルの著作権は その作者(名古屋大学 / cl-nagoya)に帰属します。

  • ライブデモ: Ruri Atlas — 日本語・英語の文の埋め込みを、ブラウザ上(WebGPU)で3Dの地図として見られます。埋め込みモデルの切り替え、リランカーのON/OFF、プロンプトインジェクションの効き方の可視化もできます。文章はサーバーに送られません。

Neural EngineとGPUの両方で動作確認済み

Neural Engine向けにゼロから書き直しています(BC1Sレイアウト、Conv2dベースの投影、 ヘッドごとの明示的attention計算)。方針は masahiroid/ruri-v3-130m-coreai と同じで、アーキテクチャの詳細(full/sliding window attentionの交互配置、attentionの種類ごとの RoPE theta、半径65の双方向windowマスク)とfp32重みdtypeの罠はそちらに記載しています。 本モデルは25層・hidden 768のModernBERT-Jaで、同梱のruri_ane.pyが実装です (ANE_MODEL_NAME=cl-nagoya/ruri-v3-310mでこのサイズを選択)。

使い方(Python, coreai.runtime)

上記English節のコードをそのまま使えます(from_pretrained("cl-nagoya/ruri-v3-310m")、 ruri-v3-310m-ane_float16_seq128.aimodel)。出力は768次元のL2正規化済みベクトルです。 固定サイズは128 / 256 / 512トークンの3種類を同梱しています(max_lengthを合わせてください)。

プレフィックスについて(重要)

ruri-v3は"1+3 prefix scheme"を採用しているため、用途に応じて以下を付けてからトークナイズしてください。

  • 空文字列: 意味的な類似度計算全般
  • トピック: : 分類・クラスタリング用
  • 検索クエリ: : 検索クエリ側
  • 検索文書: : 検索対象の文書側

精度検証

PyTorch fp32リファレンスとのコサイン類似度(実際の日本語文1件、パディング含む):

ターゲット seq128 seq256 seq512
GPU specialization 1.0001693 1.0002791 1.0002584
Neural Engine specialization 1.0005926 1.0005926 1.0007246

Specs

Item Value
ベースモデル cl-nagoya/ruri-v3-310m(ModernBERT-Ja、25層、hidden 768)
精度 float16
入力 固定128/256/512トークン
フレームワーク Core AI(coreai-torch、ruri_ane.pyによるBC1S実装)
対応コンピュートユニット GPU, Neural Engine(両方で動作確認済み)

備考

  • 本変換は非公式のコミュニティ版です。cl-nagoyaチームによる公式リリースではありません。
  • セキュリティー監査にはmodel-audit-liteを 使用しています(詳細はSECURITY.md)。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/ruri-v3-310m-coreai

Finetuned
(6)
this model