LCO-Embedding-Omni-3B-2605-GGUF

GGUF quantizations of LCO-Embedding/LCO-Embedding-Omni-3B-2605 for use with llama.cpp.

Converted using ht-llama.cpp, a fork with added support for the Qwen2_5OmniThinkerForConditionalGeneration architecture.

For the earlier release see LCO-Embedding-Omni-3B-GGUF.

About the model

LCO-Embedding-Omni-3B is a multimodal embedding model based on the Thinker component of Qwen 2.5 Omni, fine-tuned with LoRA and contrastive learning to produce 2048-dimensional embeddings from text, images, audio, and video. Uses last-token pooling.

See Scaling Language-Centric Omnimodal Representation Learning (NeurIPS 2025) for details.

Available files

File Quant Size Description
LCO-Embedding-Omni-3B-2605-Q8_0.gguf Q8_0 3.4 GB Near-lossless quantization
LCO-Embedding-Omni-3B-2605-Q6_K.gguf Q6_K 2.7 GB Very high quality, smaller than Q8
LCO-Embedding-Omni-3B-2605-Q5_K_M.gguf Q5_K_M 2.3 GB High quality, balanced size
LCO-Embedding-Omni-3B-2605-Q4_K_M.gguf Q4_K_M 2.0 GB Good balance of quality and size
mmproj-LCO-Embedding-Omni-3B-2605-F16.gguf F16 2.5 GB Vision + audio projection (required for multimodal)

For text-only embedding, you only need one of the text model GGUFs. For multimodal (image/audio/video), you also need the mmproj file.

Usage

Build llama.cpp

git clone https://github.com/heiervang-technologies/ht-llama.cpp
cd ht-llama.cpp
cmake -B build
cmake --build build --target llama-embedding llama-server -j$(nproc)

Text embeddings (CLI)

./build/bin/llama-embedding \
  -m LCO-Embedding-Omni-3B-2605-Q8_0.gguf \
  --pooling last \
  -p "Your text here"

Text embeddings (server)

./build/bin/llama-server \
  -m LCO-Embedding-Omni-3B-2605-Q8_0.gguf \
  --embedding --pooling last

curl -s http://localhost:8080/embeddings \
  -d '{"content": "Your text here"}'

Multimodal embeddings (vision + audio)

Requires the mmproj file:

./build/bin/llama-server \
  -m LCO-Embedding-Omni-3B-2605-Q8_0.gguf \
  --mmproj mmproj-LCO-Embedding-Omni-3B-2605-F16.gguf \
  --embedding --pooling last
# Image embedding (base64-encoded image)
curl -s http://localhost:8080/embeddings \
  -d '{"content": [{"prompt_string": "<__media__>", "multimodal_data": ["<base64-image-data>"]}]}'

# Audio embedding (base64-encoded WAV)
curl -s http://localhost:8080/embeddings \
  -d '{"content": [{"prompt_string": "<__media__>", "multimodal_data": ["<base64-audio-data>"]}]}'

JSON output (for programmatic use)

./build/bin/llama-embedding \
  -m LCO-Embedding-Omni-3B-2605-Q8_0.gguf \
  --pooling last \
  --embd-output-format json \
  -p "Your text here"

Notes

  • The --pooling last flag is required — this model uses last-token pooling, not mean pooling
  • Embedding dimensions: 2048
  • Contributions and bug reports welcome at ht-llama.cpp

Citations

LCO-Embedding

@article{xiao2025scaling,
  title={Scaling Language-Centric Omnimodal Representation Learning},
  author={Xiao, Chenghao and Chan, Hou Pong and Zhang, Hao and Xu, Weiwen and Aljunied, Mahani and Rong, Yu},
  journal={arXiv preprint arXiv:2510.11693},
  year={2025}
}

Qwen 2.5 Omni

@article{Qwen2.5-Omni,
  title={Qwen2.5-Omni Technical Report},
  author={Jin Xu and Zhifang Guo and Jinzheng He and Hangrui Hu and Ting He and Shuai Bai and Keqin Chen and Jialin Wang and Yang Fan and Kai Dang and Bin Zhang and Xiong Wang and Yunfei Chu and Junyang Lin},
  journal={arXiv preprint arXiv:2503.20215},
  year={2025}
}
Downloads last month
297
GGUF
Model size
3B params
Architecture
qwen2vl
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for marksverdhei/LCO-Embedding-Omni-3B-2605-GGUF

Quantized
(1)
this model

Collection including marksverdhei/LCO-Embedding-Omni-3B-2605-GGUF

Papers for marksverdhei/LCO-Embedding-Omni-3B-2605-GGUF

Evaluation results