IQuest-Q1-FP8
FP8 blockwise quantization of IQuestLab/IQuest-Q1, prepared by ShengYan0305. Both the main model and the recursive MTP draft model are included.
The base model was developed by IQuest Research. This is a community quantization, not an official IQuest release.
Runtime: use vLLM with the official IQuest-Q1 plugin. See runtime compatibility for setup and serving details.
| Property | Value |
|---|---|
| Base model | IQuest-Q1, approximately 320B total / 15B active parameters |
| Weight quantization | FP8 E4M3, 128 × 128 blocks, RTN |
| Activations | Dynamic quantization |
| Recursive MTP | Included in mtp/, also quantized to FP8 |
| Weight files | 175 main-model shards + 1 MTP file |
| Weight size | Approximately 327 GB / 304.67 GiB |
| Validation | Tensor integrity checks and a three-request generation smoke test |
Weight file size is not an estimate of total runtime GPU memory, which also includes KV cache, activations, and runtime buffers.
Download
pip install -U huggingface_hub
hf download ShengYan0305/IQuest-Q1-FP8 --local-dir ./IQuest-Q1-FP8
Keep the mtp/ subdirectory when downloading if you intend to use recursive MTP.
Source and quantization
- Base revision:
5c21b0630586ef77d38cff1094b8cf37a417fcd8. - Quantization: round-to-nearest (RTN), FP8 E4M3, 128 x 128 weight blocks, FP32 inverse scales, dynamic activations.
- Main model and recursive MTP are both quantized; tensors excluded from quantization retain their original values. See each
config.jsonfor the exact exclusions. - Weight layout: 175 main-model Safetensors shards plus
mtp/model.safetensors. - Weight size: 327,131,664,936 bytes (approximately 327 GB / 304.67 GiB).
- Tokenizer, chat template, generation configuration, and model metadata are preserved from the source, except for the added quantization configuration.
- Recursive MTP retains the published Q/K/V weight names and uses fused
qkv_proj.weight_scale_invscales.
Validation
The quantization run verified all source shard SHA256 hashes against the Hugging Face source metadata, independently requantized all FP8 weights and scales for bitwise comparison, and checked that unquantized tensors were unchanged.
A subsequent smoke test on 2026-09-30 passed with the following configuration:
| Component | Tested configuration |
|---|---|
| GPUs | 8 × NVIDIA H20-3e |
| Tensor / expert parallelism | 8 / 8 |
| Maximum context in this test | 4,096 tokens |
| Speculative decoding | Recursive MTP, 3 speculative tokens |
Main-model and recursive-MTP CUDA graph capture passed. Three generation requests completed: Chinese self-introduction, a simple arithmetic question, and Python function generation. The reasoning parser was not enabled in that test, so responses included raw reasoning text and the </think> delimiter.
This was a smoke test, not a quality benchmark, long-context validation, or load test. Base-model benchmark scores should not be interpreted as measurements of this quantized model.
The 4,096-token setting above is the tested serving configuration, not a change to the base model's context-length configuration. Long-context quality and memory requirements have not been measured for this quantization.
Runtime compatibility
Use vLLM with the official IQuest-Q1 plugin, installed in every worker environment. Follow the plugin's installation instructions for a compatible vLLM environment.
Reference serving configuration
After installing vLLM and the official IQuest-Q1 plugin, use the following reference configuration. It reproduces the tested serving settings with the local model path and public model name substituted.
vllm serve ./IQuest-Q1-FP8 \
--served-model-name IQuest-Q1-FP8 \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--quantization fp8 \
--gpu-memory-utilization 0.65 \
--max-model-len 4096 \
--max-num-seqs 4 \
--max-num-batched-tokens 2048 \
--speculative-config '{"method":"eagle","model":"./IQuest-Q1-FP8/mtp","num_speculative_tokens":3}' \
--host 127.0.0.1 \
--port 8000
Once the server is ready:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"IQuest-Q1-FP8","messages":[{"role":"user","content":"What is 15% of 240?"}],"max_tokens":256,"temperature":0}'
License
The original IQuest-Q1 Modified MIT License and copyright notice are preserved in LICENSE. The license includes the requirement to prominently display “IQuest-Q1” in the user interface of commercial products or services using the model or its derivatives.
- Downloads last month
- 36
Model tree for ShengYan0305/IQuest-Q1-FP8
Base model
IQuestLab/IQuest-Q1