IQuest-Q1-FP8

FP8 blockwise quantization of IQuestLab/IQuest-Q1, prepared by ShengYan0305. Both the main model and the recursive MTP draft model are included.

The base model was developed by IQuest Research. This is a community quantization, not an official IQuest release.

Runtime: use vLLM with the official IQuest-Q1 plugin. See runtime compatibility for setup and serving details.

Property Value
Base model IQuest-Q1, approximately 320B total / 15B active parameters
Weight quantization FP8 E4M3, 128 × 128 blocks, RTN
Activations Dynamic quantization
Recursive MTP Included in mtp/, also quantized to FP8
Weight files 175 main-model shards + 1 MTP file
Weight size Approximately 327 GB / 304.67 GiB
Validation Tensor integrity checks and a three-request generation smoke test

Weight file size is not an estimate of total runtime GPU memory, which also includes KV cache, activations, and runtime buffers.

Download

pip install -U huggingface_hub
hf download ShengYan0305/IQuest-Q1-FP8 --local-dir ./IQuest-Q1-FP8

Keep the mtp/ subdirectory when downloading if you intend to use recursive MTP.

Source and quantization

  • Base revision: 5c21b0630586ef77d38cff1094b8cf37a417fcd8.
  • Quantization: round-to-nearest (RTN), FP8 E4M3, 128 x 128 weight blocks, FP32 inverse scales, dynamic activations.
  • Main model and recursive MTP are both quantized; tensors excluded from quantization retain their original values. See each config.json for the exact exclusions.
  • Weight layout: 175 main-model Safetensors shards plus mtp/model.safetensors.
  • Weight size: 327,131,664,936 bytes (approximately 327 GB / 304.67 GiB).
  • Tokenizer, chat template, generation configuration, and model metadata are preserved from the source, except for the added quantization configuration.
  • Recursive MTP retains the published Q/K/V weight names and uses fused qkv_proj.weight_scale_inv scales.

Validation

The quantization run verified all source shard SHA256 hashes against the Hugging Face source metadata, independently requantized all FP8 weights and scales for bitwise comparison, and checked that unquantized tensors were unchanged.

A subsequent smoke test on 2026-09-30 passed with the following configuration:

Component Tested configuration
GPUs 8 × NVIDIA H20-3e
Tensor / expert parallelism 8 / 8
Maximum context in this test 4,096 tokens
Speculative decoding Recursive MTP, 3 speculative tokens

Main-model and recursive-MTP CUDA graph capture passed. Three generation requests completed: Chinese self-introduction, a simple arithmetic question, and Python function generation. The reasoning parser was not enabled in that test, so responses included raw reasoning text and the </think> delimiter.

This was a smoke test, not a quality benchmark, long-context validation, or load test. Base-model benchmark scores should not be interpreted as measurements of this quantized model.

The 4,096-token setting above is the tested serving configuration, not a change to the base model's context-length configuration. Long-context quality and memory requirements have not been measured for this quantization.

Runtime compatibility

Use vLLM with the official IQuest-Q1 plugin, installed in every worker environment. Follow the plugin's installation instructions for a compatible vLLM environment.

Reference serving configuration

After installing vLLM and the official IQuest-Q1 plugin, use the following reference configuration. It reproduces the tested serving settings with the local model path and public model name substituted.

vllm serve ./IQuest-Q1-FP8 \
  --served-model-name IQuest-Q1-FP8 \
  --tensor-parallel-size 8 \
  --enable-expert-parallel \
  --quantization fp8 \
  --gpu-memory-utilization 0.65 \
  --max-model-len 4096 \
  --max-num-seqs 4 \
  --max-num-batched-tokens 2048 \
  --speculative-config '{"method":"eagle","model":"./IQuest-Q1-FP8/mtp","num_speculative_tokens":3}' \
  --host 127.0.0.1 \
  --port 8000

Once the server is ready:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"IQuest-Q1-FP8","messages":[{"role":"user","content":"What is 15% of 240?"}],"max_tokens":256,"temperature":0}'

License

The original IQuest-Q1 Modified MIT License and copyright notice are preserved in LICENSE. The license includes the requirement to prominently display “IQuest-Q1” in the user interface of commercial products or services using the model or its derivatives.

Downloads last month
25
Safetensors
Model size
320B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xhxlb/IQuest-Q1-FP8

Quantized
(4)
this model

Collection including xhxlb/IQuest-Q1-FP8