Instructions to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Use Docker
docker model run hf.co/notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
- Ollama
How to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with Ollama:
ollama run hf.co/notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with Docker Model Runner:
docker model run hf.co/notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
- Lemonade
How to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Step-3.7-Flash-Q4_K_M-MTP-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "notSnix/Step-3.7-Flash-Q4_K_M-MTP-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Difficulty loading due to the naming of layers (`blk.45` vs `blk.0`)
Very interested to try out the MTP draft! However, I was unable to run using the llama.cpp commit you referenced. It seems like llama.cpp wanted a draft model that started with layer 0, but you've named it after the layer where it belongs when MTP is merged.
Would it work to simply rename the layers in the draft models?
The uploaded draft files intentionally keep the original tail-layer
names (blk.45-blk.47) because the tested llama.cpp Step MTP loader treats
the draft as the MTP tail of a 48-layer model (nextn_predict_layers=3), not as
a normal standalone draft model.
If llama.cpp is asking for blk.0, it is likely using the normal speculative
draft path instead of the Step MTP path. make sure the command includes--spec-type draft-mtp and is built from the referenced/current commit. A simple
tensor rename is probably not gonna work, because the metadata still saysblock_count=48 and nextn_predict_layers=3. This is just how the Step model is
Or just ask codex/claude to figure it out and fix it
ive updated the repo with better instructions
Amazing. I did attempt to rebuild the MTP draft model using the gguf python library but it was not as simple as renaming the layers; the geometry was wrong. Thanks for taking a look.
Okay, so that patch is non-trivial. I've watched how PRs go with llama.cpp and I suspect this isn't getting merged without some feedback from the core devs. (I'm not one).
This is really cool work, though ... have you already got a PR going for this?
FYI, this PR was just merged and your MTP models work just fine now when compiling the master llama.cpp branch.