fish-s2-pro
Fish Audio S2 Pro, a 4.5B dual-autoregressive text-to-speech model (Qwen3-based semantic decoder, 400M codebook decoder, 44.1 kHz DAC codec) on Tenstorrent Blackhole, with a fish-speech-compatible /v1/tts API (streaming WAV, voice cloning via references) and an OpenAI-compatible /v1/audio/speech endpoint. Model code and weights are under the Fish Audio Research License (non-commercial).
Runs on p150 or p300 or p300x2 โ see the serve profiles below.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
At a glance
| Hardware | p150, p300, p300x2 |
Quickstart
tt-model pull jashansinghTT/s2-pro-blackhole --with-weights
tt-model serve jashansinghTT/s2-pro-blackhole
pull --with-weights downloads the Docker image and the fishaudio/s2-pro weights at 1de9996b6be38b745688de084d87a5633f714e4e (into your HF cache; they are not in the image). serve starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Try it
curl -s localhost:20000/v1/health
curl -s localhost:20000/v1/tts -H 'Content-Type: application/json' \
-d '{"text":"Hello from Tenstorrent.","format":"wav"}' -o hello.wav
curl -s localhost:20000/v1/audio/speech -H 'Content-Type: application/json' \
-d '{"model":"fishaudio/s2-pro","input":"Hello.","voice":"default","response_format":"mp3"}' -o hello.mp3
Voice cloning: POST /v1/references/add (multipart id / audio / text), then pass "reference_id" in /v1/tts.
Serve profiles
One image serves every profile below; pick one with --profile.
| profile | hardware | mesh |
|---|---|---|
p150 (default) |
p150 | P150 |
p300 |
p300 | P300 |
p300x2 |
p300x2 | P300x2 |
Using it
This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API โ its request and response shapes are the model's own. See the author's notes above for the payload it expects.
Expected performance
Measured on one p150 chip (profile p150), seeded requests to this package's server:
| metric | value |
|---|---|
| real-time factor, non-streaming (seconds of compute per second of audio) | ~1.8 (6.3 s clip in 11.1 s; 36 s clip in 67 s) |
time to first audio, streaming /v1/tts |
~0.4 s |
| real-time factor, streaming | ~1.9 |
| intelligibility (Whisper ASR word error rate on the test clips) | 0.0 |
| accuracy vs the PyTorch reference (semantic decoder output PCC) | >= 0.998 on 1, 2 and 4 chips |
Limitations
- Slower than real time: about 1.8 s of compute per second of audio on one chip.
- The audio codec runs on the host CPU, not on the Tenstorrent device.
- Tested on p150, p300 and p300x2 only. Speech quality was checked on English text only.
- Fish Audio Research License: non-commercial use only.
Feedback
Questions or problems with this package: open a discussion at https://huggingface.co/jashansinghTT/s2-pro-blackhole/discussions โ that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.
Provenance
The exact sources the image was built from โ code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | 810e9006bd03d6f5f120a41cfa4758dc7fa1cfe1 (dirty tree โ the image includes uncommitted changes) |
code/ digest |
f29564967e9fcf61 (sha256, first 16 hex digits) |
| built | 2026-09-30T16:30:39+00:00 by tt-model 0.1.0 |
Model tree for jashansinghTT/s2-pro-blackhole
Base model
fishaudio/s2-pro