fish-s2-pro

Fish Audio S2 Pro, a 4.5B dual-autoregressive text-to-speech model (Qwen3-based semantic decoder, 400M codebook decoder, 44.1 kHz DAC codec) on Tenstorrent Blackhole, with a fish-speech-compatible /v1/tts API (streaming WAV, voice cloning via references) and an OpenAI-compatible /v1/audio/speech endpoint. Model code and weights are under the Fish Audio Research License (non-commercial).

Runs on p150 or p300 or p300x2 โ€” see the serve profiles below.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

At a glance

Hardware p150, p300, p300x2

Quickstart

tt-model pull  jashansinghTT/s2-pro-blackhole --with-weights
tt-model serve jashansinghTT/s2-pro-blackhole

pull --with-weights downloads the Docker image and the fishaudio/s2-pro weights at 1de9996b6be38b745688de084d87a5633f714e4e (into your HF cache; they are not in the image). serve starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Try it

curl -s localhost:20000/v1/health
curl -s localhost:20000/v1/tts -H 'Content-Type: application/json' \
  -d '{"text":"Hello from Tenstorrent.","format":"wav"}' -o hello.wav
curl -s localhost:20000/v1/audio/speech -H 'Content-Type: application/json' \
  -d '{"model":"fishaudio/s2-pro","input":"Hello.","voice":"default","response_format":"mp3"}' -o hello.mp3

Voice cloning: POST /v1/references/add (multipart id / audio / text), then pass "reference_id" in /v1/tts.

Serve profiles

One image serves every profile below; pick one with --profile.

profile hardware mesh
p150 (default) p150 P150
p300 p300 P300
p300x2 p300x2 P300x2

Using it

This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API โ€” its request and response shapes are the model's own. See the author's notes above for the payload it expects.

Expected performance

Measured on one p150 chip (profile p150), seeded requests to this package's server:

metric value
real-time factor, non-streaming (seconds of compute per second of audio) ~1.8 (6.3 s clip in 11.1 s; 36 s clip in 67 s)
time to first audio, streaming /v1/tts ~0.4 s
real-time factor, streaming ~1.9
intelligibility (Whisper ASR word error rate on the test clips) 0.0
accuracy vs the PyTorch reference (semantic decoder output PCC) >= 0.998 on 1, 2 and 4 chips

Limitations

  • Slower than real time: about 1.8 s of compute per second of audio on one chip.
  • The audio codec runs on the host CPU, not on the Tenstorrent device.
  • Tested on p150, p300 and p300x2 only. Speech quality was checked on English text only.
  • Fish Audio Research License: non-commercial use only.

Feedback

Questions or problems with this package: open a discussion at https://huggingface.co/jashansinghTT/s2-pro-blackhole/discussions โ€” that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.

Provenance

The exact sources the image was built from โ€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal 810e9006bd03d6f5f120a41cfa4758dc7fa1cfe1 (dirty tree โ€” the image includes uncommitted changes)
code/ digest f29564967e9fcf61 (sha256, first 16 hex digits)
built 2026-09-30T16:30:39+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jashansinghTT/s2-pro-blackhole

Base model

fishaudio/s2-pro
Finetuned
(13)
this model