Commit ·
a36d000
1
Parent(s): bb055dd
Update model card for 5090 optimized Fish S2
Browse files- NOTICE.md +14 -0
- README.md +112 -1
- assets/rtx5090-benchmark-bars.svg +111 -0
NOTICE.md
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
This repository documents an RTX 5090 local serving profile for Fish Audio S2-Pro.
|
| 2 |
+
|
| 3 |
+
It does not redistribute Fish Audio model weights. Users must obtain the base model
|
| 4 |
+
from the upstream Hugging Face repository:
|
| 5 |
+
|
| 6 |
+
https://huggingface.co/fishaudio/s2-pro
|
| 7 |
+
|
| 8 |
+
Fish Audio attribution:
|
| 9 |
+
|
| 10 |
+
This model is licensed under the Fish Audio Research License, Copyright (c) 39 AI,
|
| 11 |
+
INC. All Rights Reserved.
|
| 12 |
+
|
| 13 |
+
Commercial use of Fish Audio S2-Pro or derivative works requires a separate license
|
| 14 |
+
from Fish Audio.
|
README.md
CHANGED
|
@@ -1,3 +1,114 @@
|
|
| 1 |
---
|
| 2 |
-
license:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: fish-audio-research-license
|
| 4 |
+
license_link: https://huggingface.co/fishaudio/s2-pro/blob/main/LICENSE.md
|
| 5 |
+
base_model: fishaudio/s2-pro
|
| 6 |
+
pipeline_tag: text-to-speech
|
| 7 |
+
library_name: sglang
|
| 8 |
+
tags:
|
| 9 |
+
- text-to-speech
|
| 10 |
+
- fish-audio
|
| 11 |
+
- s2-pro
|
| 12 |
+
- sglang
|
| 13 |
+
- realtime
|
| 14 |
+
- rtx-5090
|
| 15 |
+
- no-weights
|
| 16 |
+
inference: false
|
| 17 |
---
|
| 18 |
+
|
| 19 |
+
# Fish Audio S2-Pro Realtime Optimized for RTX 5090
|
| 20 |
+
|
| 21 |
+
This repository is a model-card-only release for an RTX 5090 local serving profile for
|
| 22 |
+
[Fish Audio S2-Pro](https://huggingface.co/fishaudio/s2-pro).
|
| 23 |
+
|
| 24 |
+
It does **not** redistribute Fish Audio S2-Pro weights. The underlying model files are
|
| 25 |
+
the official upstream `fishaudio/s2-pro` checkpoint files, so users should accept the
|
| 26 |
+
upstream license and download the weights directly from Fish Audio.
|
| 27 |
+
|
| 28 |
+
## What This Is
|
| 29 |
+
|
| 30 |
+
This release documents a local realtime optimization profile:
|
| 31 |
+
|
| 32 |
+
- Base model: `fishaudio/s2-pro`
|
| 33 |
+
- Target GPU: NVIDIA GeForce RTX 5090, 32 GB VRAM
|
| 34 |
+
- Serving path: Docker-backed SGLang Omni
|
| 35 |
+
- Intended use: single-user realtime local TTS
|
| 36 |
+
- Example workstation split: RTX 5090 handles TTS; a separate RTX 3090 can handle realtime STT/ASR
|
| 37 |
+
|
| 38 |
+
No fine-tuning, quantization, or checkpoint conversion is claimed here. The optimization
|
| 39 |
+
work is in the serving/runtime profile around the official S2-Pro weights.
|
| 40 |
+
|
| 41 |
+
## Why There Are No Weights Here
|
| 42 |
+
|
| 43 |
+
The local checkpoint files match the official upstream `fishaudio/s2-pro` release. To
|
| 44 |
+
avoid duplicating gated model files or confusing the license boundary, this repository
|
| 45 |
+
does not upload:
|
| 46 |
+
|
| 47 |
+
- `model-00001-of-00002.safetensors`
|
| 48 |
+
- `model-00002-of-00002.safetensors`
|
| 49 |
+
- `codec.pth`
|
| 50 |
+
- tokenizer/config files from the upstream model
|
| 51 |
+
|
| 52 |
+
Download those files from:
|
| 53 |
+
|
| 54 |
+
```text
|
| 55 |
+
https://huggingface.co/fishaudio/s2-pro
|
| 56 |
+
```
|
| 57 |
+
|
| 58 |
+
## Local Benchmark Summary
|
| 59 |
+
|
| 60 |
+
Measured locally on an RTX 5090. Lower is better.
|
| 61 |
+
|
| 62 |
+

|
| 63 |
+
|
| 64 |
+
| Metric | Before: Python Fish server | After: SGLang Omni + cached reference | Change |
|
| 65 |
+
|---|---:|---:|---:|
|
| 66 |
+
| First audio | 25.1s | 0.36s | 98.6% lower latency |
|
| 67 |
+
| Total request time | 25.1s | 2.10s | 91.6% lower latency |
|
| 68 |
+
| Estimated RTF | 5.51 | 0.48 | 91.3% lower RTF |
|
| 69 |
+
|
| 70 |
+
Benchmark notes:
|
| 71 |
+
|
| 72 |
+
- Baseline was the local Python Fish server path.
|
| 73 |
+
- Optimized path was the Docker-backed SGLang Omni path with a warm server and cached reference voice.
|
| 74 |
+
- Warm cached SGLang generation was below realtime for the measured short live TTS sample.
|
| 75 |
+
- The chart is a local serving benchmark, not an upstream Fish Audio benchmark claim.
|
| 76 |
+
|
| 77 |
+
## Optimization Profile
|
| 78 |
+
|
| 79 |
+
The measured realtime profile used:
|
| 80 |
+
|
| 81 |
+
- SGLang Omni instead of the eager Python Fish server path.
|
| 82 |
+
- The Docker image's pinned Torch/SGLang/FlashInfer stack.
|
| 83 |
+
- SGLang CUDA graph replay enabled.
|
| 84 |
+
- RTX 5090 / SM120-safe Fish audio-decoder path by disabling the incompatible `sgl-kernel` KV-cache flash-attention path.
|
| 85 |
+
- Graph-safe fixed-cache SDPA fallback for the Fish audio decoder.
|
| 86 |
+
- `flashinfer` text attention backend.
|
| 87 |
+
- Single-user live memory profile:
|
| 88 |
+
- `mem_fraction_static=0.50`
|
| 89 |
+
- `chunked_prefill_size=2048`
|
| 90 |
+
- `max_running_requests=4`
|
| 91 |
+
- Preloaded/cached reference VQ codes for repeated voice-reference requests.
|
| 92 |
+
- Docker model/runtime volumes to avoid repeated slow checkpoint reads through Windows `/mnt/d` bind mounts.
|
| 93 |
+
|
| 94 |
+
Measured local VRAM after the tuned live restart was about 24.6 GB on the RTX 5090,
|
| 95 |
+
down from an earlier near-full 32.2 GB SGLang container reservation.
|
| 96 |
+
|
| 97 |
+
## License
|
| 98 |
+
|
| 99 |
+
The base model is governed by the
|
| 100 |
+
[Fish Audio Research License](https://huggingface.co/fishaudio/s2-pro/blob/main/LICENSE.md).
|
| 101 |
+
Research and non-commercial use are permitted by Fish Audio under that license.
|
| 102 |
+
Commercial use requires a separate license from Fish Audio.
|
| 103 |
+
|
| 104 |
+
This repository does not grant additional rights to the Fish Audio model weights.
|
| 105 |
+
|
| 106 |
+
## Attribution
|
| 107 |
+
|
| 108 |
+
Built with Fish Audio S2-Pro. Fish Audio S2-Pro is developed by Fish Audio / 39 AI, INC.
|
| 109 |
+
|
| 110 |
+
Upstream model:
|
| 111 |
+
|
| 112 |
+
```text
|
| 113 |
+
fishaudio/s2-pro
|
| 114 |
+
```
|
assets/rtx5090-benchmark-bars.svg
ADDED
|
|