项目文件夹

文件
wehub-resource-sync eec33d25b2
Build Wheel / build (3.11) (push) Failing after 1s
Build Wheel / build (3.12) (push) Failing after 0s
pre-commit / pre-commit (push) Failing after 1s
chore: import upstream snapshot with attribution
2026-07-13 12:29:08 +08:00

5.1 KiB

Stable-Audio-Open Text-To-Audio Generation on 1x GPU

Text-to-audio recipe for Stable Audio Open with offline inference and OpenAI-compatible online serving on 1x RTX 4090 24GB.

Summary

  • Vendor: Stability AI
  • Model: stabilityai/stable-audio-open-1.0
  • Task: Text-to-audio generation
  • Mode: Offline inference and online serving
  • Maintainer: Community

When to use this recipe

Use this recipe when you want to run Stable Audio Open on a single RTX 4090 24GB GPU for music or sound-effect generation. The recipe covers a 10-second offline validation sample with TeaCache and online serving through the /v1/audio/generate endpoint.

References

Hardware Support

GPU

1x RTX 4090 24GB

Environment
  • OS: Ubuntu 22.04.5
  • Python: 3.12
  • GPU: NVIDIA GeForce RTX 4090, 24564 MiB VRAM
  • Driver / runtime: NVIDIA driver 595.80, CUDA-capable runtime matching the repository build
  • vLLM version: 0.22.0
  • vLLM-Omni version: source checkout
  • PyTorch: 2.11.0+cu130
  • Model path used in the commands below: /path/to/stable-audio-open-1.0

Stable Audio Open is a gated Hugging Face model. Accept the model license on the Hugging Face model card before downloading the checkpoint.

hf auth login

hf download stabilityai/stable-audio-open-1.0 \
  --local-dir /path/to/stable-audio-open-1.0
Commands

Run a 10-second offline validation sample from the repository root:

python examples/offline_inference/text_to_audio/text_to_audio.py \
  --model /path/to/stable-audio-open-1.0 \
  --prompt "A gentle piano melody with soft room ambience" \
  --negative-prompt "Low quality, distorted, noisy" \
  --seed 42 \
  --guidance-scale 7.0 \
  --audio-length 10.0 \
  --num-inference-steps 50 \
  --cache-backend tea_cache \
  --output examples/offline_inference/text_to_audio/stable_audio_10s.wav

Start the online serving endpoint:

vllm-omni serve /path/to/stable-audio-open-1.0 \
  --host 0.0.0.0 \
  --port 8091 \
  --gpu-memory-utilization 0.9 \
  --trust-remote-code \
  --enforce-eager \
  --omni

Generate a 10-second WAV file from the repository root in another terminal:

curl http://localhost:8091/health

curl -X POST http://localhost:8091/v1/audio/generate \
  -H "Content-Type: application/json" \
  -d '{
    "input": "A gentle piano melody with soft room ambience",
    "audio_length": 10.0,
    "num_inference_steps": 50,
    "guidance_scale": 7.0,
    "negative_prompt": "Low quality, distorted, noisy",
    "seed": 42,
    "response_format": "wav"
  }' \
  --output examples/online_serving/stable_audio/piano_10s.wav
Verification

Check that:

  • The offline command writes a valid WAV file.
  • The server responds on http://localhost:8091/health.
  • The online request writes a valid WAV file.
  • The generated audio sample rate is 44.1 kHz.
  • The generated duration is approximately 10 seconds.
  • Peak sampled GPU memory is within the 24GB RTX 4090 budget. In the validated run, offline and online generation each peaked at about 12.6 GiB.

Validate the offline outputs:

ls -lh examples/offline_inference/text_to_audio/stable_audio_10s.wav

python - <<'PY'
import soundfile as sf

path = "examples/offline_inference/text_to_audio/stable_audio_10s.wav"
audio, sample_rate = sf.read(path)
print("sample_rate:", sample_rate)
print("shape:", audio.shape)
print("duration:", len(audio) / sample_rate)
PY

Validate the online output:

ls -lh examples/online_serving/stable_audio/piano_10s.wav

python - <<'PY'
import soundfile as sf

path = "examples/online_serving/stable_audio/piano_10s.wav"
audio, sample_rate = sf.read(path)
print("sample_rate:", sample_rate)
print("shape:", audio.shape)
print("duration:", len(audio) / sample_rate)
PY
Notes
  • stable-audio-open-1.0 can generate up to about 47 seconds of 44.1 kHz stereo audio. This recipe validates 10-second WAV outputs.
  • --cache-backend tea_cache is supported and was used for the 10-second offline validation command.
  • The model is gated on Hugging Face and requires license acceptance before download.
  • If online serving fails while importing torchaudio, make sure the torchaudio wheel matches the installed PyTorch and CUDA build. The validated environment used torch==2.11.0+cu130 and torchaudio==2.11.0+cu130.
  • The NIXL is not available, GLOO_SOCKET_IFNAME, and torchsde boundary warnings observed during validation did not prevent successful generation.
  • This recipe was validated on one RTX 4090 24GB GPU only. Other GPU counts, ROCm, XPU, and NPU setups are not covered here.
  • Long generations, higher inference-step counts, and non-WAV response formats were not benchmarked in this recipe.