0
Generation failed. Please try again.: Gradio pipeline not available after 2 minutes
What happened?
when trying to generate it fails and says "Generation failed. Please try again.: Gradio pipeline not available after 2 minutes"
Steps to reproduce
- Try to generate any audio.
Your system (OS / GPU / RAM / VRAM / etc.)
Windows 11, 32GB RAM, RTX 4070ti Super with 16gb vram.
All models downloaded.
Logs / full error output
Auto-enabling CPU offload (GPU 16.0GB < 20.0GB threshold)
Output directory: D:/Pinokio/api/ACE-Step-Studio-pinokio.git/app/ACE-Step-1.5/gradio_outputs
[Gradio] LM Backend: vllm (requested: vllm)
Initializing service from command line...
[Gradio] Initializing DiT model: marcorez8/acestep-v15-xl-turbo-bf16 on auto...
2026-09-25 15:35:45.245 | INFO | acestep.core.generation.handler.init_service_loader:_load_main_model_from_checkpoint:181 - [initialize_service] Attempting to load model with attention implementation: flash_attention_2
[Pipeline] Startup timeout (300000ms) — pipeline still loading, Express will start anyway
2026-09-25 15:36:09.403 | INFO | acestep.core.generation.handler.init_service_loader:_load_main_model_from_checkpoint:207 - [initialize_service] Keeping main model on cuda (persistent)
2026-09-25 15:38:21.628 | INFO | acestep.core.generation.handler.init_service_loader:_apply_dit_quantization:120 - [initialize_service] DiT quantized with: int8_weight_only
There are modules in AutoencoderOobleck that should be kept in float32: []. Casting directly with `to()` can lead to inconsistent results; set `torch_dtype` in `from_pretrained()` instead to keep these modules in float32.```
Replies (0)
Up to 10 files, 25MB each. Images are optimized; GIFs -> MP4; videos 720p (max 120s).
