[LINUX + NVIDIA ONLY] Real-time interactive world model. Drive an infinite, action-conditioned world rollout at 720p/16fps on a single desktop GPU (~19GB VRAM). https://github.com/amap-cvlab/ABot-World
6Morpheus6/stable-diffusion-webui-forgev2.0updated 19d ago
[NVIDIA ONLY] The most efficient way to run FLUX (Optimized to run even on low memory machines, as low as 3GB VRAM with 512x512 resolution) https://github.com/lllyasviel/stable-diffusion-webui-forge
Blizaine/Qwen3-TTS-MLX-WebUI-Enhancedv5.0updated 19d ago
High-quality text-to-speech with Beautiful Web UI & API, optimized for Apple Silicon using MLX. Features include Custom Voice (preset speakers), Voice Design (natural language), and Voice Cloning. With enhanced features for saving custom voices and long-form / endless TTS streaming.
BazedFrog/SongGeneration-Studiov3.7updated 21d ago
AI Song Generation with Full Style Control - Generate complete songs with lyrics, vocals, and instrumental tracks using Tencent AI Lab's SongGeneration (LeVo) model. [NVIDIA ONLY]
cocktailpeanut/stabledaw.pinokiov7.0updated 24d ago
Browser-based AI audio DAW for Stable Audio 3 with text-to-audio, inpainting, LoRA training, FFmpeg effects, waveform editing, sequencer, piano roll, and persistent library. https://github.com/gantasmo/stabledaw
Unified audio production dashboard for Volt Records - integrates StableDAW, Stable Audio 3, TASCAR, and catalog intelligence into a single command center.
Unified audio production dashboard for Volt Records - integrates StableDAW, Stable Audio 3, TASCAR, and catalog intelligence into a single command center.
Transform lyric transcriptions into karaoke-style MP4 videos. Built on Python-Lyric-Transcriber, this Gradio UI uses Whisper for transcription, an LLM for lyric edits, and Demucs for vocal separation. A fun tool for karaoke fans, though outputs may vary.