An all-in-one, 100% local AI video, image & music studio. Its Director mode turns a single prompt into a full music video or short film — LLM-planned, shot by shot. Built on the WanGP pipeline (Wan 2.1/2.2, LTX-2.3, Qwen, Hunyuan Video, Flux). Requires an NVIDIA GPU (6GB+ VRAM).
Global radar
Interpolate, Upscale, Decompress, and Denoise videos Locally on Linux/Windows/MacOS.
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface. https://github.com/comfyanonymous/ComfyUI
MiniMax H3 omni-modal video generation in ComfyUI. Text/image/video/audio in, video with native 32kHz stereo audio out (768p default, 1080p+ supported). Disk-optimized: pruned INT8 + NVFP4 weights (~63GB instead of ~290GB). NVIDIA only.
ElevenLabs at home. Multilingual TTS with Voice Design, Voice Cloning, and end-to-end LoRA fine-tuning straight from a video or podcast. Built on VoxCPM2 by OpenBMB. 30 languages incl. Russian.
[NVIDIA ONLY] Advanced Web UI for CogVideo (text to video, image to video, video to video, extend video, etc) -- Generate videos with less than 10GB VRAM
NVIDIA launcher for LingBot-Map, a streaming 3D reconstruction viewer from Robbyant.
[AMD ONLY] Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video, Flux and more. (On Windows supported by all dedicated AMD GPUs from RDNA 2 - RDNA 4)
AuraSR-v2 - An open reproduction of the GigaGAN Upscaler from fal.ai https://huggingface.co/spaces/gokaygokay/AuraSR-v2
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching https://huggingface.co/spaces/mrfakename/E2-F5-TTS
Apple Silicon video studio — text-to-video and video-to-video with LTX-Video, Wan 2.2, HunyuanVideo, and CogVideoX. Powered by PyTorch (MPS) + Diffusers.
A professional, Suno-like music generation studio for HeartLib. https://github.com/fspecii/HeartMuLa-Studio
[NVIDIA, ROCM] One app to train them all. LORA training and Model finetuning for Z-Image, Qwen Image, FLUX.1, Flux.2 Dev and Klein, Chroma, SD 1.5 - 3.5, SDXL, Würstchen-v2, Stable Cascade, PixArt-Alpha, PixArt-Sigma, Sana, Hunyuan Video and inpainting models.
One-click installer for Microsoft TRELLIS.2: High-quality 3D asset generation from images with PBR textures.
A simple, high-quality voice conversion tool focused on ease of use and performance.
1-click WanGP Launcher. Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/deepbeepmeep/Wan2GP
open source chat UI for Ollama https://github.com/ivanfioravanti/chatbot-ollama
