MiniMax H3 omni-modal video generation in ComfyUI. Text/image/video/audio in, video with native 32kHz stereo audio out (768p default, 1080p+ supported). Disk-optimized: pruned INT8 + NVFP4 weights (~63GB instead of ~290GB). NVIDIA only.
Global radar
New projects people are discovering or following across Pinokio.
Kokoro, KittenTTS, Higgs audio, Chatterbox/Multi, Fish-Speech, F5 & index-tts & indextts2, VoxCPM and VibeVoice in one app
[NVIDIA ONLY] Super Optimized Gradio UI for Wan2.1 video for GPU poor machines (5GB+ VRAM). Generate up to 12 sec videos https://github.com/deepbeepmeep/Wan2GP
[AMD ONLY] Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video, Flux and more. (On Windows supported by all dedicated AMD GPUs from RDNA 2 - RDNA 4)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching https://huggingface.co/spaces/mrfakename/E2-F5-TTS
An all-in-one, 100% local AI creative studio, director, and multi-track editor. Generate with MiniMax H3, LTX-2.5/2.3, Wan, Flux, Qwen, and more; turn an idea or song into a planned production; then finish it on the timeline. Requires an NVIDIA GPU (6GB+ VRAM).
Style Aligned Image Generation via Shared Attention https://style-aligned-gen.github.io/
Prompt Orchestrator that turns module-based game design (genre, mechanics, visuals, menus, audio) into a complete, playable HTML5 game generated by your chosen AI provider. Supports OpenAI, Gemini, Claude, Ollama, and LM Studio. Every game ships as a single self-contained HTML file.
Fast AI Video Generation per GPU poor (Wan2.1, Hunyuan, LTV). Gradio UI su http://127.0.0.1:7860
Estúdio IA para Criação de Letras, Estruturação de Prompts Suno/Udio/Mureka e Geração Musical YuE2 com Transcrição de Covers SheetSage2.
All-in-one Gradio UI for the MOSS-TTS Family: voice cloning, dialogue generation, voice design from text, and sound effects.
🎙️ Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning. High-quality text-to-speech synthesis supporting zero-shot voice cloning and streaming inference with natural emotional expression.
Uncensored deepfakes for images and videos, no training required. Advanced masking, batch processing, and face enhancement powered by InsightFace and ONNX. Supports NVIDIA (CUDA/TensorRT), AMD (DirectML/ROCm), Apple Silicon, and CPU.
Uncensored Deepfakes for images and videos without training and an easy-to-use GUI.
[NVIDIA ONLY] Generate Video Progressively. FramePack is a next-frame (next-frame-section) prediction neural network structure that generates videos progressively. https://github.com/lllyasviel/FramePack
Expressive TTS with voice cloning, prompt-driven speech synthesis built on LTX-2.3 by Resemble AI
