Describe an image, get a 100% schema-valid Ideogram 4 JSON prompt — generated fully locally with an embedded llama.cpp (no Ollama or LM Studio required).
Global radar
New projects people are discovering or following across Pinokio.
One-click launcher for Stable Diffusion web UI (AUTOMATIC1111/stable-diffusion-webui)
Free & unlimited AI video and image watermark remover. Uses Florence-2 and LaMA to seamlessly erase watermarks and logos offline
A professional, Suno-like music generation studio for HeartLib. https://github.com/fspecii/HeartMuLa-Studio
One-click installer for Microsoft TRELLIS.2: High-quality 3D asset generation from images with PBR textures.
A simple, high-quality voice conversion tool focused on ease of use and performance.
Official inference code for SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
Text-to-Video (T2V) generation framework from Vchitect https://github.com/Vchitect/LaVie
Upload a clean 20 seconds WAV file of the vocal persona you want to mimic, type your text-to-speech prompt and hit submit! A local version of https://huggingface.co/spaces/fffiloni/instant-TTS-Bark-cloning
An all-in-one, 100% local AI video, image & music studio. Its Director mode turns a single prompt into a full music video or short film — LLM-planned, shot by shot. Built on the WanGP pipeline (Wan 2.1/2.2, LTX-2.3, Qwen, Hunyuan Video, Flux). Requires an NVIDIA GPU (6GB+ VRAM).
[NVIDIA ONLY] AllTalk-TTS is a unified UI for E5-TTS, XTTS, Vite TTS, Piper TTS, Parler TTS and RVC, based on CoquiTTS, including a finetune mode.
Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/Blizaine/Maestro
Automatically clip videos and generate captions for LoRA training using advanced vision models like Gemma-3, Qwen3-VL, and Qwen2-VL.
clone voices into different languages by using just a quick 3-second audio clip. (a local version of https://huggingface.co/spaces/coqui/xtts)
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion WebUI (based on Gradio) to make development easier, optimize resource management, and speed up inference. https://github.com/lllyasviel/stable-diffusion-webui-forge?tab=readme-ov-file
Customizing Realistic Human Photos via Stacked ID Embedding https://github.com/TencentARC/PhotoMaker
AutoGPT is a powerful tool that lets you create and run intelligent agents https://github.com/Significant-Gravitas/AutoGPT
Minimal Flux Web UI powered by Gradio & Diffusers (Flux Schnell + Flux Merged)
Remove backgrounds from videos and images with precision AI matting. Runs locally on 12GB VRAM — Windows, Linux, and macOS.
