Full-featured DAW, DJ, and VJ app interoperable w/Ableton, Reaper, Resolume & more. Stable Audio 3, Magenta RT2, Lyria 3 Pro, Suno API, Chimera track fusion, Demucs stems, MIDI generate/notate, native Audima Labs Sway support, img > spectrogram > music, drawing > music, VST3 & .gan plugins, automix & key-lock, GLSL shaders, volumetric video, Quest 3 XR interface, MIDI auto-map, RAG assistant, asset library with example projects
Global radar
New projects people are discovering or following across Pinokio.
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
Local video manager with AI autotagging and optional SeedVR2 upscale, powered by VLC
Enhanced background remove and replace app built around BRIA-RMBG-2.0 https://huggingface.co/briaai/RMBG-2.0
The open-source ElevenLabs alternative. Local voice cloning, video dubbing, and real-time dictation — 646 languages, no API keys.
Stable Diffusion web UI UX: https://github.com/anapnoe/stable-diffusion-webui-ux
Image inpainting tool powered by SOTA AI models. Remove any unwanted object, defect, or even people from your pictures, and replace (powered by stable diffusion) anything in your pictures. https://www.iopaint.com/
A fully local, cross-platform audio visualizer editor. Create reactive music videos with layered graphics, AI-transcribed lyrics, and frame-perfect MP4 exports — all running in your browser
YUE2 // GROOVE is the latest music studio built on the open-source Yue2 model and its inference stack. Generate high-quality full songs from style and lyrics with an editable score plan — powered by the latest YuE model — cover from audio with SheetSage2 and MERT2, refine and compare edits, and keep your works in a reusable, easy-to-manage library. You get high-quality creation with real creative control. Hardware: an NVIDIA GPU with 24 GB VRAM on Linux (YuE2's recommended setup, validated end-to-end on NVIDIA L4 hosts; a 16 GB memory budget runs everything the app can produce, and 12 GB runs the unquantized model at CFG 1.0 or for shorter songs) or an Apple Silicon Mac with 32 GB+ unified memory (where this app is developed and tested). Windows is best effort: install, launch, Cover and a full-length song verified on Windows 11 (RTX 2070, 8 GB); song generation there runs through the GGUF engine. Cards under 16 GB (and Windows) get the optional GGUF engine: the same model through yue2.cpp with an 8-bit backbone, 8.2 GB peak for a full song, measured indistinguishable from the reference rendering in a blind ABX, a different take for the same seed; the reference PyTorch configuration stays the default wherever it fits.
An all-in-one, 100% local AI video, image & music studio. Its Director mode turns a single prompt into a full music video or short film — LLM-planned, shot by shot. Built on the WanGP pipeline (Wan 2.1/2.2, LTX-2.3, Qwen, Hunyuan Video, Flux). Requires an NVIDIA GPU (6GB+ VRAM).
create a story by generating consistent images https://github.com/HVision-NKU/StoryDiffusion
A gradio web UI for running Large Language Models like LLaMA, llama.cpp, GPT-J, Pythia, OPT, and GALACTICA.
Fast Image generator using Latent consistency models https://replicate.com/blog/run-latent-consistency-model-on-mac
Zero-shot multilingual voice cloning with cross-lingual synthesis, disentangled emotion control, pronunciation guidance, and speaking-speed control.
open source chat UI for Ollama https://github.com/ivanfioravanti/chatbot-ollama
A unified image generation model that you can use to perform various tasks, including but not limited to text-to-image generation, subject-driven generation, Identity-Preserving Generation, and image-conditioned generation. https://huggingface.co/spaces/Shitao/OmniGen
Convert your videos to densepose and use it on MagicAnimate https://github.com/Flode-Labs/vid2densepose
Janus Pro 7B is a powerful multimodal AI model designed for advanced image understanding and text-to-image generation.
Real-time network traffic visualizer — privacy-first, all data stays local. / Visualiseur de trafic réseau en temps réel — vie privée d'abord, toutes les données restent locales.
