Global radar
New projects people are discovering or following across Pinokio.
Run Qwen3-TTS text-to-speech locally on Mac (M1/M2/M3/M4). Voice cloning, voice design, custom voices. 100% offline using MLX.
Enter text and upload a reference audio to create a synthesized speech that matches the speaker's voice and chosen style. Supports English and Chinese.
Upload a source image and audio file to create a video of the image's face moving and speaking as if it were saying the audio. You can also use reference videos to enhance the animation.
Code for a proof of concept for OMR (Optical Mark Recognition) using opensource tools
基于AI的图片/视频硬字幕去除、文本水印去除,无损分辨率生成去字幕、去水印后的图片/视频文件。无需申请第三方API,本地实现。AI-based tool for removing hard-coded subtitles and text-like watermarks from videos or Pictures.
ARGO is an open-source AI Agent platform that brings Local Manus to your desktop. With one-click model downloads, seamless closed LLM integration, and offline-first RAG knowledge bases, ARGO becomes a DeepResearch powerhouse for autonomous thinking, task planning, and 100% of your data stays locally. Support Win/Mac/Docker.
Open-source, accurate and easy-to-use video speech recognition & clipping tool, LLM based AI clipping intergrated.
Contribute to cocktailpeanut/kohya_ss development by creating an account on GitHub.
[ICCV'25 Best Paper Candidate] Official Implementations for Paper: Dynamic Typography: Bringing Text to Life via Video Diffusion Prior
