Global radar
Hyper-fast, local, high-quality TTS based on Kokoro-82M. PySide6 GUI included.
Qwen-Image-Edit instruction-based image editing studio — fork of crispz-studio
Contribute to jtydhr88/pentrado development by creating an account on GitHub.
Transform YouTube videos into stunning animated GIFs with perfectly-timed, stylized subtitles and eye-catching effects.
Contribute to facok/comfyui-krea2-controlnet development by creating an account on GitHub.
Instruction-based, identity-preserving image editing for Krea 2 in ComfyUI — nodes + workflows for the Krea 2 Identity Edit LoRA
Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
The best vocal remover application on the internet, and it's totally free and open source!
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Linux-tested Pinokio launcher for FluxRT with automatic model downloads.
🎙️ Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning. High-quality text-to-speech synthesis supporting zero-shot voice cloning and streaming inference with natural emotional expression.
Upload an image file or a PDF document, and the app will read the text contained in it. PDFs are first split into separate page images, then each page is processed to recognize the words, sending t...
Experimental demonstration for the Qwen/Qwen-Image-Edit-2511 model with lazy-loaded LoRA adapters supporting multi-image input editing. Users can upload one or more images (gallery format) and apply advanced edits such as pose transfer, anime conversion, or camera angle changes via natural language prompts. Features integrated Rerun SDK.
DreamID-V: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server, 65 tok/s Qwen 3.5 122B, Llama 3.3 70B, Gemma 4 31B. Private, offline, airgap-ready. Built for NDA / legal / healthcare workflows.
