Global radar
Python wrapper for ByteDance's Seedance 2.5 API — Text-to-Video, Image-to-Video, realistic human faces, native 4K, consistent character generation.
Janus Pro 7B is a powerful multimodal AI model designed for advanced image understanding and text-to-image generation.
Apple Silicon local LLM chat via MLX — OpenAI-compatible API, runs as a background service. Catalog of ~34 mlx-community quantized chat models across 14 families (Llama, Qwen, Qwen3, Mistral, Ministral, Gemma 4/3/2, Phi, Phi-4, DeepSeek, Devstral, LFM, Nemotron). Includes RAM-fit filter, search, and capability tags (starter/code/reasoning).
Zero-shot multilingual voice cloning with cross-lingual synthesis, disentangled emotion control, pronunciation guidance, and speaking-speed control.
[NVIDIA ONLY] Stable Video Diffusion Streamlit App. Currently supports Nvidia GPU machines only.
OpenAI-compatible Speech-to-Text and Text-to-Speech server. Powered by Faster-Whisper, Kokoro, and Piper.
An AI audiobook generator built on Qwen3-TTS. Annotate your book with an LLM, assign voices, per-line style instructions for delivery, clone voices from reference audio, design new voices from text descriptions, and export to MP3 or Audacity multi-track projects
Local GPU-accelerated music video generator: Gradio UI, analysis, SDXL backgrounds, NVENC output.
Turn any image into a video! (Web UI created by fffiloni: https://huggingface.co/spaces/fffiloni/MS-Image2Video)
Improving Diffusion Models for Authentic Virtual Try-on in the Wild https://huggingface.co/spaces/yisol/IDM-VTON
Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/deepbeepmeep/Wan2GP
