Swap faces in videos by providing a source image and a target video. The application will generate a new video with the face from the source image replacing the face in the target video.
Global radar
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching https://huggingface.co/spaces/mrfakename/E2-F5-TTS
[NVIDIA ONLY] A minimal Gradio interface for Automatic Speech Recognition. Transcribe Audio in Malayalam language.
Fully offline document-to-speech converter supporting 9 TTS models across 5 engine groups: Coqui (XTTS v2, Bark, VITS, YourTTS), StyleTTS 2, Resemble Chatterbox, Tortoise TTS, Sesame CSM-1B, and OpenVoice V2. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio with voice cloning support. No API keys or cloud services required.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
FLUX.1 Krea [dev] txt2img studio (Fooocus-style, 100% local): high-aesthetic text-to-image, ESRGAN+Flux refine upscale, single-file/Civitai Flux models, LoRA, styles, Describe/Improve & Vision Mix (Ollama), Remove BG, Reframe/outpaint, Face Swap. Fork of crispz-studio. https://github.com/mikecastrodemaria/crispz-krea
A web interface for the Moondream3 vision-language model featuring image captioning, visual question answering, object detection, and object pointing.
[NVIDIA ONLY - WINDOWS ONLY] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity [LoRA support fork] https://github.com/petermg/InfiniteYou
Contribute to Wan-Video/Wan-Animate-2 development by creating an account on GitHub.
SearXNG is a free internet metasearch engine which aggregates results from various search services and databases. Users are neither tracked nor profiled.
Kokoro, KittenTTS, Higgs audio, Chatterbox/Multi, Fish-Speech, F5 & index-tts & indextts2, VoxCPM and VibeVoice in one app
Stable Diffusion web UI UX: https://github.com/anapnoe/stable-diffusion-webui-ux
Irodori-TTS is a Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
an open vision-language model by Google. PaliGemma is designed as a versatile model for transfer to a wide range of vision-language tasks such as image and short video caption, visual question answering, text reading, object detection and object segmentation https://huggingface.co/spaces/google/paligemma
A simple, high-quality image generation tool to create stunning illusions.
[CVPR2024] StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On
Download and transcribe many YouTube videos or whole playlists at once with faster-whisper. Outputs txt, srt, vtt, or json.
