Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/deepbeepmeep/Wan2GP
Global radar
🌍 TranslateGemma - Google's open-source multilingual translation AI. Translate text across 55+ languages and extract/translate text from images. Powered by Gemma 3 architecture.
🔥 AI Image & Video Suite — Photo Upscale, Video Upscale, Background Removal, Image Enhancement & Tools
Local web UI for orchestrating 3D generation pipelines via ComfyUI / Tripo / Tencent. https://github.com/visualbruno/3DGenStudio
Upload a clean 20 seconds WAV file of the vocal persona you want to mimic, type your text-to-speech prompt and hit submit! A local version of https://huggingface.co/spaces/fffiloni/instant-TTS-Bark-cloning
Image generation using zai-org/GLM-Image with Gradio UI. Supports text-to-image and image-to-image generation.
This is a piano software that analyzes what chords you are playing in real time by algorithms based on music theory. This piano software supports MIDI keyboard, computer keyboard, play and analyze MIDI files and so on.
Contribute to gnipbao/whiteboard-video-engine development by creating an account on GitHub.
One-click launcher for Stable Diffusion web UI (AUTOMATIC1111/stable-diffusion-webui)
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Suno at home. Local AI music generation studio — full songs with vocals, lyrics, covers, and music videos. Built on ACE-Step 1.5 XL.
restore low-res images, restore broken images, recreate a new version of the image with a prompt https://huggingface.co/spaces/fffiloni/InstantIR
MAGNeT is a text-to-music and text-to-sound model capable of generating high-quality audio samples conditioned on text descriptions https://github.com/facebookresearch/audiocraft/blob/main/docs/MAGNET.md
All in one Gradio interface for chatterbox. Voice cloning from uploaded audio samples, automatic text processing for long content and real-time speech generation with configurable parameters. (Minimum Requirements 4GB VRAM / Recommended Requirements 8GB VRAM)
Unified AI VFX pipeline with CPE prompt engineering, storyboard canvas, and multi-node orchestrator. https://github.com/NickPittas/DirectorsConsole
clone voices into different languages by using just a quick 3-second audio clip. (a local version of https://huggingface.co/spaces/coqui/xtts)
