An open source urdu/hindi text-to-speech system with voice cloning capabilities
Global radar
A fine-tuned SpeechT5 Urdu TTS model with voice cloning that converts both Urdu and Roman Urdu text into natural speech. Trained on diverse Urdu and Zia Mohiuddin recordings, it offers expressive, speaker-specific synthesis with a FastAPI demo for easy testing.
Upload two images to swap the face from the first image into the second. Optionally enhance the face in the resulting image.
Official implementation of "Sonic: Shifting Focus to Global Audio Perception in Portrait Animation"
ComfyUI-cpu is a trimmed down version of ComfyUI that uses the cpu only. AI text to image generation with no GPU. it can be used on CPU only servers, laptops, and older computers.
JoyCaption is an image captioning Visual Language Model (VLM) being built from the ground up as a free, open, and uncensored model for the community to use in training Diffusion models.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Free AI Creative Tools - www.humangen.ai
A ComfyUI custom node for 3D camera angle control. Provides an interactive Three.js viewport to adjust camera angles and outputs formatted prompt strings for multi-angle image generation.
Upload an image to create a 3D model. Adjust settings like background removal and foreground ratio to refine the output. Get a 3D model as a result.
Official Python toolkit for the Qwen3-ASR API. Parallel high‑throughput calls, robust long‑audio transcription, multi‑sample‑rate support.
Extension for Automatic1111's Stable Diffusion WebUI, using Microsoft DirectML to deliver high performance result on any Windows GPU.
The first automated AI tool discovery platform using the revolutionary .awesome-ai.md standard. Real-time GitHub scanning, automated curation, and live leaderboards for AI tools.
