Generate stunning illusion artwork with StableDiffusion (A space by @angrypenguinPNGAP - created with Monster Labs QR ControlNet.
Global radar
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
A Gradio WebUI for voice cloning powered by Qwen3-TTS. Provide reference audio via YouTube URL, microphone recording, or file upload — the app transcribes it with Whisper and clones the voice for TTS. Save/load voice profiles for reuse. Optimized for Apple Silicon (MPS).
Contribute to dotneet/image-to-3d development by creating an account on GitHub.
NSFW expansion module for ComfyUI Photoreal Prompt Builder
One-click model liberation + chat playground
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics,
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
Remove long silences, ghost sounds and glitches from TTS audiobooks using Silero VAD
Frame is an AI-powered, open-source vibe video editor, offering a Professional VIDEO cuting alternative for creators. With Cursor-like interaction, it automates editing, enhances videos, and delivers a seamless vibe video editing experience.
Demo showcasing ~real-time Latent Consistency Model pipeline with Diffusers and a MJPEG stream server (https://github.com/radames/Real-Time-Latent-Consistency-Model)
Enter or upload plain text or an SRT file, pick a voice and tweak speed, pitch, and volume, then get an MP3 audio file. You can also produce a synchronized subtitle (.srt) file, and the app offers ...
