We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Global radar
Foundational Models for State-of-the-Art Speech and Text Translation
Select a portrait, click to move the head around (please use your own space / GPU!)
⚡️ Blazing-fast batch subtitle translation for SRT/ASS/VTT/LRC — 70+ languages, AI-powered 批量字幕翻译
Automatically generate, translate, and overlay subtitles for any video.
AnimateDiff for AUTOMATIC1111 Stable Diffusion WebUI - continue-revolution/sd-webui-animatediff
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
Run Qwen3-TTS text-to-speech locally on Mac (M1/M2/M3/M4). Voice cloning, voice design, custom voices. 100% offline using MLX.
Enter text and upload a reference audio to create a synthesized speech that matches the speaker's voice and chosen style. Supports English and Chinese.
Upload a source image and audio file to create a video of the image's face moving and speaking as if it were saying the audio. You can also use reference videos to enhance the animation.
Code for a proof of concept for OMR (Optical Mark Recognition) using opensource tools
