Transcription & Audio Cleanup Tool with noise removal and stutter detection
Global radar
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Text-to-Video (T2V) generation framework from Vchitect https://github.com/Vchitect/LaVie
Contribute to MusicLang/musiclang development by creating an account on GitHub.
A Web UI for easy subtitle using whisper model (https://github.com/jhj0517/Whisper-WebUI)
Convert scanned PDFs into searchable text locally using Vision LLMs (olmOCR). 100% private, offline, and free. Features a modern Web UI & CLI.
Run a minimal IOC sweep for the March 30-31, 2026 axios npm compromise.
A web interface for the Moondream3 vision-language model featuring image captioning, visual question answering, object detection, and object pointing.
The Open-Source Elevenlabs alternative AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.

Pure C inference engine for Qwen3-TTS text-to-speech. No Python, no PyTorch — just C and BLAS. Supports 0.6B and 1.7B models, 9 voices, 10 languages.
[CVPR 2025] Learning Flow Fields in Attention for Controllable Person Image Generation
[NeurIPS 2022] Towards Robust Blind Face Restoration with Codebook Lookup Transformer