[AMD ONLY] Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video, Flux and more. (On Windows supported by all dedicated AMD GPUs from RDNA 2 - RDNA 4)
Build, curate, caption and clean LoRA training datasets in your browser, then train, compare and iterate — local Flask app, your files stay on your disk. https://github.com/perfectgf/lora-dataset-studio
Fully offline speech-to-text transcription with 11 local AI models. Generates styled subtitles (SRT, ASS) and burns them directly onto video. Supports Whisper, Parakeet, Canary, Moonshine, SenseVoice, Vosk, and more. No API keys required.
Fully offline document-to-speech converter using Tortoise TTS. High-quality autoregressive synthesis with voice cloning. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
Fully offline document-to-speech converter using Sesame CSM-1B. Conversational speech model with Llama backbone and voice cloning. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
Fully offline document-to-speech converter using OpenVoice V2. Tone color conversion and voice cloning via MeloTTS. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
Fully offline document-to-speech converter using Coqui TTS models (XTTS v2, Bark, VITS, YourTTS) and StyleTTS 2. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio with voice cloning support. No API keys or cloud services required.
C0m3b4ck/DocToSpeech-Chatterboxv7.0updated 15d ago
Fully offline document-to-speech converter using Resemble Chatterbox TTS. 350M parameter model with paralinguistic tags and voice cloning. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio. No API keys or cloud services required.
Finrandojin/alexandria-audiobookv5.0updated 15d ago
A multi-voice AI audiobook generator built on Qwen3-TTS — annotate scripts with an LLM, assign unique voices to each character, per-line style instructions for delivery, clone voices from reference audio, design new voices from text descriptions, train custom voices with LoRA fine-tuning, and export to MP3 or Audacity multi-track projects
🗣️ Generative text-to-speech optimized for dialogue. Natural, expressive English and Chinese speech with fine-grained control over laughter, pauses and prosody, plus reusable speaker embeddings.
⚡️ Efficient 6B parameter image generation model with sub-second inference. Generate high-quality, photorealistic images with only 8 inference steps. Features bilingual text rendering (Chinese & English) and Single-Stream Diffusion Transformer architecture.