Supertonic is a lightning-fast, on-device text-to-speech system designed for extreme performance with minimal computational overhead. Powered by ONNX Runtime,
@sup3rmass1ve
Projects by @sup3rmass1ve
50 totalThis project is an enhanced version of the IC-Light repository, designed for advanced image relighting and enhancement using Stable Diffusion and deep learning techniques
Forget everything you thought you knew about AI art generation - RuinedFooocus is here to completely reinvent the game!
SongBloom, a novel framework for full-length song generation
Fast and High-Quality Zero-Shot voice clone Text-to-Speech with Flow Matching
Audio Transcription App with Parakeet-TDT-0.6b-v2
NeuTTS Air is the world’s first super-realistic, on-device, TTS speech language model with instant voice cloning. Built off a 0.5B
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
Dedicated gradio based WebUI for ComfyUI-AdvancedLivePortrait
Higgs Audio Text-to-Speech Playground
Real Time Speech Transcription
Generate realistic and expressive speech with natural language voice design.
BEN2 for background removal
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics,
(WINDOWS)NVIDIA, Hallo2: Long-Duration and High-Resolution Audio-driven Portrait Image Animation
interacting with the Ovis2-8B model. The script allows users to load the model, process image and video inputs, and generate text-based responses using a conversational chatbot.
