Followed3d ago
HiDream O1 Image FP8
github.com/cocktailpeanut

One-click launcher for the original HiDream-O1-Image web UI using lazy-downloaded drbaph Dev or Full FP8 checkpoints through a root FP8 runner. Requires an NVIDIA CUDA GPU.

Followed3d ago
ace-step-ui
github.com/fspecii

🎵 The Ultimate Open Source Suno Alternative - Professional UI for ACE-Step 1.5 AI Music Generation. Free, local, unlimited. Stop paying for Suno!

Followed3d ago
Invoke
github.com/pinokiofactory

The Gen AI Platform for Pro Studios https://github.com/invoke-ai/InvokeAI

Followed3d ago
IndicTrans2
github.com/ai4bharat

Translation models for 22 scheduled languages of India

Followed3d ago
Hy-MT2
github.com/PierrunoYT

Hy-MT2 multilingual translation — Gradio UI with 38 language and variant choices for Hy-MT2-1.8B, Hy-MT2-7B, and Hy-MT2-30B-A3B.

Discovered3d ago
real-avatar-local-open-source
github.com/efwoods

This is a local version of real-avatar that enables the creation of avatar's locally on my personal computer

Followed3d ago
Wan2GP
github.com/cmfreak3-jpg

Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/deepbeepmeep/Wan2GP

Followed3d ago
DocToSpeech
github.com/C0m3b4ck

Fully offline document-to-speech converter supporting 9 TTS models across 5 engine groups: Coqui (XTTS v2, Bark, VITS, YourTTS), StyleTTS 2, Resemble Chatterbox, Tortoise TTS, Sesame CSM-1B, and OpenVoice V2. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio with voice cloning support. No API keys or cloud services required.

Followed3d ago
dengjen-tts
github.com/zirekhq

A cross-platform inference engine for neural TTS models.

Followed3d ago
PhotoMaker2
github.com/pinokiofactory

Customizing Realistic Human Photos via Stacked ID Embedding https://huggingface.co/spaces/TencentARC/PhotoMaker-V2

Followed3d ago
Wan2GP
github.com/jillesmc

Super Optimized Gradio UI for AI video creation for GPU poor machines (6GB+ VRAM). Supports Wan 2.1/2.2, Qwen, Hunyuan Video, LTX Video and Flux. https://github.com/deepbeepmeep/Wan2GP

Followed3d ago
Parakeet-TDT
github.com/SUP3RMASS1VE

Audio Transcription App with Parakeet-TDT-0.6b-v2

Followed3d ago
Transcribr
github.com/PierrunoYT

Bulk transcribe many YouTube videos, whole playlists, or your own uploaded audio/video files at once with faster-whisper. Outputs txt, srt, vtt, or json.

Followed3d ago
DiffRhythm_2
github.com/quymao

Diffusion-based rhythm generation with DiffRhythm2 model

Followed3d ago
Audio Splitter & Transcriber
github.com/Minhaemun

Upload MP3 files, split into chunks, and transcribe to text using Whisper.

Followed3d ago
fluxgym
github.com/cocktailpeanut

[NVIDIA Only] Dead simple web UI for training FLUX LoRA with LOW VRAM support (From 12GB)

Followed3d ago
candy-machine
github.com/Feedjer

Image Dataset Tagger for Stable Diffusion / Lora / DreamBooth Training: https://github.com/mikeknapp/candy-machine

Followed3d ago
muscriptor
github.com/muscriptor

MuScriptor is a multi-instrument music transcription model developed by Kyutai and Mirelo.