Followed3h ago
Hermes Agent
github.com/cocktailpeanut

Self-improving CLI agent by Nous Research with local memory, skills, and messaging workflows.

Followed3h ago
StableDAW
github.com/cocktailpeanut

Browser-based AI audio DAW for Stable Audio 3 with text-to-audio, inpainting, LoRA training, FFmpeg effects, waveform editing, sequencer, piano roll, and persistent library. https://github.com/gantasmo/stabledaw

Followed3h ago
ChatterBox
github.com/PierrunoYT

AI-Powered Text-to-Speech with Voice Cloning using Chatterbox TTS and a Gradio interface. Includes Turbo, Multilingual (23+ languages), and Original models. Runs locally; CUDA GPU recommended, CPU supported. Windows, Mac, and Linux.

Followed4h ago
InstantIR
github.com/pinokiofactory

restore low-res images, restore broken images, recreate a new version of the image with a prompt https://huggingface.co/spaces/fffiloni/InstantIR

Followed4h ago
MuScriptor
github.com/cocktailpeanut

Local multi-instrument audio-to-MIDI transcription from Kyutai and Mirelo.

Followed4h ago
God's Eye View
github.com/theandychang

A live 3D intelligence console for planet Earth.

Followed4h ago
ai-toolkit
github.com/pinokiofactory

AI Toolkit by Ostris

Followed4h ago
HunyuanVideo
github.com/pinokiofactory

[NVIDIA ONLY] Super Optimized Gradio UI for Hunyuan Video Generator that works on GPU poor machines. Generate up to 10~14 sec videos https://github.com/deepbeepmeep/HunyuanVideoGP

Followed4h ago
SwarmUI
github.com/SUP3RMASS1VE

A Modular AI Image Generation Web-User-Interface, with an emphasis on making powertools easily accessible, high performance, and extensibility. Supports Stable Diffusion, Flux, etc. AI image models, with plans to support AI video, audio, and more in the future.

Followed4h ago
pyramidflow
github.com/pinokiofactory

Pyramd Flow Video Generation AI (text-to-video & image-to-video) https://github.com/jy0205/Pyramid-Flow

Followed4h ago
Qwen3-TTS MLX WebUI Enhanced
github.com/Blizaine

High-quality text-to-speech with Beautiful Web UI & API, optimized for Apple Silicon using MLX. Features include Custom Voice (preset speakers), Voice Design (natural language), and Voice Cloning. With enhanced features for saving custom voices and long-form / endless TTS streaming.

Followed5h ago
OpenAudio
github.com/pinokiofactory

Multilingual Text-to-Speech with Voice Cloning (Supports: English, Japanese, Korean, Chinese, French, German, Arabic, and Spanish) https://github.com/fishaudio/fish-speech

Followed5h ago
CogStudio
github.com/pinokiofactory

[NVIDIA ONLY] Advanced Web UI for CogVideo (text to video, image to video, video to video, extend video, etc) -- Generate videos with less than 10GB VRAM

Followed5h ago
Forge
github.com/pinokiofactory

[NVIDIA ONLY] The most efficient way to run FLUX (Optimized to run even on low memory machines, as low as 3GB VRAM with 512x512 resolution) https://github.com/lllyasviel/stable-diffusion-webui-forge

Followed5h ago
ReClip
github.com/krynsky

Self-hosted, open-source video and audio downloader with a clean web UI. Supports YouTube, TikTok, Instagram, X, and 1000+ other sites via yt-dlp.

Followed5h ago
Alexandria
github.com/DokeDev

A multi-voice AI audiobook generator built on Qwen3-TTS — annotate scripts with an LLM, assign unique voices to each character, per-line style instructions for delivery, clone voices from reference audio, design new voices from text descriptions, train custom voices with LoRA fine-tuning, and export to MP3 or Audacity multi-track projects

Followed5h ago
LiquidAI-LFM2.5 Playground
github.com/TheAwaken1

Local multimodal app powered by Liquid AI LFM2.5-Audio-1.5B and LFM2.5-VL-1.6B models, delivering real-time voice chat, text-to-speech synthesis, long-form audio transcription, and multi-image vision reasoning.

Followed5h ago
Comfyui
github.com/pinokiofactory

The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface. https://github.com/comfyanonymous/ComfyUI

Followed6h ago
Roop Ultimate
github.com/rishabh4496

Face swapping for images and video, with a React UI. Independent project; AGPL-3.0. Private — access is by invitation.

Followed6h ago
Jevthoven
github.com/cocktailpeanut

Prompt-first symbolic-music studio powered by live TypeSafe Jev decisions. Describe music in words, listen to an editable composition, keep shaping the same piece. Requires a TYPESAFE_API_KEY for live mode; a no-credit fixture mode is built in.