Prompt Orchestrator that turns module-based game design (genre, mechanics, visuals, menus, audio) into a complete, playable HTML5 game generated by your chosen AI provider. Supports OpenAI, Gemini, Claude, Ollama, and LM Studio. Every game ships as a single self-contained HTML file.
Global radar
seed-vc voice conversion adapted for Maestro (GPL-3.0). Cloned automatically at install time by github.com/Blizaine/Maestro - not for standalone use.
Contribute to Hakaze/wan2gp-flashvsr development by creating an account on GitHub.
Contribute to GeekatplayStudio/Image-Express development by creating an account on GitHub.
[NVIDIA GPU REQUIRED] Realtime world generator by Overworld Waypoint world model
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
Text-to-audio with SA3 Medium / Small Music / Small SFX.
Video to Openpose & DWPose (All OS supported) https://github.com/sdbds/vid2pose
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
AI outpainting and 2D game-art studio powered by Gemini image models through OpenRouter.
Z-Image txt2img + upscaler/detailer studio (Fooocus-style, 100% local): txt2img, ESRGAN+Z-Image refine upscale, single-file/Civitai models, LoRA, styles, Describe/Improve & Vision Mix (Ollama), Remove BG, Reframe/outpaint, Face Swap. https://github.com/mikecastrodemaria/crispz-studio
Branded text-to-video studio for long-form documentary YouTube videos. Local LLM scripts via LM Studio, web image sourcing with lightweight AI recreation (uniform 1080p, watermark/text removal), local Kokoro TTS narration, cinematic Ken Burns rendering, automatic GPU load/unload, and a modern web UI to manage everything.
Contribute to MiniMax-AI/MiniMax-Music3 development by creating an account on GitHub.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Uncensored Image editing tool Edit image with a text prompt
Unified audio production dashboard for Volt Records - integrates StableDAW, Stable Audio 3, TASCAR, and catalog intelligence into a single command center.
Customizing Realistic Human Photos via Stacked ID Embedding https://huggingface.co/spaces/TencentARC/PhotoMaker-V2
