Global radar

New projects people are discovering or following across Pinokio.
Followed23d ago
IllusionDiffusion
github.com/cocktailpeanut

Generate stunning illusion artwork with StableDiffusion (A space by @angrypenguinPNGAP - created with Monster Labs QR ControlNet.

Followed23d ago
Open-Hivemind
github.com/matthewhand

Run the Open-Hivemind multi-agent orchestrator locally with Pinokio.

Discovered23d ago
qwen-tts-studio
github.com/danmoreng

Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.

Discovered23d ago
qwen-voice-clone-webui
github.com/shuichi346

A Gradio WebUI for voice cloning powered by Qwen3-TTS. Provide reference audio via YouTube URL, microphone recording, or file upload — the app transcribes it with Whisper and clones the voice for TTS. Save/load voice profiles for reuse. Optimized for Apple Silicon (MPS).

Discovered24d ago
image-to-3d
github.com/dotneet

Contribute to dotneet/image-to-3d development by creating an account on GitHub.

Followed24d ago
Spark-TTS
github.com/SUP3RMASS1VE
Followed24d ago
Step-Audio-Edit-LOWVRAM
github.com/SUP3RMASS1VE

A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics,

Followed24d ago
Qwen3-TTS-Openai-Fastapi
github.com/dingausmwald

Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.

Discovered24d ago
frame
github.com/aregrid

Frame is an AI-powered, open-source vibe video editor, offering a Professional VIDEO cuting alternative for creators. With Cursor-like interaction, it automates editing, enhances videos, and delivers a seamless vibe video editing experience.

Discovered24d ago
silero-vad
github.com/snakers4

Silero VAD: pre-trained enterprise-grade Voice Activity Detector

Followed24d ago
Kronos
github.com/shiyu-coder

Kronos: A Foundation Model for the Language of Financial Markets

Followed24d ago
Realtime StableDiffusion
github.com/cocktailpeanut

Demo showcasing ~real-time Latent Consistency Model pipeline with Diffusers and a MJPEG stream server (https://github.com/radames/Real-Time-Latent-Consistency-Model)

Followed24d ago
Unlimited-OCR
github.com/baidu

Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

Discovered24d ago
Text To Speech - a Hugging Face Space by walidadebayo
huggingface.co/spaces

Enter or upload plain text or an SRT file, pick a voice and tweak speed, pitch, and volume, then get an MP3 audio file. You can also produce a synchronized subtitle (.srt) file, and the app offers ...