Global radar

New projects people are discovering or following across Pinokio.
Discovered7d ago
Roop_unleashed4.3.2 - a Hugging Face Space by prince24j
huggingface.co/spaces

Swap faces in videos by providing a source image and a target video. The application will generate a new video with the face from the source image replacing the face in the target video.

Followed7d ago
e2-f5-tts
github.com/EdAlXGoAm

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching https://huggingface.co/spaces/mrfakename/E2-F5-TTS

Followed7d ago
Nemoml
github.com/6Morpheus6

[NVIDIA ONLY] A minimal Gradio interface for Automatic Speech Recognition. Transcribe Audio in Malayalam language.

Followed7d ago
DocToSpeech
github.com/C0m3b4ck

Fully offline document-to-speech converter supporting 9 TTS models across 5 engine groups: Coqui (XTTS v2, Bark, VITS, YourTTS), StyleTTS 2, Resemble Chatterbox, Tortoise TTS, Sesame CSM-1B, and OpenVoice V2. Converts EPUB, PDF, DOCX, HTML, and TXT files to audio with voice cloning support. No API keys or cloud services required.

Followed7d ago
crispz-krea
github.com/mikecastrodemaria

FLUX.1 Krea [dev] txt2img studio (Fooocus-style, 100% local): high-aesthetic text-to-image, ESRGAN+Flux refine upscale, single-file/Civitai Flux models, LoRA, styles, Describe/Improve & Vision Mix (Ollama), Remove BG, Reframe/outpaint, Face Swap. Fork of crispz-studio. https://github.com/mikecastrodemaria/crispz-krea

Followed7d ago
Moondream3 Gradio UI
github.com/PierrunoYT

A web interface for the Moondream3 vision-language model featuring image captioning, visual question answering, object detection, and object pointing.

Followed7d ago
InfiniteYou
github.com/petermg

[NVIDIA ONLY - WINDOWS ONLY] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity [LoRA support fork] https://github.com/petermg/InfiniteYou

Discovered7d ago
Wan-Animate-2
github.com/wan-video

Contribute to Wan-Video/Wan-Animate-2 development by creating an account on GitHub.

Followed7d ago
searxng
github.com/searxng

SearXNG is a free internet metasearch engine which aggregates results from various search services and databases. Users are neither tracked nor profiled.

Followed7d ago
n8n
github.com/SUP3RMASS1VE

Secure Workflow Automation for Technical Teams

Followed7d ago
Ultimate-TTS-Studio-SUP3R-Edition
github.com/SUP3RMASS1VE

Kokoro, KittenTTS, Higgs audio, Chatterbox/Multi, Fish-Speech, F5 & index-tts & indextts2, VoxCPM and VibeVoice in one app

Followed7d ago
stable-diffusion-webui-ux
github.com/Feedjer

Stable Diffusion web UI UX: https://github.com/anapnoe/stable-diffusion-webui-ux

Followed7d ago
Irodori-TTS
github.com/ncore7

Irodori-TTS is a Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control

Followed7d ago
paligemma
github.com/cocktailpeanutlabs

an open vision-language model by Google. PaliGemma is designed as a versatile model for transfer to a wide range of vision-language tasks such as image and short video caption, visual question answering, text reading, object detection and object segmentation https://huggingface.co/spaces/google/paligemma

Followed7d ago
Illusion Diffusion HQ
github.com/6Morpheus6

A simple, high-quality image generation tool to create stunning illusions.

Discovered7d ago
StableVITON
github.com/rlawjdghek

[CVPR2024] StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On

Followed7d ago
Hello World Gradio
github.com/halr9000

Simple hello world app using Gradio and uv sync

Followed7d ago
ScribeTube
github.com/PierrunoYT

Download and transcribe many YouTube videos or whole playlists at once with faster-whisper. Outputs txt, srt, vtt, or json.