A Web UI for easy subtitle using whisper model.
cocktailpeanut
@cocktailpeanutCreations by @cocktailpeanut
157 totalmoondream1 is a tiny (1.6B parameter) vision language model trained by @vikhyatk that performs on par with models twice its size. It is trained on the LLaVa training dataset, and initialized with SigLIP as the vision tower and Phi-1.5 as the text encoder. https://huggingface.co/spaces/vikhyatk/moondream1
Focus on prompting and generating
Next generation face swapper and enhancer
an image to mp4 workflow, created by morpheus
Run Latent Consistency Models on your Mac
Open Vocabulary Image Segmentation using Segment Anything Model and MetaCLIP combo
Contribute to cocktailpeanut/kohya_ss development by creating an account on GitHub.
Next generation face swapper and enhancer
Stable Diffusion web UI
[ICCV 2023] StableVideo: Text-driven Consistency-aware Diffusion Video Editing https://rese1f.github.io/StableVideo/
A webui for different audio related Neural Networks
Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
AUTOMATIC1111/stable-diffusion-webui
Text to audio, open sourced by Meta

