Drag in an image, get a description and video prompt from a local Ollama vision model, then render it on your ComfyUI instance.