Phosphene 3.4.0 — two engines: LTX-2.3 and Hailuo H3 (it talks)
Sound on — the pancake sizzle and "Dinner is served" were generated with the pixels, locally. That's the new engine.
Phosphene 3.4.0 adds an engine picker. Two engines, one app, both 100% local on Apple Silicon:
- LTX-2.3 — everything you already use: trained characters, Remix/Ingredients, keyframes, extend. Runs on every supported Mac.
- Hailuo H3 (new, optional) — MiniMax's 33B omni-modal model: video with joint dialogue + sound, in Text and Image modes. Image mode = first-frame conditioning: drop in a still (like your character's face) and H3 animates it — speaking.
H3 tiers, measured on an M4 Max 64 GB (optimized settings out of the box):
- Draft · 3s — 640×384 — ~3 min
- HQ · 3s — 768×448 — ~5 min
- HQ · 5s — 768×448 — ~8 min
- Long · 10s — 768×448 — ~36 min (batch territory)
Requirements & install: H3 is an optional pack — hit Update, then "Install Hailuo H3" in the sidebar (~75 GB; needs 64 GB unified memory — the picker tells you honestly if your Mac can't run it, and LTX is unaffected either way). Weights are under the MiniMax Community License (territory restrictions apply).
Under the hood: H3 runs as an isolated engine — the LTX pipeline is untouched, your gallery/queue/logs work identically across both, and the whole thing was validated end-to-end with real renders (the clip above is the actual validation render, straight from the queue).
Built on the PipeNetwork MLX port of H3 — credit where due. The practical-tier work and the measured configs: https://github.com/mrbizarro/minimax-h3-mlx/tree/codex/h3-engine
v3.4.0 is live on main now. Update and pick your engine. 🎛️
