Interesting approach to generate a persistent world.
Instead of just generating and displaying everything in realtime, it separates what the world is, from what the world looks like. An LLM tracks objects and state externally, while the video model renders them. New objects discovered in generated frames get added back into that state, so they can be interacted with later and changes can persist even after the camera looks away.
None of the pieces are individually new. But it packages them in a single workflow where persistent interactive worlds are built by giving video models an external world state than by asking one model to handle appearance, memory, and interaction all at once.
The problem is, this is yet another empty github repo:
Code is still under preparation. All code will be available later.
But the idea itself is simple enough that anyone could probably ask an agent to implement this. Bookmarking so I can revisit when they do release code, or if someone takes a stab at implementing this on their own.
Interesting approach to generate a persistent world.
Instead of just generating and displaying everything in realtime, it separates what the world is, from what the world looks like. An LLM tracks objects and state externally, while the video model renders them. New objects discovered in generated frames get added back into that state, so they can be interacted with later and changes can persist even after the camera looks away.
None of the pieces are individually new. But it packages them in a single workflow where persistent interactive worlds are built by giving video models an external world state than by asking one model to handle appearance, memory, and interaction all at once.
The problem is, this is yet another empty github repo:
But the idea itself is simple enough that anyone could probably ask an agent to implement this. Bookmarking so I can revisit when they do release code, or if someone takes a stab at implementing this on their own.