Learning to Retrieve:
Internalizing Memory Retrieval for Video World Models

JiaKui Hu1,2*, Tailai Chen3*, Yuqi Pan3, Xuerui Qiu3
Jialun Liu4, Xiao Cao5, Zhenxin Zhu2, Guang Chen2, Hangjun Ye2, Bing Wang2#, Yanye Lu1†
1PKU,  2Xiaomi,  3CASIA,  4UQ,  5NUS
#Project leader,  †Corresponding author

We introduce Learning-to-Retrieve (L2R), which can generate consistent videos without external memory modules or 3D reconstruction pipelines. Red boxes mark content observed from matched camera poses separated by a long temporal interval, illustrating that L2R preserves scene identity when the camera revisits a previously observed region.

L2R

Method


Demos

Demo #1: Turn 180 degrees to the left and then turn back.

Demo #2: Turn 180 degrees to the left and then turn back.

Demo #3: (After a teddy bear is inserted into the input image) Turn 90 degrees to the left, then turn back.

Demo #4: (After a teddy bear is inserted into the input image) Turn 45 degrees to the right, then turn back.

Results on Memory

memory

Results on Visual Quality

visual

BibTeX

@article{hu2026l2r,
      title={Learning to Retrieve: Internalizing Memory Retrieval for Video World Models},
      author={Hu, JiaKui and Chen, Tailai and Pan, Yuqi and Qiu, Xuerui and Liu, Jialun and Cao, Xiao and Zhu, Zhenxin and Chen, Guang and Ye, Hangjun and Wang, Bing and Lu, Yanye},
      journal={arXiv preprint arXiv:2610.11444},
      year={2026}
}