PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation Paper • 2609.34759 • Published 3 days ago • 107
PAWBench: How Far Are We from Probabilistically Aligned World Modeling? Paper • 2608.27345 • Published Aug 27 • 77
PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation Paper • 2609.34759 • Published 3 days ago • 107
On-Policy Self-Distillation for Multi-Turn Image Editing Paper • 2609.35611 • Published 3 days ago • 3
JEV-Star: Fast, Low-Cost StarCraft II Control with Language-Model Planning Paper • 2609.27331 • Published 8 days ago • 5
On-Policy Self-Distillation for Multi-Turn Image Editing Paper • 2609.35611 • Published 3 days ago • 3
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation Paper • 2609.30221 • Published 7 days ago • 46
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence Paper • 2604.19858 • Published Apr 23
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation Paper • 2605.26525 • Published May 26
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions Paper • 2506.23361 • Published Jun 29, 2025 • 1
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published Aug 13 • 47
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World Paper • 2605.13169 • Published May 13 • 20
GDRO: Group-level Reward Post-training Suitable for Diffusion Models Paper • 2601.02036 • Published Jan 5
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions Paper • 2506.23361 • Published Jun 29, 2025 • 1
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives Paper • 2512.14699 • Published Dec 16, 2025 • 29
DiffDoctor: Diagnosing Image Diffusion Models Before Treating Paper • 2501.12382 • Published Jan 21, 2025