Combining next-token prediction and video diffusion in computer vision and robotics Coworky, Fadi Souilem 23 ott 2024 0 332