Combining next-token prediction and video diffusion in computer vision and robotics

17.10.2024

Combining next-token prediction and video diffusion in computer vision and robotics

By Robotics News - Robot News, Robotics, Robots, Robotics Sciences in News, robotics, Robotics Classification, robots, robots in business, Robots Podcast Tag news

In the current AI zeitgeist, sequence models have skyrocketed in popularity for their ability to analyze data and predict what to do next. For instance, you've likely used next-token prediction models like ChatGPT, which anticipate each word (token) in a sequence to form answers to users' queries. There are also full-sequence diffusion models like Sora, which convert words into dazzling, realistic visuals by successively "denoising" an entire video sequence.