500,000 Hours of First-Person Video: How to Add 'Touch' to Embodied World Models
Zhiwu Wujie releases Being-H0.8, a latent tactile world-action model trained on over 500,000 hours of first-person video. By generating pseudo-tactile labels from visual data and unifying robot sensor inputs, it aims to bridge the gap between visual observation and physical contact for embodied AI.