学术前沿

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

首次构建统一视觉-语言-动作模型,增强空间细节和动态信息理解。