星尘发布在线强化学习框架SmoothRL,实现与大模型的异步推理

新浪科技
Sep 04

  新浪科技讯 9月4日上午消息,近日,星尘智能(Astribot) 基座模型团队发布能异步执行的在线强化学习框架 SmoothRL(Online Reinforcement Learning During Asynchronous Execution)。解决的是大模型异步推理成为真实部署常态后,与之适配的online RL到底该从哪些动作里学。

  SmoothRL 首次将这类异步online RL放到真实高动态投掷任务中验证,进一步证明在线学习不仅能用于高精度修正,也能适配连续加速、精确释放、不能停顿的动态操作。

  过去几年,机器人基础模型主要解决的是“会不会”:用更多数据、更强模型,把能力做宽。但机器人真正进入真实环境以后,新的问题往往不是完全不会,而是差几毫米、差半拍,或者换一个物体位置就不够稳定。

  SmoothRL 关注的是这之后的一层:一个已经具备基础能力的 pretrained policy,进入真实世界以后,能不能继续根据真实执行结果修正自己。SmoothRL 指向的是一个正在变得越来越具体的问题:预训练让机器人学会怎么做,真实世界里的后训练,再把这些能力练得更准、更稳、更可靠。

  在模型全面发展时,星尘“AI模型-具身OS-绳驱本体”的全栈系统也在持续迭代,推动Physical AI的应用与规模化部署。

  下一步,团队计划进一步探索更大范围的策略更新、更广泛的任务分布,以及异步执行与生成式策略端到端优化之间的结合。

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10