SmoothRL Framework by Astribot Enables Asynchronous Online Reinforcement Learning for Large Models

Deep News
昨天

Astribot's foundational model team has recently unveiled SmoothRL, an online reinforcement learning framework designed to support asynchronous execution, enabling large models to refine their actions during real-time operations. This innovation addresses the critical challenge of determining which actions an online RL system should learn from once asynchronous inference becomes the standard in real-world deployments.

SmoothRL marks the first time such asynchronous online RL has been validated in real-world, high-dynamic throwing tasks. This demonstrates that online learning is not only effective for high-precision corrections but also adaptable to continuous acceleration, precise release, and dynamic operations that cannot afford pauses. Over the past few years, robotics foundation models have primarily focused on expanding capabilities by leveraging more data and stronger models to broaden what robots can do.

However, when robots transition to real-world environments, the persistent issues often aren't a lack of capability but rather minor deviations—being off by a few millimeters or a split second, or failing to maintain stability when object positions shift. SmoothRL targets this next layer: whether a pretrained policy, already equipped with foundational abilities, can continue to self-correct based on actual execution outcomes once deployed in the real world.

SmoothRL points to an increasingly concrete problem: pretraining teaches robots how to perform tasks, but post-training in the real world is what sharpens those skills to be more precise, stable, and reliable. As the model develops comprehensively, Astribot's full-stack system, which integrates AI models, embodied OS, and cable-driven hardware, continues to evolve, driving the application and large-scale deployment of Physical AI. Moving forward, the team plans to explore broader policy updates, wider task distributions, and the integration of asynchronous execution with end-to-end optimization of generative policies.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10