跳到主要导航 跳到搜索 跳到主要内容

A Deep Reinforcement Learning Model for Humanoid Robot Navigation and Task Scheduling in Automated Supply Chains

  • Jiadong Zhang*
  • , Wei Wang
  • *此作品的通讯作者
  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Humanoid robots must navigate, decide, and schedule efficiently to boost automated supply chain efficiency. Traditional rule-based techniques fail in dynamic situations, especially with task dependencies. A unique Puma Optimizer-mutated Twin-Stage Adaptive Twin-Delayed Deep Deterministic Policy Gradient (PO-TSATD3) method is used in this deep reinforcement learning (DRL) system. Training datasets imitate real-world logistics situations with dynamic impediments, many robots, and varying workloads. Data preparation cleans and normalizes for quality. The Puma Optimizer optimizes convergence and operating efficiency, while the PO-TSATD3 framework improves navigation and scheduling adaptive learning. Python simulations show considerable gains in navigation accuracy, collision reduction, and schedule optimization over conventional methods. The model’s outstanding performance metrics proved its scalability and durability in complicated situations. This research validates the application of DRL, augmented by PO-TSATD3, as a powerful solution for intelligent humanoid robot operations in future supply chain systems.

源语言英语
文章编号2640014
期刊International Journal of Humanoid Robotics
DOI
出版状态已接受/待刊 - 2026

指纹

探究 'A Deep Reinforcement Learning Model for Humanoid Robot Navigation and Task Scheduling in Automated Supply Chains' 的科研主题。它们共同构成独一无二的指纹。

引用此