TY - JOUR
T1 - A Deep Reinforcement Learning Model for Humanoid Robot Navigation and Task Scheduling in Automated Supply Chains
AU - Zhang, Jiadong
AU - Wang, Wei
N1 - Publisher Copyright:
© 2026 World Scientific Publishing Company.
PY - 2026
Y1 - 2026
N2 - Humanoid robots must navigate, decide, and schedule efficiently to boost automated supply chain efficiency. Traditional rule-based techniques fail in dynamic situations, especially with task dependencies. A unique Puma Optimizer-mutated Twin-Stage Adaptive Twin-Delayed Deep Deterministic Policy Gradient (PO-TSATD3) method is used in this deep reinforcement learning (DRL) system. Training datasets imitate real-world logistics situations with dynamic impediments, many robots, and varying workloads. Data preparation cleans and normalizes for quality. The Puma Optimizer optimizes convergence and operating efficiency, while the PO-TSATD3 framework improves navigation and scheduling adaptive learning. Python simulations show considerable gains in navigation accuracy, collision reduction, and schedule optimization over conventional methods. The model’s outstanding performance metrics proved its scalability and durability in complicated situations. This research validates the application of DRL, augmented by PO-TSATD3, as a powerful solution for intelligent humanoid robot operations in future supply chain systems.
AB - Humanoid robots must navigate, decide, and schedule efficiently to boost automated supply chain efficiency. Traditional rule-based techniques fail in dynamic situations, especially with task dependencies. A unique Puma Optimizer-mutated Twin-Stage Adaptive Twin-Delayed Deep Deterministic Policy Gradient (PO-TSATD3) method is used in this deep reinforcement learning (DRL) system. Training datasets imitate real-world logistics situations with dynamic impediments, many robots, and varying workloads. Data preparation cleans and normalizes for quality. The Puma Optimizer optimizes convergence and operating efficiency, while the PO-TSATD3 framework improves navigation and scheduling adaptive learning. Python simulations show considerable gains in navigation accuracy, collision reduction, and schedule optimization over conventional methods. The model’s outstanding performance metrics proved its scalability and durability in complicated situations. This research validates the application of DRL, augmented by PO-TSATD3, as a powerful solution for intelligent humanoid robot operations in future supply chain systems.
KW - Humanoid robots
KW - Puma optimizer-mutated twin-stage adaptive twin-delayed deep deterministic policy gradient algorithm
KW - navigation efficiency
KW - supply chain automation
UR - https://www.scopus.com/pages/publications/105039270974
U2 - 10.1142/S0219843626400141
DO - 10.1142/S0219843626400141
M3 - 文章
AN - SCOPUS:105039270974
SN - 0219-8436
JO - International Journal of Humanoid Robotics
JF - International Journal of Humanoid Robotics
M1 - 2640014
ER -