Bipedal robots are underactuated systems in which balance and collision avoidance are tightly coupled, so obstacle-aware walking is substantially harder than for wheeled platforms. We study a biped with only four actuated joints per leg (hip yaw, hip pitch, knee pitch, ankle pitch). The absence of roll joints constrains frontal-plane control and makes collision-free navigation a demanding test case. We present a hierarchical reinforcement learning (HRL) framework in which a high-level (HL) navigation policy, trained with Soft Actor–Critic (SAC), observes the robot pose, raycast proximity measurements, dynamic obstacle states, and a receding-horizon local goal drawn from a global plan, and outputs a body-velocity command (vx, vy, ωyaw). A velocity-conditioned low-level (LL) SAC gait policy executes each command for the next ten control steps through joint-position targets tracked by PD controllers. The temporal abstraction separates navigation from gait generation while keeping both learned and balance-aware. Because the LL gait is task-agnostic, we replace the learned HL policy with classical planners over the identical command interface and obtain three controlled hybrid baselines, SAC+A⋆, SAC+RRT⋆, and SAC+APF. In randomized simulated environments the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, and an ablation study quantifies the contribution of each observation channel and reward component.
Our Hierarchical Reinforcement Learning (HRL) framework splits the control complexity into two distinct, specialized policies connected by a simple command velocity interface:
Trained via Soft Actor-Critic (SAC). It processes the robot pose, raycast lidar proximity measurements, dynamic obstacle states, and a receding-horizon local goal, and outputs target body velocity commands (vx, vy, ωyaw).
A velocity-conditioned gait policy trained with SAC that executes high-level velocity commands over a 10-step horizon. It outputs target joint positions tracked by joint-level PD controllers, stabilizing the underactuated biped's frontal-plane balance.
Training is conducted in physically realistic environments with randomized obstacles. The biped features only four actuated joints per leg, creating a severe control constraint in the frontal plane due to the absence of roll joints.
By separating the navigation task from the gait generator, the LL gait policy learns task-agnostic balance preservation and footprints. It can then be seamlessly paired with different high-level planning policies.
The proposed HRL framework is compared against three hybrid baselines that swap out the learned HL policy with classical planning algorithms (A*, RRT*, and Artificial Potential Fields) while utilizing the same learned LL gait policy.
| Method | Static Trials Success | Dynamic Trials Success |
|---|---|---|
| Proposed HRL (SAC+SAC) | 98.0% | 88.0% |
| Hybrid SAC + A* | 78.0% | 68.0% |
| Hybrid SAC + RRT* | ≤ 78.0% | ≤ 68.0% |
| Hybrid SAC + APF | ≤ 78.0% | ≤ 68.0% |
The proposed end-to-end HRL method significantly outperforms hybrid classical planners by coordinating balance and navigation dynamically.
An extensive ablation study was conducted to evaluate the contribution of individual sensory inputs and reward components:
Observation Channels: Removing raycast proximity sensors or dynamic obstacle state feeds drastically decreases the success rates in cluttered dynamic environments.
Reward Shaping: The action-smoothness penalty and target tracking reward terms prevent joint jerks and stabilize the underactuated biped's frontal-plane swing.