Table of Contents
- BONSYSTEMS · CYCLOIDAL TECHNOLOGY
- Frequently Asked Questions (FAQ)
- References
- Bonsystems
- 🔗 Get the technology trend materials → www.bonsystems.com

Reinforcement learning is an AI technique that enables humanoid robots to learn movements in a variety of environments. Unlike rule-based approaches in which developers must explicitly program each movement, reinforcement learning allows a robot to explore different actions and learn from their outcomes.
This characteristic is particularly relevant to humanoid robots, which must maintain balance while walking and adjust their posture in response to changes in their surroundings. In recent years, the use of AI-based learning has expanded beyond walking to object manipulation tasks such as grasping and moving objects, as well as more complex task execution.
This article examines the basic concept of reinforcement learning, how it is used in humanoid walking and motion control, the role of simulation-based learning, and the importance of drive units in executing learned movements on physical robots.
The Basic Concept of Reinforcement Learning
Key Summary Reinforcement learning is a method in which an AI agent interacts with an environment, takes actions, receives feedback, and learns which actions are most effective for achieving a defined objective.
In reinforcement learning, an AI agent interacts with a given environment and learns through the outcomes of its actions. Rather than following a fixed sequence of predefined instructions, the agent repeatedly takes actions and receives feedback that helps it determine which behaviors are more effective for achieving a specific goal.
During the learning process, the available actions and learning objective are defined. The system that learns and takes actions is called the agent, while the space in which the agent operates is called the environment.
One of the key characteristics of this learning approach is that developers do not need to explicitly program every action in advance. Instead, the agent can learn appropriate behaviors by repeatedly interacting with its environment. This makes reinforcement learning useful for problems that involve complex or changing conditions.
However, learning outcomes can vary significantly depending on how the objective, environment, and reward conditions are defined. When this method is applied to robots, the task, operating environment, and required range of motion must therefore be clearly defined in advance.
Why Reinforcement Learning Is Used in Humanoid Robots
Key Summary A humanoid robot has a human-like structure with multiple joints throughout its legs, torso, arms, and hands. Walking and posture control require coordinated movement across many joints rather than the movement of a single joint.
A humanoid robot has a human-like structure with multiple joints throughout its legs, torso, arms, and hands. To walk, maintain balance, or change posture, multiple joints must move in a coordinated manner.
For example, when the robot moves one leg forward, the torso and opposite leg must adjust simultaneously to maintain balance. When reaching for an object, the shoulder, elbow, and wrist must also move in coordination.
Humanoid motion is therefore created through coordinated movement across multiple joints. Robot learning can help the system develop coordinated joint behaviors these coordinated joint behaviors toward a defined objective, such as walking forward, recovering balance, or reaching for an object.
However, robot movement is not determined by reinforcement learning alone. Joint mechanics, range of motion, sensor configuration, actuator performance, and control systems all work together. Reinforcement learning should therefore be understood as one part of a broader robotic system rather than as a standalone solution.
Learning Using a Simulation Environment
Key Summary Humanoid motion learning requires joints to move repeatedly while the results of those movements are evaluated. In addition to physical robots, simulated robots can be used to perform this learning process in a virtual environment.
During humanoid motion learning, joints must move repeatedly while the resulting behavior is evaluated. Instead of performing every learning iteration on a physical robot, developers can use simulated robots in virtual environments.
In simulation, the robot’s joint structure, range of motion, body geometry, and environmental conditions can be modeled digitally. Walking, arm movements, posture changes, and other behaviors can then be tested and adjusted without repeatedly operating the physical robot.
Reinforcement learning is widely used in these simulation environments. A simulated robot can repeatedly explore different actions under defined conditions and learn behaviors that help it achieve a specific objective. The learned behavior can then be evaluated for transfer to a physical robot.
However, differences inevitably exist between simulation and real-world robot behavior. Factors such as mechanical tolerances, actuator characteristics, sensor noise, friction, and external disturbances may not be represented perfectly in a virtual environment.
For this reason, behaviors learned in simulation must be validated and adjusted according to the mechanical structure, components, sensors, and operating conditions of the physical robot. This process is commonly associated with sim-to-real transfer, in which behaviors learned in simulation are adapted for real-world operation.
The Relationship Between Learning-Based Control and Robot Hardware
Key Summary Even when a robot learns a motion strategy through reinforcement learning, robot hardware is ultimately responsible for executing the learned motion. Actuators convert control commands into physical joint movement.
Even when a robot learns a motion strategy through reinforcement learning, robot hardware is ultimately responsible for executing that motion. Actuators convert commands from the control system into physical joint movement.
A humanoid robot uses many joints simultaneously while walking, maintaining balance, or handling objects. Each joint may require different levels of torque and different ranges of motion depending on its location and function.
For this reason, realizing humanoid movement requires more than AI-based learning. Joint mechanics, range of motion, sensor configuration, actuator performance, drive-unit size, and component layout must all be considered as part of an integrated system.
The learning system determines and improves motion strategies, while hardware physically executes those strategies. The two perform different functions but must work together to produce stable and practical robot movement.
The Role of Actuators in Humanoid Motion
Key Summary Learning-based control can help a humanoid robot develop motion strategies for walking, posture control, and object manipulation. Actuators are required to execute these learned motions by converting control commands into physical joint motion and torque.
Reinforcement learning can be used to learn motion strategies for humanoid walking, posture control, and object manipulation. To execute these learned motions on a physical robot, actuators must convert control commands into physical joint motion and torque.
An actuator converts motor output into the torque and motion required at each joint, allowing the robot to move its arms, legs, and torso. When a humanoid supports its body on one leg, changes posture, or lifts an arm to manipulate an object, each joint must deliver the torque required for its specific position and function.
The Bonsystems BCSA is an actuator that combines a frameless motor and cycloidal reduction technology in a thin, compact form. In robots that integrate many joints within a compact body, such as humanoids, actuator size and thickness can directly affect joint design and the robot’s overall profile.
The BCSA is designed to deliver the torque required at each joint within a slim structure. A compact actuator can provide additional space for sensors, wiring, structural components, and other elements around joints such as the shoulder, elbow, hip, and knee.
A compact drive unit not only reduces installation space but can also simplify the layout of components inside the joint and support a more efficient overall robot design.
The BCSA uses Bonsystems’ Pinless cycloidal structure. This design reduces certain coupling elements used in conventional pin-based configurations and simplifies the reducer’s internal structure. A simplified component structure can help streamline assembly and maintenance while also making it easier to integrate the drive unit within a confined space.
The Bonsystems BCSA combines a slim form factor, cycloidal reduction technology, and a Pinless internal design. It is designed to deliver the required joint torque while minimizing actuator thickness and making efficient use of the internal space available in humanoid robots.
As humanoid motions become more diverse, joint drive units must deliver stable torque and movement across a wider range of operating conditions. The BCSA is Bonsystems’ integrated drive solution for development environments where both joint-space efficiency and drive performance are important.
Frequently Asked Questions (FAQ)
Q1. What role does the actuator play in a reinforcement-learning-based humanoid?
A. Reinforcement learning helps the robot learn a motion strategy, while the actuator converts control commands into physical joint movement. It delivers the torque and motion required for actions such as walking, posture adjustment, and arm movement.
Q2. Why is a slim actuator needed in humanoid joints?
A. A humanoid robot integrates many joints within a compact body, while sensors, wiring, structural components, and other devices must also be arranged around those joints. A slim actuator makes it easier to integrate the drive unit and surrounding components within a confined space.
Q3. What structure of actuator is the BCSA?
A. The BCSA is an actuator that combines a frameless motor and cycloidal reduction technology in a thin, compact structure. By integrating the motor and reduction mechanism into a single drive unit, it can help make efficient use of the limited space inside a robot joint.
Q4. What is the BCSA’s Pinless cycloidal structure?
A. The Pinless structure reduces certain coupling elements used in conventional pin-based configurations to simplify the internal reduction mechanism. This can help reduce component complexity, streamline assembly and maintenance, and make efficient use of limited internal space.
Q5. Which joints of a humanoid can the BCSA be applied to?
A. The BCSA lineup includes the RO series, a high-torque slim actuator intended for lower-body joints, and the RI series, a compact actuator for upper-body joints. The RO series is designed for lower-body joints such as the hip, knee, and ankle, while the RI series can be applied to upper-body joints such as the shoulder, elbow, and wrist, where the drive unit must fit within a relatively confined space. The appropriate product should be selected based on the torque, size, and range-of-motion requirements of each joint.
References
• [1] Bonsystems Official Website — https://www.bonsystems.com
• [2] International Federation of Robotics (IFR) — https://ifr.org
• [3] Korea Institute for Robot Industry Advancement (KIRIA) — https://www.kiria.org
• [4] Korea Association of Robot Industry (KAR) — https://www.irobotics.or.kr
• [5] Korea Institute of Robotics & Technology Convergence (KIRO) — https://www.kiro.re.kr
