Interpretable Hierarchical Reinforcement Learning for Continuous Control

A Study of NEXUS Across Skills, Language, and Vision

Ivan Smirnov · Maryana Smirnova · Anja Koroveshi · Rosela Berberi

Technical University of Darmstadt

One greedy rollout per variant. The seed is the one with the highest training return for that variant. Hopper flat is the 8× training budget, which is the highest-return flat seed. Videos loop.

Cartpole
Cheetah
Walker
Hopper
Panda
Go1 flat
Go1 rough

Cartpole

Balance an inverted pole on a cart.

Flat · seed 1 · return 999

Neural · seed 1 · return 989

NeSy · seed 0 · return 482

Symbolic · seed 2 · return 315

Cheetah

Planar cheetah runs forward as fast as it can.

Flat · seed 21 · return 902

Neural · seed 0 · return 906

NeSy · seed 0 · return 907

Symbolic · seed 1 · return 875

Walker

Planar biped walks forward at a target speed.

Flat · seed 5 · return 961

Neural · seed 0 · return 966

NeSy · seed 2 · return 942

Symbolic · seed 1 · return 732

Hopper

One-legged hopper hops forward without falling.

Flat · 8× budget · seed 5 · return 648

Neural · seed 10 · return 296

NeSy · seed 13 · return 370

Symbolic · seed 1 · return 208

Panda

Franka Panda reaches, grasps and lifts a cube.

Flat · seed 0 · return 1164

Neural · seed 2 · return 456

NeSy · seed 2 · return 458

Symbolic · seed 1 · return 479

Go1 flat

Unitree Go1 tracks a commanded linear and yaw velocity on flat ground.

Flat · seed 4 · return 11

Neural · seed 9 · return 6

NeSy · seed 7 · return 8

Symbolic · seed 2 · return 9

Go1 rough

Unitree Go1 tracks a commanded velocity over a heightfield.

Flat · seed 13 · return 6

Neural · seed 17 · return 5

NeSy · seed 2 · return 11

RGB observations

The skill actor encodes the 64×64 three-frame grayscale camera and concatenates that embedding with the full state. The critic and the meta-controller still read the state alone. The clips below are that camera. Side-by-side clips put the state-only rollout on the left and the state-plus-camera rollout on the right.

The skill actor concatenates a camera embedding with the full state.

Cartpole, in-loop camera, seed 0. 64×64 grayscale added to the state.

Cheetah, in-loop camera, seed 0. 64×64 grayscale added to the state.

Walker, in-loop camera, seed 0. 64×64 grayscale added to the state.

Walker, in-loop camera, seed 1. 64×64 grayscale added to the state.

Cartpole, seed 0. State only on the left, state plus camera on the right.

Cheetah, seed 0. State only on the left, state plus camera on the right.

Walker, seed 0. State only on the left, state plus camera on the right.

Walker. State only, seed 0, next to state plus camera, seed 1.