ChronoSRLTemporal Geometry for Self-Supervised Reinforcement Learning
1Technical University of Darmstadt 2Robotics Institute Germany (RIG), German Research Center for AI (DFKI), hessian.AI
A goal that is close in space can be far away in time. ChronoSRL trains the distances in its critic's embedding space to match the time the agent needs to reach a goal, and predicts from them when the agent arrives and how long it stays. It learns faster and reaches higher performance than CRL, AC-CRL and SRL across network depths, even with much smaller networks.
How ChronoSRL works
ChronoSRL trains the distance between state–action and goal embeddings to match the time the agent takes to reach the goal.
Goals that were never reached are pushed beyond . Two heads on the same embeddings predict when the goal is reached and how long the agent stays there, and the policy maximizes both.
Distance follows time
In Ant Hardest Maze, ChronoSRL's distance follows the steps the agent still needs more closely than the distance in space, and a critic without the geometric losses hardly follows them.
Faster across network depths
The comparison covers seven tasks, four methods and network depths from 1 to 64.
Toward self-supervised robot locomotion
The Unitree Go2 learns velocity tracking, goal-position reaching and box climbing in a sim-to-real locomotion setup.
BibTeX
@article{bohlinger2026chronosrl,
title = {ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning},
author = {Bohlinger, Nico and Peters, Jan},
journal = {arXiv preprint arXiv:2609.36238},
year = {2026}
}