This is the pre-proceedings for the RLC 2026. You may expect minor changes.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control

By Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, and Clare Lyle

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Deep RL, Visual RL, Neural Network Design, Plasticity, Normalization

Abstract:

Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly. This challenge is pronounced in visual RL, where high-dimensional inputs often obscure learning signals. While prior work in visual RL has focused on algorithmic solutions, such as better dynamics models or exploration strategies, recent advances in state-based RL show that architectural design alone can lead to significant gains in sample efficiency. This raises an important question: Can these architectural principles transfer to visual RL? In response, we introduce V-Simba, a simple yet effective visual RL architecture inspired by the Simba architecture from state-based RL. Built on top of Soft Actor-Critic (SAC) with data augmentation, V-Simba modifies the architecture by adding normalization layers to stabilize training and using pointwise convolutions to reduce computation. Despite its simplicity, V-Simba matches or outperforms the state-of-the-art methods across DMC, Adroit, and Meta-World benchmarks, while being more computationally efficient than DrQ-v2. We make our code publicly available at https://github.com/DAVIAN-Robotics/V-Simba.


Citation Information:

Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, and Clare Lyle. "V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{kim2026vsimba,
    title={V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control},
    author={Donghu Kim and Youngdo Lee and Hojoon Lee and Johan Obando-Ceron and Byungkun Lee and Aaron Courville and Pablo Samuel Castro and Jaegul Choo and Clare Lyle},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}