Sign in

marcelvangerven.bsky.social

@marcelvangerven.bsky.social
17 followers 29 following 3 posts
PostsRepliesMedia
marcelvangerven.bsky.social @marcelvangerven.bsky.social · 08/05/2026
Proud to share a new pre-print by @davidleeftink.bsky.social and with @maxhinne.bsky.social. We make a formal connection between the Pontryagin minimum principle and deep recurrent reinforcement learning, improving performance on challenging control tasks! arxiv.org/pdf/2605.05373
Recurrent Reinforcement Learning via Neural co-state policies. NCP s are regularized during training to mirror the optimality conditions implied by the minimum principle. By structuring the hidden states as representations of the underlying optimal control co-states, the read-out layer acts as a control-Hamiltonian minimizer.
041
Reposted by @marcelvangerven.bsky.social
Max Hinne @maxhinne.bsky.social · 01/12/2025
Proud to share a new pre-print by @davidleeftink.bsky.social and with @marcelvangerven.bsky.social on Bayesian optimization for automatic laser dicing in semiconductor manufacturing! arxiv.org/abs/2511.23141
012
marcelvangerven.bsky.social @marcelvangerven.bsky.social · 19/09/2025
New preprint out on neuromorphic intelligence: arxiv.org/pdf/2509.11940
arxiv.org
000
Reposted by @marcelvangerven.bsky.social
David Leeftink @davidleeftink.bsky.social · 03/09/2025
Our recent paper “Optimal Control of Probabilistic Dynamics Models via Mean Hamiltonian Minimization,” has been accepted to CDC 2025! We demonstrate how applying optimal control principles can significantly improve planning in deep model-based reinforcement learning with epistemic uncertainty.
181
marcelvangerven.bsky.social @marcelvangerven.bsky.social · 05/04/2025
New preprint out! arxiv.org/abs/2504.02543
arxiv.org
Probabilistic Pontryagin's Maximum Principle for Continuous-Time Model-Based Reinforcement Learning
Without exact knowledge of the true system dynamics, optimal control of non-linear continuous-time systems requires careful treatment of epistemic uncertainty. In this work, we propose a probabilistic...
030