Siddhant Haldar @haldarsiddhant.bsky.social · 28/02/2025The most frustrating part of imitation learning is collecting huge amounts of teleop data. But why teleop robots when robots can learn by watching us? Introducing Point Policy, a novel framework that enables robots to learn from human videos without any teleop, sim2real, or RL. 110
Reposted by Siddhant HaldarLerrel Pinto @lerrelpinto.com · 26/02/2025We just released AnySense, an iPhone app for effortless data acquisition and streaming for robotics. We leverage Apple’s development frameworks to record and stream: 1. RGBD + Pose data 2. Audio from the mic or custom contact microphones 3. Seamless Bluetooth integration for external sensors 23510
Reposted by Siddhant Haldargaoyuezhou.bsky.social @gaoyuezhou.bsky.social · 31/01/2025Can we extend the power of world models beyond just online model-based learning? Absolutely! We believe the true potential of world models lies in enabling agents to reason at test time. Introducing DINO-WM: World Models on Pre-trained Visual Features for Zero-shot Planning. 1208
Reposted by Siddhant HaldarRaunaq Bhirangi @raunaqb.bsky.social · 10/12/2024BAKU is fully open source and surprisingly effective. We found it easily adaptable for a host of visuotactile tasks in visuoskin.github.iovisuoskin.github.ioLearning Precise, Contact-Rich Manipulation through Uncalibrated Tactile SkinsLearning Precise, Contact-Rich Manipulation through Uncalibrated Tactile Skins 082
Siddhant Haldar @haldarsiddhant.bsky.social · 11/12/2024I will be presenting BAKU at the #NeurIPS2024 poster session on Thursday, December 12, from 11 a.m. to 2 p.m. PST at East Exhibit Hall A-C #4206! Do drop in to chat about efficient robot policy architectures as well as some of the more recent work using BAKU. 030
Reposted by Siddhant HaldarRaunaq Bhirangi @raunaqb.bsky.social · 10/12/2024P3-PO is a great example of how simple human priors can facilitate significantly better generalizability for robot policies. 032
Siddhant Haldar @haldarsiddhant.bsky.social · 11/12/2024Turns out that replacing images with keypoint-based representations can enable enhanced generalization across spatial positions and orientations and novel object instances! We just released P3-PO, a method for learning generalizable policies with minimal data. 🚀 110
Reposted by Siddhant HaldarLerrel Pinto @lerrelpinto.com · 09/12/2024Modern policy architectures are unnecessarily complex. In our #NeurIPS2024 project called BAKU, we focus on what really matters for good policy learning. BAKU is modular, language-conditioned, compatible with multiple sensor streams & action multi-modality, and importantly fully open-source! 1309