Sign in

Sai Kumar Dwivedi

@saidwivedi.in
994 followers 646 following 24 posts

PhD Candidate at @MPI-IS || 3D Vision & Digital Avatars || Ex: @Meta, @Daimler, @Intel Webpage: saidwivedi.in

PostsRepliesMedia
Sai Kumar Dwivedi @saidwivedi.in · 24/09/2026
There’s: • no relay server
 • no account
 • no app to install on the device Free and open source. pockettui.com
 github.com/saidwivedi/p... [7/7]
github.com
GitHub - saidwivedi/PocketTUI: Your terminal on any browser and any device.
Your terminal on any browser and any device. Contribute to saidwivedi/PocketTUI development by creating an account on GitHub.
010
Sai Kumar Dwivedi @saidwivedi.in · 24/09/2026
Setup is one command. Open pockettui.com/app from any browser, enter the machine address and pairing code, and you’re in. It remembers the machine after that, so getting back to a session is one tap. [6/7]
100
Sai Kumar Dwivedi @saidwivedi.in · 24/09/2026
Everything keeps running in tmux on my machine. I can close the tab, lose the connection, or switch devices. When I come back, the same sessions are exactly where I left them. [5/7]
100
Sai Kumar Dwivedi @saidwivedi.in · 24/09/2026
I wanted this to be comfortable to use from my phone. The terminal keys I need are built in: Esc, Ctrl, Tab, arrows. For longer instructions, I can just speak. Transcription runs on my machine and uses terminal context, so it gets filenames, flags and commands right. [4/7]
100
Sai Kumar Dwivedi @saidwivedi.in · 24/09/2026
But the terminal alone doesn’t cover everything I need. Sometimes I want to browse files, check what an agent changed, read Markdown, or open an image or video a job just produced. So I built the UI I needed around the terminal, in the same browser tab. [3/7]
210
Sai Kumar Dwivedi @saidwivedi.in · 24/09/2026
PocketTUI connects to the terminal itself, so there’s nothing special to integrate. Claude Code, Codex, scripts, servers, builds, training jobs. They all stay in their real tmux sessions on your machine. If it runs in your terminal, it works with PocketTUI. [2/7]
100
Sai Kumar Dwivedi @saidwivedi.in · 24/09/2026
I use Claude Code in one terminal, Codex in another, and a training job beside them. Next month, I’ll use something else. I realized I don’t want remote access to an agent. I want remote access to where they all already live: my terminal. So I built PocketTUI.com [1/7]
120
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
InteractVLM (#CVPR2025) is a great collaboration MPI-IS, UvA and Inria. Authors: @saidwivedi.in, @anticdimi.bsky.social, S. Tripathi, O. Taheri, C. Schmid, @michael-j-black.bsky.social and @dimtzionas.bsky.social. Code & models available at: interactvlm.is.tue.mpg.de (10/10)
030
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
InteractVLM is the first method that infers 3D contacts on both humans and objects from in-the-wild images, and exploits these for 3D reconstruction via an optimization pipeline. In contrast, existing methods like PHOSA rely on handcrafted or heuristic-based contacts. (9/10)
120
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
With just 5% of DAMON’s 3D body contact annotations, InteractVLM surpasses the fully-supervised DECO baseline trained on 100% of 3D annotations. This is promising for minimizing reliance on costly 3D data by using foundational models. (8/10)
110
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
InteractVLM also shows strong outperformance on object affordance prediction on the PIAD dataset. Here affordance is defined as contact probabilities on the object. (7/10)
120
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
InteractVLM significantly outperforms prior work, both qualitatively and quantitatively, on in-the-wild 3D human (binary & semantic) contact prediction on the DAMON dataset. (6/10)
110
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
To bridge this 2D-to-3D gap, we propose "Render-Localize-Lift": - Render: 3D human/object meshes into multiview 2D images. - Localize: A Multiview Localization (MV-Loc) model, guided by VLM tokens, predicts 2D contact masks. - Lift: 2D contact masks to 3D. (5/10)
111
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
How can we infer 3D contact with limited 3D data? InteractVLM exploits foundational models—a VLM & localization model fine tuned to reason about contact. Given an image & prompt, the VLM outputs tokens for localization. But these models work in 2D, while contact is 3D. (4/10)
111
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
Furthermore, simple binary contact (touching “any” object) misses the rich semantics of real multi-object interactions. Thus, we introduce a novel task - "Semantic Human Contact" estimation: predicting contact points on a human related to a specified object. (3/10)
120
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
Precisely inferring where humans contact objects from an image is hard due to occlusion & depth ambiguity. Current datasets of images with 3D contact are small as they’re costly & tedious to create (mocap/manual labeling), limiting performance of contact predictors. (2/10)
110
Sai Kumar Dwivedi @saidwivedi.in · 15/06/2025
Why does 3D human-object reconstruction fail in the wild or get limited to a few object classes? A key missing piece is accurate 3D contact. InteractVLM (#CVPR2025) uses foundational models to infer contact on humans & objects, improving reconstruction from a single image. (1/10)
152
Sai Kumar Dwivedi @saidwivedi.in · 12/05/2025
✨ Happy to be recognised again as an Outstanding Reviewer for #CVPR2025!
020
Sai Kumar Dwivedi @saidwivedi.in · 03/04/2025
Thanks to the workshop organizers: @yixinchen.bsky.social, Baoxiong Jia, @yaoyaofeng.bsky.social, @songyoupeng.bsky.social, Chuhang Zou, @saidwivedi.in, Yixin Zhu, Siyuan Huang! 🙌 And the challenge organizers: Xiongkun Linghu, Tai Wang, Jingli Lin, Xiaojian Ma
010
Sai Kumar Dwivedi @saidwivedi.in · 03/04/2025
📢 Excited to announce the 5th Workshop on 3D Scene Understanding for Vision, Graphics & Robotics at #CVPR2025! We’ll dive into multimodal 3D scene understanding & reasoning with amazing speakers and challenges. @cvprconference.bsky.social More Details: scene-understanding.com.
132
Sai Kumar Dwivedi @saidwivedi.in · 12/03/2025
I've been using GitHub's Lists feature for over a year, and it's seriously underrated! ⭐ It lets you assign labels to all your starred repos, making it super easy to find projects later based on specific fields or topics. No more endless scrolling! Link to my list: github.com/saidwivedi?t...
030
Reposted by Sai Kumar Dwivedi
Dimitris Tzionas @dimtzionas.bsky.social · 26/01/2025
📢 I am #hiring 2x #PhD candidates to work on Human-centric #3D #ComputerVision at the University of #Amsterdam! The positions are funded by an #ERC #StartingGrant. For details and for submitting your application please see: werkenbij.uva.nl/en/vacancies... 🆘 Deadline: Feb 16 🆘
werkenbij.uva.nl
Vacancy — PhD Positions, Project 'Spatiotemporal Reconstruction of Interacting People for Perceiving Systems'
Do you want to help computers see, understand, and assist us, humans, in everyday life?  Are you excited with 3D Machine Perception, 3D Human and Object Understanding, 3D Human Avatars, and Machine Le...
1156
Sai Kumar Dwivedi @saidwivedi.in · 09/01/2025
Thanks for sharing :) @chrisoffner3d.bsky.social can you also please add me to the list? I work on 3D human avatar.
030
Sai Kumar Dwivedi @saidwivedi.in · 08/12/2024
One of the best tutorials for understanding Transformers! 📽️ Watch here: www.youtube.com/watch?v=bMXq... Big thanks to @giffmana.ai for this excellent content! 🙌
youtube.com
[M2L 2024] Transformers - Lucas Beyer
YouTube video by Mediterranean Machine Learning (M2L) summer school
0548
Sai Kumar Dwivedi @saidwivedi.in · 24/11/2024
Would love to be in the list 😃
110
Reposted by Sai Kumar Dwivedi
Michael J. Black @michael-j-black.bsky.social · 20/11/2024
For those who missed this post on the-network-that-is-not-to-be-named, I made public my "secrets" for writing a good CVPR paper (or any scientific paper). I've compiled these tips of many years. It's long but hopefully it helps people write better papers. perceiving-systems.blog/en/post/writ...
perceiving-systems.blog
Writing a good scientific paper
426065