Sign in

Yash Bhalgat

@ysbhalgat.bsky.social
349 followers 214 following 56 posts

PhD at VGG, Oxford w/ Andrew Zisserman, Andrea Vedaldi, Joao Henriques, Iro Laina. Past: Senior RS Qualcomm #AI #Research, UMich, IIT Bombay. I occasionally post AI memes. yashbhalgat.github.io

PostsRepliesMedia
Yash Bhalgat @ysbhalgat.bsky.social · 23/03/2025
Excited to announce the 1st Workshop on 3D-LLM/VLA at #CVPR2025! 🚀 @cvprconference.bsky.social Topics: 3D-VLA models, LLM agents for 3D scene understanding, Robotic control with language. 📢 Call for papers: Deadline – April 20, 2025 🌐 Details: 3d-llm-vla.github.io #llm #3d #Robotics #ai
061
Reposted by Yash Bhalgat
Andreas Geiger @andreasgeiger.bsky.social · 22/02/2025
Our beginner's oriented accessible introduction to modern deep RL is now published in Foundations and Trends in Optimization. It is a great entry to the field if you want to jumpstart into RL! @bernhard-jaeger.bsky.social www.nowpublishers.com/article/Deta... arxiv.org/abs/2312.08365
26214
Yash Bhalgat @ysbhalgat.bsky.social · 18/02/2025
I think a few things will happen soon: 🚀 Scale beyond 8B 🎯 Multi-modal capabilities ⚡️Faster inference 🔄 Reinforcement learning integration Exciting to see alternatives to autoregressive models succeeding at scale! Paper: ml-gsai.github.io/LLaDA-demo/ (8/8)
ml-gsai.github.io
SOCIAL MEDIA TITLE TAG
SOCIAL MEDIA DESCRIPTION TAG TAG
000
Yash Bhalgat @ysbhalgat.bsky.social · 18/02/2025
Results vs LLaMA3 8B: - Matches/exceeds on most tasks - Better at math & Chinese tasks - Strong in-context learning - Improved dialogue capabilities (7/8) 🧵
100
Yash Bhalgat @ysbhalgat.bsky.social · 18/02/2025
A major result: LLaDA breaks the "reversal curse" that plagues autoregressive models. 🔄 On tasks requiring bidirectional reasoning, it outperforms GPT-4 and maintains consistent performance in both forward/reverse directions. (6/8) 🧵
100
Yash Bhalgat @ysbhalgat.bsky.social · 18/02/2025
For generation, they introduce clever remasking strategies: - Low-confidence remasking: Remask tokens the model is least sure about - Semi-autoregressive: Generate in blocks left-to-right while maintaining bidirectional context (5/8) 🧵
100
Yash Bhalgat @ysbhalgat.bsky.social · 18/02/2025
Training uses random masking ratio t ∈ [0,1] for each sequence. The model learns to predict original tokens given partially masked sequences. No causal masking used. Also enables instruction-conditioned generation with the same technique. No modifications. (4/8) 🧵
100
Yash Bhalgat @ysbhalgat.bsky.social · 18/02/2025
💡Core insight: Generative modeling principles, not autoregression, give LLMs their power. LLaDA's forward process gradually masks tokens while reverse process predicts them simultaneously. This enables bidirectional modeling. (3/8) 🧵
110
Yash Bhalgat @ysbhalgat.bsky.social · 18/02/2025
Key highlights: - Successful scaling of masked diffusion to LLM scale (8B params) - Masking with variable ratios for forward/reverse process - Smart remasking strategies for generation, incl. semi-autoregressive - SOTA on reversal tasks, matching Llama 3 on others (2/8) 🧵
100
Yash Bhalgat @ysbhalgat.bsky.social · 18/02/2025
"LLaDA: Large Language Diffusion Models" Nie et al. Just read this fascinating paper. Scaled up Masked Diffusion Language Models to 8B params, and show that it can match #LLMs (including Llama 3) while solving some key limitations! Let's dive in... 🧵 (1/8) #genai
111
Yash Bhalgat @ysbhalgat.bsky.social · 16/02/2025
Project page: bujiazi.github.io/light-a-vide... Code: github.com/bcmi/Light-A... Could be a game-changer for quick video mood/lighting adjustments without complicated VFX pipelines! 🎬
bujiazi.github.io
Light-A-VideoClick to Play and Loop VideoClick to Play and Loop VideoClick to Play and Loop VideoClick to Play and Loop VideoClick to Play and Loop VideoClick to Play and Loop Video
000
Yash Bhalgat @ysbhalgat.bsky.social · 16/02/2025
The results are pretty good ✨ They can transform regular videos into moody noir scenes, add sunlight streaming through windows, or create cyberpunk neon vibes -- works on everything from portrait videos to car commercials! 🚗
100
Yash Bhalgat @ysbhalgat.bsky.social · 16/02/2025
Technical highlights 🔍: - Consistent Light Attention (CLA) module for stable lighting across frames - Progressive Light Fusion for smooth temporal transitions - Works with ANY video diffusion model (AnimateDiff, CogVideoX) - Zero-shot - no fine-tuning needed!
100
Yash Bhalgat @ysbhalgat.bsky.social · 16/02/2025
New work introduces a training-free method to relight entire videos, while maintaining temporal consistency! 📽️🌅 "Light-A-Video: Training-free Video Relighting via Progressive Light Fusion" Zhou et al. (1/n) 🧵 #genai #ai #research #video
1102
Yash Bhalgat @ysbhalgat.bsky.social · 15/02/2025
Project page: liuisabella.com/RigAnything/ Code: not available yet Really excited to try this out once the code is available!
liuisabella.com
RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets
000
Yash Bhalgat @ysbhalgat.bsky.social · 15/02/2025
Authors claim that the model generalizes well across diverse shapes - from humanoids to marine creatures! And works with real-world images & arbitrary poses. 🤩
100
Yash Bhalgat @ysbhalgat.bsky.social · 15/02/2025
Technical highlights: - BFS-ordered skeleton sequence representation - Autoregressive joint prediction with diffusion sampling - Hybrid attention masking: full self-attention for shape tokens, causal attention for skeleton - e2e trainable pipeline without clustering/MST ops
120
Yash Bhalgat @ysbhalgat.bsky.social · 15/02/2025
Need to rig 3D models? 🦖 New work from UCSD and Adobe: "RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets" Liu et al. tl;dr: reduces rigging time from 2 mins to 2 secs, works on any shape category & doesn't need predefined templates! 🚀
150
Yash Bhalgat @ysbhalgat.bsky.social · 14/02/2025
@sebastianraschka.com this is such an interesting discussion! I haven't tried this myself, but I think this can be analyzed theoretically by looking at the rank of the attention matrix in both cases. I have posted my thoughts on the discussion here: github.com/rasbt/LLMs-f...
github.com
Self attention: Merge Query matrix and Key matrix into a single covariance matrix? · rasbt LLMs-from-scratch · Discussion #517
When compute the context vector in the attention algorithm, three weight matrices were introduced. It has discussed in #454 that the value matrix W_V is not necessary. For the rest two, query matri...
010
Yash Bhalgat @ysbhalgat.bsky.social · 14/02/2025
Interesting how they handle the domain gap between 2D latent space and 3D representations through their three-stage pipeline. The correspondence-aware encoding significantly reduces high-frequency noise while preserving geometry. Project: latent-radiance-field.github.io/LRF/
latent-radiance-field.github.io
Latent Radiance Fields with 3D-aware 2D Representations
010
Yash Bhalgat @ysbhalgat.bsky.social · 14/02/2025
Technical approach: - Correspondence-aware autoencoding to enhance 3D consistency in VAE latent space - Builds 3D representations from 3D-aware 2D features - VAE-Radiance Field alignment to bridge domain gap between latent and image space #nerf #ai #research
130
Yash Bhalgat @ysbhalgat.bsky.social · 14/02/2025
"Latent Radiance Fields with 3D-aware 2D Representations" Zhou et al., #ICLR2025 tl;dr: Novel framework that integrates 3D awareness into VAE latent space using correspondence-aware encoding, enabling high-quality rendered images with ~50% memory savings. (1/n) 🧵
120
Yash Bhalgat @ysbhalgat.bsky.social · 13/02/2025
Project: research.nvidia.com/labs/dir/edg... Training and inference code available here: github.com/NVlabs/EdgeR...
000
Yash Bhalgat @ysbhalgat.bsky.social · 13/02/2025
The architecture uses a lightweight encoder and auto-regressive decoder to compress variable-length meshes into fixed-length codes, enabling point cloud and single-image conditioning. Their ArAE model controls face count for varying detail while preserving mesh topology.
100
Yash Bhalgat @ysbhalgat.bsky.social · 13/02/2025
"EdgeRunner" (#ICLR2025) from #Nvidia & PKU introduces an auto-regressive auto-encoder for mesh generation, supporting up to 4000 faces at 512³ resolution. 🤩 Their mesh tokenization algorithm (adapted from EdgeBreaker) achieves ~50% compression (4-5 tokens per face vs 9), making training efficient.
100
Yash Bhalgat @ysbhalgat.bsky.social · 12/02/2025
Technical highlight: They combine 3D latent diffusion with multi-view conditioning for the base shape, then use 2D normal maps for refinement. The results look way cleaner than previous methods.
000
Yash Bhalgat @ysbhalgat.bsky.social · 12/02/2025
Their two-stage approach: First generate coarse geometry (5s), then add fine details (20s) using normal maps based refinement. Smart way to balance speed and quality.
100
Yash Bhalgat @ysbhalgat.bsky.social · 12/02/2025
Just came across this fascinating paper "CraftsMan3D" - a practical approach to text/image-to-3D generation that mimics how artists actually work! Code available (pretrained models too) 🤩: github.com/wyysf-98/Cra... (1/n) 🧵
110
Yash Bhalgat @ysbhalgat.bsky.social · 10/02/2025
Got me excited for a second here 🫠
000
Yash Bhalgat @ysbhalgat.bsky.social · 29/01/2025
So, what happened this week in #AI?
010
Yash Bhalgat @ysbhalgat.bsky.social · 23/01/2025
(3/n) ⚡️ Speed matters: GSLoc adds just ~180ms overhead while providing substantial accuracy gains. We also provide GSLoc_rel variant for even faster refinement when runtime is critical.
000
Yash Bhalgat @ysbhalgat.bsky.social · 23/01/2025
(2/n) 📈 Results: GSLoc achieves new SOTA on indoor datasets (7Scenes & 12Scenes) and significantly improves accuracy on Cambridge Landmarks. Our one-shot refinement outperforms methods requiring 50+ optimization steps!
100
Yash Bhalgat @ysbhalgat.bsky.social · 23/01/2025
(1/n) 🔑 Key idea: We use 3DGS to render high-quality synthetic images & depth maps, enabling efficient one-shot pose refinement of existing APR and SCR methods. No need for iterative optimization or training specialized feature extractors!
100
Yash Bhalgat @ysbhalgat.bsky.social · 23/01/2025
📢 Paper accepted to #ICLR2025 🎉 "GSLoc: Efficient Camera Pose Refinement via 3D Gaussian Splatting" TL;DR: a novel test-time camera pose refinement framework leveraging 3DGS as the scene representation and MASt3R for 2D matching. 🔗: arxiv.org/abs/2408.11085
172
Yash Bhalgat @ysbhalgat.bsky.social · 21/01/2025
Switzerland is rolling out solar panels... on railway tracks! 🇨🇭🚄 Swiss startup Sun-ways will run a pilot project turning train lines into clean energy highways. #renewable #energy for the win 🤓 www.pv-magazine.com/2024/10/04/s...
pv-magazine.com
Switzerland authorizes removable PV plant on railway track
Swiss startup Sun-ways is planning to build a 18 kW pilot PV system between the racks of a 100-m linear section of a railway line in the Swiss canton of Neuchâtel.
1134
Reposted by Yash Bhalgat
Philip Oldfield @sustainabletall.bsky.social · 13/01/2025
Solar panels are becoming so cheap the Swiss are looking at installing them *between train tracks* !!! www.pv-magazine.com/2024/10/04/s...
38118
Reposted by Yash Bhalgat
Maurice Fallon @mauricefallon.bsky.social · 11/01/2025
🚨🚨🚨 Reminder: closing in 3 weeks time 🚨🚨🚨 Please re-post! Note: Oxford recruits faculty at Associate Professor level - we have no Assistant Professor level.
031
Yash Bhalgat @ysbhalgat.bsky.social · 09/01/2025
Came across this LLM visualisation tool today: bbycroft.net/llm Cool stuff! Let's you visualize each operation or layer in different Transformer architectures, and also explains them on the side. 😍 #llm #visualisation #gpt #ai #transformers
040
Yash Bhalgat @ysbhalgat.bsky.social · 09/01/2025
Didn't know it was that good! Will have to watch it now
000
Reposted by Yash Bhalgat
Michael Niemeyer @miniemeyer.bsky.social · 08/01/2025
For other 3D vision newcomers to blue sky: I highly recommend joning @chrisoffner3d.bsky.social 's list to follow the right people ;): go.bsky.app/Cfm9XFe
2194
Yash Bhalgat @ysbhalgat.bsky.social · 08/01/2025
Color and aspect-ratio control.
000
Yash Bhalgat @ysbhalgat.bsky.social · 08/01/2025
(2/2) This method enables inference-time control over SVG attributes like color palettes, aspect ratios, and stroke density, leveraging the learned representation. Really neat approach for editable text-to-vector generation. :)
100
Yash Bhalgat @ysbhalgat.bsky.social · 08/01/2025
"NeuralSVG: An Implicit Representation for Text-to-Vector Generation" (1/2) Encodes SVGs as implicit neural representations using a small MLP trained with Score Distillation Sampling (SDS). Maps 2D coordinates to shape/color outputs. Dropout-like technique ensures ordered, layered structures.
100
Yash Bhalgat @ysbhalgat.bsky.social · 08/01/2025
Tempted to ask what the idea is 👀
100
Yash Bhalgat @ysbhalgat.bsky.social · 06/01/2025
"AR4D: Autoregressive 4D Generation from Monocular Videos" *without* SDS. Autoregressively generate "3D frames" (aka 3DGS) starting from a canonical space, and using a local deformation field for each frame -- high-quality prompt-aligned generations. #ai #nerf #GenAI #video
010
Yash Bhalgat @ysbhalgat.bsky.social · 06/01/2025
Another gem from Bill Freeman, Katie Bouman & team 🌌 A differentiable rendering framework for direct #exoplanet imaging, leveraging wavefront sensing to refine starlight subtraction. Tested on JWST, it approaches noise limits and reveals faint planets like never before! 🚀 #ai #astronomy
031
Yash Bhalgat @ysbhalgat.bsky.social · 05/01/2025
Uh-oh 😅
020
Yash Bhalgat @ysbhalgat.bsky.social · 05/12/2024
Doesn't opencv provide PnP?
100
Yash Bhalgat @ysbhalgat.bsky.social · 30/11/2024
🙃
010
Yash Bhalgat @ysbhalgat.bsky.social · 29/11/2024
I needed this reminder lol
000