Sign in

Jim RB

@jbohnslav.bsky.social
512 followers 1.1K following 628 posts

computer vision + machine learning. Perception at Zoox. Prev: Cobot, PhD. Arxiv every day.

PostsRepliesMedia
Jim RB @jbohnslav.bsky.social · 24/07/2025
ZEBRA-CoT Dataset for vision-language reasoning where the model *generates images during the CoT*. Example: for geometry problems, it's helpful to draw lines in image space. 182K CoT labels: math, visual search, robot planning, and more. Only downside: cc-by-nc license :(
140
Jim RB @jbohnslav.bsky.social · 23/07/2025
Franca Fully open vision encoder. Masks image, encodes patches, then trains student to match teacher's clusters. Key advance: Matryoshka clustering. Each slice of the embedding gets its own projection head and clustering objective. Fewer features == fewer clusters to match.
140
Jim RB @jbohnslav.bsky.social · 15/07/2025
VRU-Accident New benchmark of 1K videos, 1K captions, and 6K MCQs from accidents involving VRUs. Example: "why did the accident happen?" "(B): pedestrian moves or stays on the road." Current VLMs get ~50-65% accuracy, much worse than humans (95%).
230
Jim RB @jbohnslav.bsky.social · 15/07/2025
BlindSight AMD paper: they find attention heads often have stereotyped sparsity patterns (e.g. only attending within an image, not across). They generate sparse attention variants for each prompt. Theoretically saves ~35% FLOPs for 1-2% worse on benches.
110
Jim RB @jbohnslav.bsky.social · 11/07/2025
Long-RL Nvidia paper scaling RL to long videos. First trains with SFT on a synthetic long CoT dataset, then does GRPO with up to 512 video frames. Uses cached image embeddings + sequence parallelism, speeding up rollouts >2X. Bonus: code is already up!
120
Jim RB @jbohnslav.bsky.social · 09/07/2025
Skywork-R1V3: new reasoning VLM with 76% MMMU. InternViT-6B stitched with QwQ-32B. SFT warmup, GRPO on math, then a small SFT fine-tune at the end. Good benches, actual ablations, and interesting discussion. Details: 🧵
110
Jim RB @jbohnslav.bsky.social · 09/07/2025
MGPO: multi-turn grounding-based policy optimization. I've been waiting for a paper like this! Trains the LLM to iteratively crop regions of interest to answer a question, and the only reward is the final answer. Details in thread 👇
131
Jim RB @jbohnslav.bsky.social · 08/07/2025
DriveMRP: interesting method to get a VLM to understand BEV maps + driving scenarios They synthesize high-risk scenes derived from NuPlan. They render it as both a bird's eye view image and a front camera view. 👇
100
Jim RB @jbohnslav.bsky.social · 08/07/2025
SeqGrowGraph Instead of segment + postprocess, generate lane graphs autoregressively. Node == vertex in BEV space, edge == control point for Bezier curves. At each step, a vertex is added and the adjacency matrix adds one row + column. They formulate this process as next token prediction. Neat!
110
Jim RB @jbohnslav.bsky.social · 02/07/2025
GLM-4.1V-Thinking: new reasoning VLM with heavy emphasis on RL. Tons of hints but few ablations 😞 eg they upweight difficult-but-learnable samples every iteration, but don't show how it compares to baseline. 9B variant beats Qwen2.5-VL-7B on many standard benchmarks. Details in thread 👇
100
Jim RB @jbohnslav.bsky.social · 01/07/2025
DenseWorld-1M: insanely detailed + grounded caption dataset. Synthetic data only: SAM, APE for segmentation. Each crop is captioned and verified. VLMs stitch object captions into huge image captions. Beats Sa2VA on referring expression segmentation. Dataset improves Qwen2.5-VL on VQA benchmarks
arxiv.org
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations ...
231
Jim RB @jbohnslav.bsky.social · 01/07/2025
SAM4D: promptable 4D seg Data engine uses Grounding-Dino + SAM2 for segmentation + tracking in images. Lidar -> voxels -> ray casting to match to pixels. Clustering to stitch visual masklets into 3D instances. Model is like SAM2 but with a lidar encoder + motion-aware cross attention.
110
Jim RB @jbohnslav.bsky.social · 01/07/2025
MiCo: multi-image contrast Simple idea: Input is multiple augmented images, either from video or image edits. Prompt: "are these images the same or different?" Train with GRPO. Large bump in multi-image benchmarks, minor bump in general VQA / hallucination benches.
110
Jim RB @jbohnslav.bsky.social · 01/07/2025
SpatialReasoner-R1 Generates long CoT data with Multi-Model Monte Carlo Tree Search -- multiple candidate models for each step, evaluated by multiple LLM judges. DPO with separate losses on the "descriptive" caption and reasoning. Huge improvements on spatial datasets, good performance on VQA.
121
Jim RB @jbohnslav.bsky.social · 25/06/2025
ScaleCap: synthetic image captioning pipeline. 2542 characters per cap. Given a caption, generates followup questions and answers with a VLM. Compute P(sentence | image,prompt) - P(sentence|prompt). Sentences with low scores are only using their text prior, so filter them out.
100
Jim RB @jbohnslav.bsky.social · 25/06/2025
UniVLA: VQA, world modeling, and robotic controls all-in-one transformer VQ-quantize images, DCT robotic actions, standard text tokens. Train the whole interleaved sequence with NTP. This type of approach would benefit from serious scaling. Anyone have a few thousand H100s laying around?
163
Jim RB @jbohnslav.bsky.social · 08/04/2025
ST-Kit: new dataset + benchmark for kinematic understanding for VLMs. Examples: estimate the total distance covered by <object> in the video? In what direction does <object> move in BEV coordinates? Open source and closed source models do poorly but their fine-tune does well.
100
Jim RB @jbohnslav.bsky.social · 08/04/2025
Vision-R1: another RLVR for VLMs, this time for detection. Rewards: formatting, precision, recall. Adds +9mAP to Qwen-2.5-VL on COCO + ODINW 🤯. They change the threshold in a curriculum (easy -> hard) which adds ~2 points. Repo has training code based on R1-V and TRL 👍
100
Jim RB @jbohnslav.bsky.social · 08/04/2025
CAR-1000 Move over Stanford Cars, new dataset with 1000 hierarchical classes and 140,312 samples! Scraped from a Chinese car enthusiast forum, there's a wide variety of cars from around the world.
100
Jim RB @jbohnslav.bsky.social · 08/04/2025
NuPlanQA: uses nuPlan annotations + GPT4o to 1 million QA pairs for training and an 8K multiple-choice benchmark. BEV-LLM: baseline model. Multi-view video -> BEVFusion -> cross-attend with image features -> projector -> LLaMA3.2. Some helpful ablations on # views and frames.
100
Jim RB @jbohnslav.bsky.social · 03/04/2025
Hydra-NeXt: strong E2E closed-loop driving performance with only open-loop training Images -> encoder -> 4096 discrete trajectory vocab -> transformer -> bicycle model -> denoising diffusion refinement -> best trajectory selection. Much better closed-loop perf than UniAD, VAD
100
Jim RB @jbohnslav.bsky.social · 03/04/2025
Nice methods paper (/ advertisement) from Nvidia on training video models with NeMo. Nuts and bolts: use Ray + NVDEC for curation, S3 + WebDataset for dataloading, FSDP + TP + Context Parallelism + PP for video DiT training. 48.2% MFU
100
Jim RB @jbohnslav.bsky.social · 03/04/2025
ChatBEV: new dataset + benchmark for VQA on BEV maps for autonomous driving. Data from nuPlan. 116K train, 21K test. Render a BEV map + generate VQA with LLMs(?) Applications: VQA, use the model to condition a scene generator.
100
Jim RB @jbohnslav.bsky.social · 27/03/2025
TrajHF: diffusion-based planner fine-tuned with RLHF. Simple idea: humans have preferences for candidate trajectories. Collect human feedback on key frames from 4.5K clips mined for aggressive maneuvers. Improves these aggressive scenarios but worsens some open-loop metrics.
100
Jim RB @jbohnslav.bsky.social · 27/03/2025
DriveLMM-o1: new dataset + benchmark for autonomy VQA. Dataset is 18k QA pairs, each with step-by-step reasoning. Generated with GPT4o and human verified. Model is a fine-tuned InternVL2.5-8b. Nit: don't call your model o1 if you don't use RLVR!
100
Jim RB @jbohnslav.bsky.social · 20/03/2025
SimLingo: Wayve paper on VLAs for self driving Front camera + prompts -> InternVL2-0.5B. They add GPS targets with an MLP. Output language, and decode queries for waypoints + speed. Very strong CARLA performance, but CARLA-only training data.
110
Jim RB @jbohnslav.bsky.social · 20/03/2025
DeCapBench: new benchmark on super-detailed captioning. DCScore: new evaluation method for captioning, great correlation with human ratings. FeedQuill: VLM trained using PPO with a DScore-trained reward model. Great captioning specialist.
110
Jim RB @jbohnslav.bsky.social · 20/03/2025
DriveTransformer: E2E driving with a single (complex) transformer, not modular like UniAD. Beats UniAD, VAD in open-loop but by much more in closed-loop in CARLA. It's also more robust to camera perturbations.
100
Jim RB @jbohnslav.bsky.social · 19/03/2025
CLIPGrader: fine-tune CLIP to evaluate detection labels. Draw a magenta bbox on an image. Artificially perturb the box and make the caption "the magenta bounding box is a bad bounding box of a..." 91% accuracy in quality classification. Examples show lots of promise.
110
Jim RB @jbohnslav.bsky.social · 13/03/2025
AlphaDrive: Trains a reasoning VLM to output multiple discrete action plans (accelerate, turn left) for autonomous driving. Much better than zero-shot or SFT on MetaAD, a new dataset of 110K 3s clips. In ablations, SFT < RL < SFT + RL. Looks like the days of pure SFT are over!
100
Jim RB @jbohnslav.bsky.social · 12/03/2025
Visual-RFT: RL + verifiable rewards + GRPO works for vision tasks. Apply to detection (~IoU reward) and fine-grained classification (accuracy reward). Huge gains over Qwen2VL + SFT: +~25 accuracy, +11-27 mAP in open-vocab / few-shot detection.
141
Jim RB @jbohnslav.bsky.social · 12/03/2025
Sce2DriveX: uses multi-view videos + rendered BEV maps to drive with VLMs. Uses a DriveLM-style graph to encode scene information. Makes an SFT dataset focused on 3D understanding. Good performance on NuScenes motion planning.
100
Jim RB @jbohnslav.bsky.social · 11/03/2025
Jumbo: improving ViTs with a jumbo class token. Instead of registers, which are multiple "dummy" tokens that go through the same FFN, they make one large learnable register. Split it at the beginning, then concat it at the end. Goes through its own FFN instead of the shared one
110
Jim RB @jbohnslav.bsky.social · 11/03/2025
Magma: First paper I've seen to focus on both GUI agents and robotics. Trains on lots of egocentric video + robotics simulation. Makes heavy use of set-of-mark and trace-of-mark visual prompting. Decent performance in VQA, GUI use, robotics, but beaten by experts in every case.
110
Jim RB @jbohnslav.bsky.social · 11/03/2025
DiCeption: diffusion generalist for multiple dense tasks. Trains a single DiT to do monodepth, normals, segmentation, point-prompted segmentation. Strong zero-shot performance, even beating specialists like DepthAnythingv2, albeit with 28 diffusion steps at inference
100
Jim RB @jbohnslav.bsky.social · 11/03/2025
New paper introduces two new visual perception tokens 1. Region Selection Token: instructs the model to crop a bbox, encode it, and pass it back through the LLM 2. Vision Re-Encoding Token: instructs the model to re-encode the image with DINOv2 and pass it back through the LLM
100
Jim RB @jbohnslav.bsky.social · 11/03/2025
VaViM: an open-source autoregressive video model VaVAM: video action model for self-driving VaVAM is a flow-matching VLA head on the VaViM backbone. Attends to a high level text command, input waypoints, and causal video features via attention masking.
100
Jim RB @jbohnslav.bsky.social · 11/03/2025
BEVDiffuser: BEV map denoising significantly improves 3D detection performance. Adds noise to off-the-shelf BEV maps, eg BEVFusion. You condition on labels, so the model learns to focus on real objects. At inference time, you either use the denoiser or distill the denoised maps
100
Jim RB @jbohnslav.bsky.social · 25/02/2025
ChatVLA VLAs catastrophically forget general knowledge. They add VQA data to training to mitigate this. Second, they use a two-expert MoE, VQA and control. The head is a DiVLA text-conditioned-diffusion head. Strong improvements over OpenVLA / Octo.
110
Jim RB @jbohnslav.bsky.social · 25/02/2025
SigLIP 2: another gift from GDM Zurich. Goal was to keep zero-shot / retrieval capabilities while improving localization + dense features. Adds captioning loss, masked loss, & self-distillation but keeps arch. the same for drop-in weights. Great improvements across the board.
100
Jim RB @jbohnslav.bsky.social · 25/02/2025
Qwen2.5-VL technical report * from-scratch ViT with window attention * dynamic FPS for video * scaled pretraining from 1.2T to 4.1T tokens The lack of detail is a little sad: "throughout the training process, we... adjusted the composition and proportions of these data types."
100
Jim RB @jbohnslav.bsky.social · 19/02/2025
DART: duplication-aware reduction of tokens Turns out using importance to select tokens for VLMs is often worse than random. Instead, drop duplicates! They pick top 2% "pivot tokens" and select the N tokens with lowest cosine similarity. 2X throughput at 97.1% of original POPE
100
Jim RB @jbohnslav.bsky.social · 19/02/2025
CloCKDistill: knowledge distillation for DETRs Instead of naively distilling feature maps, they use GT boxes to upweight object features, more for small objects. They also distill the decoded bounding boxes. Improves deformable DETR by +2AP, and up to +6 for small objects
100
Jim RB @jbohnslav.bsky.social · 19/02/2025
V2V-LLM Combines vehicle-to-vehicle communication + LLMs for autonomous driving. Each AV detects objects + encodes them LLaVA-style + can ask and answer each others' questions. They also make a new dataset with 500K+ VQA pairs.
100
Jim RB @jbohnslav.bsky.social · 19/02/2025
SimDINO: simplifying DINO/ DINOv2. They use L2 loss on global <-> local features and a penalty on feature covariance to prevent collapse. With these additions, they can get rid of a lot of the bells and whistles in DINO/v2 training while improving representations (ImageNet-KNN).
110
Jim RB @jbohnslav.bsky.social · 19/02/2025
MM-RLHF, new dataset of 120K human ratings + explanation of responses. Using this to train a 7B reward model beats LlavaCritic, even GPT4o on RewardBench. Using that reward model + DPO improves LLaVA-OV baselines improves both performance and safety.
111
Jim RB @jbohnslav.bsky.social · 12/02/2025
Google trained SigLIP models with 1B, 10B, and 100B samples🤯 Performance saturates from 10B to 100B on Western evals (imagenet) while improving nonwestern evals (e.g. img2txt recall in Tegulu, a language representing <0.04% of the internet). Checkpoints on huggingface when?
110
Jim RB @jbohnslav.bsky.social · 11/02/2025
DexVLA: Diffusion Expert for Vision-Language-Action models Images -> Qwen2-VL -> reasoning + action tokens -> projector -> 1B diffusion model. Automatically breaks down tasks into substeps that are easier for the diffusion policy. Beats OpenVLA. Hard to compare to pi0 🤖
100
Jim RB @jbohnslav.bsky.social · 11/02/2025
OccGS: new method for zero-shot semantic 3D occupancy. Image + lidar -> grounding dino + sam for zero-shot panoptic seg -> gaussian splat + P(classes) -> splat to voxels. Results beat many fully supervised approaches on Occ3D-nuscenes!
191
Jim RB @jbohnslav.bsky.social · 07/02/2025
SMART Approach to use OpenStreetMap + satellite imagery to learn a prior for HD maps. When integrated with on-vehicle sensor data, improves OpenLaneV2 performance by huge margins.
121