Sign in

Jim RB

@jbohnslav.bsky.social
512 followers 1.1K following 628 posts

computer vision + machine learning. Perception at Zoox. Prev: Cobot, PhD. Arxiv every day.

PostsRepliesMedia
Jim RB @jbohnslav.bsky.social · 24/07/2025
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning arxiv: arxiv.org/abs/2507.16746 data: huggingface.co/datasets/mul...
arxiv.org
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
Humans often use visual aids, for example diagrams or sketches, when solving complex problems. Training multimodal models to do the same, known as Visual Chain of Thought (Visual CoT), is challenging ...
040
Jim RB @jbohnslav.bsky.social · 24/07/2025
ZEBRA-CoT Dataset for vision-language reasoning where the model *generates images during the CoT*. Example: for geometry problems, it's helpful to draw lines in image space. 182K CoT labels: math, visual search, robot planning, and more. Only downside: cc-by-nc license :(
140
Jim RB @jbohnslav.bsky.social · 23/07/2025
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning arxiv: arxiv.org/abs/2507.14137 code: github.com/valeoai/Franca
arxiv.org
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance of state-of-the-art...
010
Jim RB @jbohnslav.bsky.social · 23/07/2025
Cool technique: RASA, Removal of Absolute Spatial Attributes. They decode grid coords and find the plane in feature space that encodes position. They basically subtract this off, baking it into the last linear layer to leave the forward pass unchanged.
110
Jim RB @jbohnslav.bsky.social · 23/07/2025
Beats or is competitive to SigLIP/2, DinoV2 on linear eval, OOD detection, linear segmentation.
100
Jim RB @jbohnslav.bsky.social · 23/07/2025
Franca Fully open vision encoder. Masks image, encodes patches, then trains student to match teacher's clusters. Key advance: Matryoshka clustering. Each slice of the embedding gets its own projection head and clustering objective. Fewer features == fewer clusters to match.
140
Jim RB @jbohnslav.bsky.social · 15/07/2025
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding arxiv: arxiv.org/abs/2507.098... project: vru-accident.github.io
arxiv.org
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, is a critical challenge for autonomous driving systems, as crashes involving VRUs often result in severe or fatal...
010
Jim RB @jbohnslav.bsky.social · 15/07/2025
VRU-Accident New benchmark of 1K videos, 1K captions, and 6K MCQs from accidents involving VRUs. Example: "why did the accident happen?" "(B): pedestrian moves or stays on the road." Current VLMs get ~50-65% accuracy, much worse than humans (95%).
230
Jim RB @jbohnslav.bsky.social · 15/07/2025
BlindSight: Harnessing Sparsity for Efficient VLMs arxiv: arxiv.org/abs/2507.090...
arxiv.org
BlindSight: Harnessing Sparsity for Efficient VLMs
Large vision-language models (VLMs) enable the joint processing of text and images. However, the inclusion of vision data significantly expands the prompt length. Along with the quadratic complexity o...
000
Jim RB @jbohnslav.bsky.social · 15/07/2025
Side note: I've always liked Pali/Gemma's Prefix-LM masking. Why have causal attention for image tokens?
100
Jim RB @jbohnslav.bsky.social · 15/07/2025
BlindSight AMD paper: they find attention heads often have stereotyped sparsity patterns (e.g. only attending within an image, not across). They generate sparse attention variants for each prompt. Theoretically saves ~35% FLOPs for 1-2% worse on benches.
110
Jim RB @jbohnslav.bsky.social · 11/07/2025
Scaling RL to Long Videos arxiv: arxiv.org/abs/2507.07966 code: github.com/NVlabs/Long-RL
arxiv.org
Scaling RL to Long Videos
We introduce a full-stack framework that scales up reasoning in vision-language models (VLMs) to long videos, leveraging reinforcement learning. We address the unique challenges of long video reasonin...
000
Jim RB @jbohnslav.bsky.social · 11/07/2025
Long-RL Nvidia paper scaling RL to long videos. First trains with SFT on a synthetic long CoT dataset, then does GRPO with up to 512 video frames. Uses cached image embeddings + sequence parallelism, speeding up rollouts >2X. Bonus: code is already up!
120
Jim RB @jbohnslav.bsky.social · 09/07/2025
Skywork-R1V3 Technical Report arxiv: arxiv.org/abs/2507.06167 code: github.com/SkyworkAI/Sk...
arxiv.org
Skywork-R1V3 Technical Report
We introduce Skywork-R1V3, an advanced, open-source vision-language model (VLM) that pioneers a new approach to visual reasoning. Its key innovation lies in effectively transferring reasoning skills f...
000
Jim RB @jbohnslav.bsky.social · 09/07/2025
They identify entropy of "wait" or "alternatively" to be strongly correlated with MMMU. Neat!
120
Jim RB @jbohnslav.bsky.social · 09/07/2025
Fine-tuning the connector at the end gives a point or two of MMMU. I wonder how much of this is benchmaxxing--I haven't seen an additional SFT stage after RL before.
100
Jim RB @jbohnslav.bsky.social · 09/07/2025
They construct their warm-start SFT data with synthetic traces from Skywork-R1V2. GRPO is pretty standard, interesting that they just did math instead of math, grounding, other possible RLVR tasks. Qwen-2.5-Instruct 32B to judges the accuracy of the answer in addition to rule-based verification.
100
Jim RB @jbohnslav.bsky.social · 09/07/2025
Skywork-R1V3: new reasoning VLM with 76% MMMU. InternViT-6B stitched with QwQ-32B. SFT warmup, GRPO on math, then a small SFT fine-tune at the end. Good benches, actual ablations, and interesting discussion. Details: 🧵
110
Jim RB @jbohnslav.bsky.social · 09/07/2025
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning arxiv: arxiv.org/abs/2507.05920 code: github.com/EvolvingLMMs...
arxiv.org
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous visual tokens, many of which are irrelevant to the ...
000
Jim RB @jbohnslav.bsky.social · 09/07/2025
Training: use verl with vLLM for rollouts. Limit image resolution to 1280 visual tokens. Train on 32 H100s. Results: +18 points better on V* compared to Qwen2.5-VL, and +5 points better than GRPO alone.
110
Jim RB @jbohnslav.bsky.social · 09/07/2025
RL: GRPO. Reward: only correct answer, not valid grounding coordinates. Seems weird to not add that though. Data: training subset of MME-RealWorld. Evaluate on V*.
100
Jim RB @jbohnslav.bsky.social · 09/07/2025
Uses Qwen2.5-VL as a base model. The NaViT encoder makes it easy to have many images of different shapes. They use a SFT warm-start, as the VLMs struggled to output good grounding coordinates. They constructed two-turn samples for this.
100
Jim RB @jbohnslav.bsky.social · 09/07/2025
MGPO: multi-turn grounding-based policy optimization. I've been waiting for a paper like this! Trains the LLM to iteratively crop regions of interest to answer a question, and the only reward is the final answer. Details in thread 👇
131
Jim RB @jbohnslav.bsky.social · 08/07/2025
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction arxiv: arxiv.org/abs/2507.02948 code: github.com/hzy138/Drive...
arxiv.org
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
Autonomous driving has seen significant progress, driven by extensive real-world data. However, in long-tail scenarios, accurately predicting the safety of the ego vehicle's future motion remains a ma...
000
Jim RB @jbohnslav.bsky.social · 08/07/2025
Using automatically generated risk category labels and the front-facing view, they have GPT4o caption the scenarios. The metrics are based on caption similarity + classification metrics on riskiness-type.
100
Jim RB @jbohnslav.bsky.social · 08/07/2025
DriveMRP: interesting method to get a VLM to understand BEV maps + driving scenarios They synthesize high-risk scenes derived from NuPlan. They render it as both a bird's eye view image and a front camera view. 👇
100
Jim RB @jbohnslav.bsky.social · 08/07/2025
SeqGrowGraph: Learning Lane Topology as a Chain of Graph Expansions arxiv: arxiv.org/abs/2507.048...
arxiv.org
SeqGrowGraph: Learning Lane Topology as a Chain of Graph Expansions
Accurate lane topology is essential for autonomous driving, yet traditional methods struggle to model the complex, non-linear structures-such as loops and bidirectional lanes-prevalent in real-world r...
000
Jim RB @jbohnslav.bsky.social · 08/07/2025
SeqGrowGraph Instead of segment + postprocess, generate lane graphs autoregressively. Node == vertex in BEV space, edge == control point for Bezier curves. At each step, a vertex is added and the adjacency matrix adds one row + column. They formulate this process as next token prediction. Neat!
110
Jim RB @jbohnslav.bsky.social · 02/07/2025
GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning arxiv: arxiv.org/abs/2507.01006 code: github.com/THUDM/GLM-4....
arxiv.org
GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
We present GLM-4.1V-Thinking, a vision-language model (VLM) designed to advance general-purpose multimodal reasoning. In this report, we share our key findings in the development of the reasoning-cent...
000
Jim RB @jbohnslav.bsky.social · 02/07/2025
Excitingly, in one of their few shown results, multi-domain RL shows positive cross-task-transfer. Training on GUI agent data improves STEM answers, OCR, and grounding.
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
Fun new reward hacks: "for a counting problem, the model would answer “a correct number between 0 and 10”, and for a relativity question about speed, it would answer “a velocity very close to the speed of light” – responses that successfully fool some LLM-based reward models and get high reward."
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
They use a ton of specialized verifiers, judges, and reward models for RL.
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
They find half of all prompts get ~90% accuracy after only 200 training steps. Therefore, they train with an adaptive curriculum, continuously adjusting the difficulty on each iteration. Really wish they had ablations here.
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
They add an additional "boxed" answer format to make it more explicit where the final, final answer is <think> reasoning </think> <answer> answer text <begin_of_box> final answer <end_of_box> </answer> Next paper: <final_final_v2>. LLMs are just like us 😅
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
RL: both RLHF and RLVR with GRPO. They use humans and pass@k from prior checkpoints to judge difficulty. When training with many RL tasks, they found weakness in any one task leads to model collapse for all tasks: "effective RL demands finely tuned, hack-resistant verifiers in every domain"
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
SFT: explicitly a warm start for RL, they only include long CoT data here. They iteratively generate new samples from good RL checkpoints. Train at 32K context length. Later, they say that quality is crucial here, with low quality SFT leading to collapse with RL.
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
Pretraining recipe: stage 1, train all parameters for 120K steps at 8192 sequence length. Stage 2, interleaved + video data at 32K sequence length with tensor parallelism + context parallelism.
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
Pretraining data 📚📷 * Captioning: 10B image-text pairs from the web * Interleaved data: websites, papers, and 100 million digitized books * OCR: 220 million images * Grounding: use GLIPv2 for images, playwright for GUIs * Video: "academic, web, and proprietary sources" * Instruction: 50M samples
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
Model details: Uses AIMv2 as the vision encoder and GLM for the LLM, both unique choices. They add 3D convs to the vision encoder to downsample videos by 2X like Qwen 2-VL.
100
Jim RB @jbohnslav.bsky.social · 02/07/2025
GLM-4.1V-Thinking: new reasoning VLM with heavy emphasis on RL. Tons of hints but few ablations 😞 eg they upweight difficult-but-learnable samples every iteration, but don't show how it compares to baseline. 9B variant beats Qwen2.5-VL-7B on many standard benchmarks. Details in thread 👇
100
Jim RB @jbohnslav.bsky.social · 01/07/2025
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World arxiv: arxiv.org/abs/2506.24102 code: github.com/lxtGH/DenseW...
arxiv.org
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations ...
000
Jim RB @jbohnslav.bsky.social · 01/07/2025
Image-level caption details: stitch all the object captions into one global caption. You could just image + object captions in prompt -> unified caption, but there are too many objects for this to work. So, they crop the image into sub-images first then concat the outputs.
100
Jim RB @jbohnslav.bsky.social · 01/07/2025
Object-level caption details: RAM + APE + SAM for seg. Crop each object + caption with InternVL2.5-78B. Then caption the object in context of its original image. Qwen2.5-VL-72B to verify. Because the object captioning pipeline is so expensive, they train a 3B captioner using 600K examples.
100
Jim RB @jbohnslav.bsky.social · 01/07/2025
DenseWorld-1M: insanely detailed + grounded caption dataset. Synthetic data only: SAM, APE for segmentation. Each crop is captioned and verified. VLMs stitch object captions into huge image captions. Beats Sa2VA on referring expression segmentation. Dataset improves Qwen2.5-VL on VQA benchmarks
arxiv.org
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations ...
231
Jim RB @jbohnslav.bsky.social · 01/07/2025
SAM4D: Segment Anything in Camera and LiDAR Streams arxiv: arxiv.org/abs/2506.21547 project: sam4d-project.github.io
arxiv.org
SAM4D: Segment Anything in Camera and LiDAR Streams
We present SAM4D, a multi-modal and temporal foundation model designed for promptable segmentation across camera and LiDAR streams. Unified Multi-modal Positional Encoding (UMPE) is introduced to alig...
000
Jim RB @jbohnslav.bsky.social · 01/07/2025
SAM4D: promptable 4D seg Data engine uses Grounding-Dino + SAM2 for segmentation + tracking in images. Lidar -> voxels -> ray casting to match to pixels. Clustering to stitch visual masklets into 3D instances. Model is like SAM2 but with a lidar encoder + motion-aware cross attention.
110
Jim RB @jbohnslav.bsky.social · 01/07/2025
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning arxiv: arxiv.org/abs/2506.22434
arxiv.org
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
This work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule-based reinforcement learning for Vision-Language Mo...
000
Jim RB @jbohnslav.bsky.social · 01/07/2025
MiCo: multi-image contrast Simple idea: Input is multiple augmented images, either from video or image edits. Prompt: "are these images the same or different?" Train with GRPO. Large bump in multi-image benchmarks, minor bump in general VQA / hallucination benches.
110
Jim RB @jbohnslav.bsky.social · 01/07/2025
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs arxiv: arxiv.org/abs/2506.21656 project: plan-lab.github.io/projects/spa...
arxiv.org
000
Jim RB @jbohnslav.bsky.social · 01/07/2025
SpatialReasoner-R1 Generates long CoT data with Multi-Model Monte Carlo Tree Search -- multiple candidate models for each step, evaluated by multiple LLM judges. DPO with separate losses on the "descriptive" caption and reasoning. Huge improvements on spatial datasets, good performance on VQA.
121