Sign in

Yingtian (David) Tang

@davidtyt.bsky.social
25 followers 15 following 30 posts
PostsRepliesMedia
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
2. Our prompt search space currently contains many people-related attribute categories, so human elements are easily introduced. We're working to fix this in future versions.
010
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
When optimizing unnatural stimuli for pSTS (as shown in thread 8/10), we even tried to suppress human traits, yet eyes and mouths still emerged occasionally.
110
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
Thanks for the question! Two reasons: 1. Many ROIs genuinely prefer human bodies. Meanwhile, we also used lenient ROI masks (the union across subjects), so even canonical motion ROIs like MT include some human-sensitive voxels.
110
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
Other resources: 🤗 Basic encoding model (used for NEvo): huggingface.co/epfl-neuroai... 🤗 Enhanced encoding model: huggingface.co/epfl-neuroai...
huggingface.co
epfl-neuroai/vjepa2-encoder-basic · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
100
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
Acknowledgments w/ Ming Zhou, Sogand Salehi, Amir Zamir, @lisik.bsky.social , @mschrimpf.bsky.social (10/10)
100
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
Looking ahead, two directions excite us: 🗺️ Map how dynamic selectivity transitions across the whole visual system 🎭 Generate *impossible stimuli* (motion + interactions not found in nature) to probe visual selectivity beyond handcrafted videos (9/10)
120
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
Controlled test: fix a non-naturalistic image anchor (two plasticine discs 🟢🟣), optimize only motion. Target pSTS → faces and coordinated interaction emerge Target MT → pure translation/oscillation, no social content Same anchor. Different content. (8/10)
100
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
Ablations confirm every design choice: ✅ dynamic > static (biggest gap in MT/EBA) ✅ two-stage search > single-stage ✅ V-JEPA 2 > CLIP (0.44 vs 0.35 predictivity) ✅ genetic search > hill-climbing / random ❌ BrainDiVE (gradient-based) collapses entirely in this setting (7/10)
110
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
Beyond fixed ROIs, NEvo runs as a searchlight 🔦 scoring overlapping cortical patches along a continuous cortical trajectory. We recover a clean lateral stream gradient: texture (V1) → motion/interaction (EBA/pSTS) → social scenes (aSTS) (6/10)
100
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
The results recover known selectivity: 🙂 faces → FFA 🏞️ scenes → PPA 🚶 bodies → EBA 🌀 motions → MT/V3A 👥 interaction → pSTS Interactive demo at nevo-project.epfl.ch They *outperform prior stimuli* — reaching 📈 99.8% of Moments in Time responses 📈 95.8% of handcrafted localizers (5/10)
120
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
Two design choices matter: 1️⃣ **Two-stage search**: optimize a static anchor image first, then animate it → much more sample-efficient. 2️⃣ **Multi-block V-JEPA 2 features**: each voxel selects its best-predicting block → hierarchical feature modeling (4/10)
110
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
How it works: Given a target ROI, we evolve text prompts over a structured search space (30 attribute categories, 614 options). The optimization loop: 🎬 prompts → videos 🧠 videos → predicted ROI response 🧬 ROI response → evolved prompts (3/10)
110
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
🔗 Website: nevo-project.epfl.ch 📄 arXiv: arxiv.org/abs/2607.02317 🤗 Model card: huggingface.co/epfl-neuroai... (2/10)
nevo-project.epfl.ch
NEvo: Neural-Guided Evolutionary Video Synthesis
110
Yingtian (David) Tang @davidtyt.bsky.social · 07/07/2026
🚨 NEW PREPRINT Videos strongly shape activity across the visual cortex. But can we design videos that maximally drive specific brain regions? We present NEvo 🧬🧠 — a neural-guided evolutionary framework that synthesizes videos to maximally activate target visual ROIs. (1/10)
13010
Yingtian (David) Tang @davidtyt.bsky.social · 02/07/2026
We had a very comprehensive study: 4 neural modalities, 3 aspects of improving brain modeling, 8 large datasets, and 600+ models. Resources are publicly available and off-the-shelf !! Please check the project website.
061
Reposted by Yingtian (David) Tang
Abdulkadir Gokce @akgokce.bsky.social · 02/07/2026
🧠 What's the best way to spend compute & neural data to build brain-aligned models? Our new ICML 2026 paper charts the scaling laws across 8 neural datasets & 600+ vision models. 👇 🌐 multimodal-brain-scaling.epfl.ch w/ @davidtyt.bsky.social & @mschrimpf.bsky.social
194
Yingtian (David) Tang @davidtyt.bsky.social · 26/06/2026
4/ 🔬 MODEL-BASED DISCOVERY vjepa2-basic has already powered model-based stimulus synthesis, leading to scientific discovery in the lateral visual stream. 💥 Our full model-based synthesis preprint drops next week. Stay tuned.
010
Yingtian (David) Tang @davidtyt.bsky.social · 26/06/2026
3/ ⚡ MODEL INPUT / OUTPUT Both models are trained on 3-sec video fMRI data. Input: short video clips
Output: normalized fsaverage5 cortical responses across both hemispheres 20,484 mesh vertices. Short video in.
Whole-cortex prediction out.
110
Yingtian (David) Tang @davidtyt.bsky.social · 26/06/2026
2/ 🚀 TWO MODELS vjepa2-basic
→ Our strong baseline
→ Based on our previous large-scale study of dynamic vision models: dynamic-vision.epfl.ch vjepa2-enhanced
→ Our best encoding model
→ Adds lightweight decoders + bagging for improved prediction performance
110
Yingtian (David) Tang @davidtyt.bsky.social · 26/06/2026
🧠🎬 RELEASE We are releasing two Hugging Face model cards for video brain encoding. 🎥 short video clips → 🧠 brain responses MODEL 1: vjepa2-basic
🔗 huggingface.co/epfl-neuroai... MODEL 2: vjepa2-enhanced
🔗 huggingface.co/epfl-neuroai... Brain encoders. Off the shelf !
huggingface.co
epfl-neuroai/vjepa2-encoder-basic · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
172
Yingtian (David) Tang @davidtyt.bsky.social · 04/08/2025
UPDATE project page: yingtiandt.github.io/dynamic-visi...
yingtiandt.github.io
Many-Two-One
000
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
8/🖼️ Big Picture
 Optimizing to model world dynamics leads to brain-like representations.
 🧠 The visual system isn't a patchwork of modules — it’s a unified system built on shared core principles.
120
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
7/🧠 Finding 4 We introduce task-based functional localization.
 It: 1. Recovers many prior neuroscience results in a unified way 2. Reveals new structure in action understanding pathways A novel scalable approach to functional brain mapping.
110
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
6/🌀 Finding 3 Putting observations together:
 • Single-objective models align with all regions and behaviors • Cortex shows hybrid, smooth representation transitions 💡 A new perspective: the brain may implement a shared feature backbone — reused for diverse tasks, just like a “foundation model”.
110
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
5/🌐 Finding 2.2 These two aren’t isolated — they’re: • Blended across ventral & dorsal streams
 • Smoothly mapped across the cortex So, the visual system isn’t modular — it’s highly distributed, and the classic stream separation theory appears oversimplified.
110
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
4/🌐 Finding 2.1
 So, what does the brain actually compute during dynamic vision? Across 10 cognitive tasks (e.g., pose, social cues, action), just two suffice to explain brain-like representations: • Object form • Appearance-free motion
120
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
3/📊 Finding 1
 ✅ Dynamic models > static image models > classic vision models
 ✅ Across both dorsal & ventral regions
 ✅ Across neural & behavioral alignment Best match to brain: V-JEPA. In general, learning world dynamics give alignment to the whole visual system.
110
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
2/🧪 Approach
 We benchmarked diverse video models, each with a different pretraining objective.
 Then: tested how well they predict human fMRI responses to natural movies.
 🧠 ~10,000 voxels, whole visual system.
110
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
1/🔍 Motivation
 The brain is thought to process vision through two streams:
 🖼 Ventral — objects, form, identity
 🧭 Dorsal — motion, spatial layout, actions Image models explain ventral well.
But: what about dorsal? Can one model do both?
120
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
🚨 New research: Can the brain's complex visual system — ventral & dorsal processing streams — arise from a single goal? We study dynamic vision and reveal how object and motion recognition — long thought to be separate — could emerge from the same underlying goal.
110
Yingtian (David) Tang @davidtyt.bsky.social · 30/07/2025
🧠 NEW PREPRINT Many-Two-One: Diverse Representations Across Visual Pathways Emerge from A Single Objective www.biorxiv.org/content/10.1...
biorxiv.org
Many-Two-One: Diverse Representations Across Visual Pathways Emerge from A Single Objective
How the human brain supports diverse behaviours has been debated for decades. The canonical view divides visual processing into distinct "what" and "where/how" streams – however, their origin and inde...
22811