Sign in

Hokin

@hokin.bsky.social
48 followers 70 following 92 posts

Philosopher, Scientist, Engineer hokindeng.github.io

PostsRepliesMedia
Hokin @hokin.bsky.social · 02/10/2026
Website: object-permanence.world Paper: arxiv.org/abs/2609.28654 Code: github.com/hokindeng/ob... Training data: huggingface.co/datasets/Hok... Eval data: huggingface.co/datasets/Hok... Model: huggingface.co/Hokin/PWM-WROP Leaderboard: object-permanence.world/leaderboard
object-permanence.world
Training Object Permanence in World Models
150 object-permanence tasks · 1.5M training samples · 300-question exam · 14 models · PWM-WROP
010
Hokin @hokin.bsky.social · 02/10/2026
A 3.5-month-old knows that a ball rolling behind a wall is still there. Today we're releasing a full-stack data infrastructure to teach object permanence to AI: 🧠 150 cognitive tasks scaling to 10K+ diverse samples 📦 A 1.5M-sample training set 🚀 A fine-tuned 16B model that validates the data
111
Hokin @hokin.bsky.social · 20/11/2025
You are such a monster
000
Hokin @hokin.bsky.social · 18/11/2025
Congratulations to @tomerullman.bsky.social on official release! For everyone, my disagreements on this paper has also already been accepted by NeurIPS SpaVLE this year. link: arxiv.org/abs/2510.20835
arxiv.org
Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling
Spatial world models, representations that support flexible reasoning about spatial relations, are central to developing computational models that could operate in the physical world, but their precis...
000
Hokin @hokin.bsky.social · 18/11/2025
Congratulations on the official release! My disagreements on this paper is also already accepted by NeurIPS SpaVLE this year link: arxiv.org/abs/2510.20835
arxiv.org
Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling
Spatial world models, representations that support flexible reasoning about spatial relations, are central to developing computational models that could operate in the physical world, but their precis...
020
Hokin @hokin.bsky.social · 07/11/2025
developmental embodiedment 😎 #DevelopmentalEmbodiedment #GrowAI
020
Hokin @hokin.bsky.social · 07/11/2025
congratulations
010
Hokin @hokin.bsky.social · 07/11/2025
what type of pen are you using
100
Hokin @hokin.bsky.social · 04/11/2025
VMEvalKit is 100% open source. We're building this in public with everyone. Plz join us ‼️ 👉 Slack: join.slack.com/t/growingail... 👉 Early Results: grow-ai-like-a-child.com/video-reason/ 📄 Paper: github.com/hokindeng/VM... 👉 GitHub: github.com/hokindeng/VM... The age of video reasoning is here 🎬🧠
github.com
GitHub - hokindeng/VMEvalKit: This is a framework for evaluating reasoning in foundational Video Models.
This is a framework for evaluating reasoning in foundational Video Models. - GitHub - hokindeng/VMEvalKit: This is a framework for evaluating reasoning in foundational Video Models.
000
Hokin @hokin.bsky.social · 04/11/2025
VMEvalKit is 100% open source. We're building this in public with everyone. Plz join us ‼️ 👉 Slack: join.slack.com/t/growingail... 👉 GitHub: github.com/hokindeng/VM... 👉 Early Results: grow-ai-like-a-child.com/video-reason/ 📄 Paper: github.com/hokindeng/VM... The age of video reasoning is here 🎬🧠
github.com
GitHub - hokindeng/VMEvalKit: This is a framework for evaluating reasoning in foundational Video Models.
This is a framework for evaluating reasoning in foundational Video Models. - GitHub - hokindeng/VMEvalKit: This is a framework for evaluating reasoning in foundational Video Models.
020
Hokin @hokin.bsky.social · 04/11/2025
While failure cases clearly show idiosyncratic patterns 🧩🤔, we currently lack a principled framework to systematically analyze or interpret them 🔍. We invite everyone to explore these examples 🧪, as they may offer valuable clues for future research directions 💡🧠🚀.
120
Hokin @hokin.bsky.social · 04/11/2025
Here is a generated video for solving the Raven's Matrices from video models. For more, checkout grow-ai-like-a-child.com/video-reason/
120
Hokin @hokin.bsky.social · 04/11/2025
Raven's Matrices is the one of standard tasks in testing IQ in humans, which require subjects to find patterns and regularities. Intriguingly, video models are able to solve them quite well !
120
Hokin @hokin.bsky.social · 04/11/2025
Here is an example of testing mental rotation in video models. For more, checkout grow-ai-like-a-child.com/video-reason/
120
Hokin @hokin.bsky.social · 04/11/2025
For testing mental rotation, we give them an {n}-voxel structure with some tilted camera views (20-40° elevation) and ask them to horizontally rotate with exactly 180° azimuth change. The hard part is 1) don't deform 2) rotate the right degree. Interesting, some models are able to do it quite well.
120
Hokin @hokin.bsky.social · 04/11/2025
Here is a video example. For more, checkout grow-ai-like-a-child.com/video-reason/
120
Hokin @hokin.bsky.social · 04/11/2025
For the Sudoku problems, the video models need to fill the gap with the correct number in order to have each row and column all have 1, 2, 3. Surprisingly, this is the easiest task for video model.
120
Hokin @hokin.bsky.social · 04/11/2025
Here is an example of generated video from the models solving the maze problem. Checkout more at grow-ai-like-a-child.com/video-reason/
120
Hokin @hokin.bsky.social · 04/11/2025
In the maze problems, video models need to generate videos where navigate the green dot 🟢 to the red flags 🚩 . And they are also able to do it quite well ~
120
Hokin @hokin.bsky.social · 04/11/2025
Here is a generated video for solving the Chess problem. For more examples, checkout: grow-ai-like-a-child.com/video-reason/
120
Hokin @hokin.bsky.social · 04/11/2025
Let's see some examples. Video models are able to figure out what are the checkmate moves in the following problems.
120
Hokin @hokin.bsky.social · 04/11/2025
Idiosyncratic behavioral patterns exist. For example, Sora-2 somehow figures out how to solve Chess problems. But all other models do not have such ability. Veo 3 and 3.1 actually are able to do mental rotation quite well, but really fail on the maze problems.
120
Hokin @hokin.bsky.social · 04/11/2025
Tasks also exhibit clear difficulty hierarchy, with Sudoku being the easiest and mental rotation being the hardest, across all models.
120
Hokin @hokin.bsky.social · 04/11/2025
Models exhibit clear performance hierarchy with Sora-2 currently being the best model.
120
Hokin @hokin.bsky.social · 04/11/2025
The basic of VMEvalkit is a Task Pair unit: 1️⃣ Initial image: unsolved puzzle 2️⃣ Text instruction: “Solve this ...” 3️⃣ Final image: correct solution (hidden during generation) Models see (1)+(2), we compare their output to (3). Simple and straight-forward ✅
120
Hokin @hokin.bsky.social · 04/11/2025
‼️ Video models start to reason, let's build-in-public scaled eval together 🚀 github.com/hokindeng/VM... (Apache 2.0) offers 1⃣One-click inference across ALL available models 2⃣Unified API & datasets & auto resume + error handling + eval 3⃣Plug new models and tasks in <5 lines of code a thread (1/n)
241
Hokin @hokin.bsky.social · 03/11/2025
Our paper is now available at arxiv.org/abs/2510.20835. For anyone interested, we’d love to hang out and chat 💬🧃 #EmbodiedAI #SpatialReasoning #NeuroAI #CognitiveScience #SpatialReasoning
arxiv.org
Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling
Spatial world models, representations that support flexible reasoning about spatial relations, are central to developing computational models that could operate in the physical world, but their precise mechanistic underpinnings are nuanced by the borrowing of underspecified or misguided accounts of human cognition. This paper revisits the simulation versus rendering dichotomy and draws on evidence from aphantasia to argue that fine-grained perceptual content is critical for model-based spatial reasoning. Drawing on recent research into the neural basis of visual awareness, we propose that spatial simulation and perceptual experience depend on shared representational geometries captured by higher-order indices of perceptual relations. We argue that recent developments in embodied AI support this claim, where rich perceptual details improve performance on physics-based world engagements. To this end, we call for the development of architectures capable of maintaining structured perceptual representations as a step toward spatial world modelling in AI.
010
Hokin @hokin.bsky.social · 03/11/2025
Third, in embodied AIs, explicit simulators (MuJoCo/Isaac/Genesis) are vital but brittle alone. Implicit world models (VIP, R3M, visual pretraining) supply perceptual structure that boosts generalization, long-horizon planning, and sim-to-real.
110
Hokin @hokin.bsky.social · 03/11/2025
However, it's necessary that visual and spatial mental content co-construct conscious experiences rather than run on isolated tracks.
120
Hokin @hokin.bsky.social · 03/11/2025
Second, it makes to sound like the dorsal stream, where the "mujoco" software of our brain lies, almost becomes a "zombie" stream, namely with no participation of our conscious experience.
110
Hokin @hokin.bsky.social · 03/11/2025
This first lies in the different interpretation of us in neuro-clinical literature of aphantasia. People with aphantasia can solve mental rotation yet report no visual imagery. We interpret this as a gating/decoding issue—not absence of "rendering" in the brain.
120
Hokin @hokin.bsky.social · 03/11/2025
We argue an alternative: robust spatial reasoning needs fine-grained perceptual content and higher-order relational indices. There’s no free lunch: coarse abstractions into "language-of-thought" like representations won’t yield human-like spatial competence.
110
Hokin @hokin.bsky.social · 02/11/2025
This almost sounds like there is a "mujoco" and a "blender" in the brain, where we use "mujoco" for simulation and "blender" for rendering.
110
Hokin @hokin.bsky.social · 02/11/2025
Basically, in our brains, we have a physics engine, which support "language-of-thought" like simulations of objects, to help reason about physical problems in the real world. We also have a separate "graphics" engine, where we're able to visualize these simulations on top of our head.
120
Hokin @hokin.bsky.social · 02/11/2025
A recent paper by @tomerullman.bsky.social and Halely Balaban raises the notion of "Physics" versus "Graphics" for physical world reasoning in the human brain [1]. [1] Balaban, Halely, and Tomer D. Ullman. "Physics versus graphics as an organizing dichotomy ..." Trends in Cognitive Sciences (2025).
110
Hokin @hokin.bsky.social · 02/11/2025
Excited to share my essay with @carrot0817.bsky.social, Kaia Gao on the representational substrate of world-reasoning in both humans and machines has been accepted to the SpaVLE Workshop at #NeurIPS2025✨ a thread (1/n)
131
Hokin @hokin.bsky.social · 01/11/2025
Very cool work !
010
Hokin @hokin.bsky.social · 01/11/2025
best Halloween costume this year
010
Hokin @hokin.bsky.social · 30/10/2025
what a precise joke hhhhh
000
Hokin @hokin.bsky.social · 08/09/2025
I would love to review more papers. Could you send me some invites at times?
010
Hokin @hokin.bsky.social · 20/08/2025
I totally agree with you!
010
Reposted by Hokin
Tobias Gerstenberg @tobigerstenberg.bsky.social · 02/08/2025
Josh Tenenbaum's inspiring keynote at #cogsci2025 on growing vs scaling AI, the big questions of cognitive science, and the many open questions for the field.
16414
Reposted by Hokin
Jake Quilty-Dunn @quiltydunn.bsky.social · 02/08/2025
this was a hilarious final slide that just hung up there during q&a
092
Reposted by Hokin
David Barner @drbarner.bsky.social · 02/08/2025
Susan Carey mic-drop at #cogsci2025. "There are no innate concepts: Discuss"
2358
Hokin @hokin.bsky.social · 22/07/2025
I had two Private Equity internships in freshman & sophomore summers and a BCG PTA internship in my junior year. But it's the same thing for people who went straight into a research lab since the first semester of their college.
020
Hokin @hokin.bsky.social · 30/06/2025
🚀 Dive in 👇 🌐 Project: williamium3000.github.io/core-knowledge/ 📄 Paper: arxiv.org/abs/2410.10855 📝 OpenReview: openreview.net/forum?id=EIK6xxIoCB 📊 Dataset: huggingface.co/datasets/williamium/CoreCognition 👨‍💻 Code: tinyurl.com/4ak2ryat 🐣 Twitter: x.com/DengHokin/status/1939207699910058405
030
Hokin @hokin.bsky.social · 30/06/2025
🙌 Super grateful to @fredashi.bsky.social, @marstin.bsky.social @tomerullman.bsky.social @siyuansong.bsky.social @gershbrain.bsky.social @zoryzhang.bsky.social @tonyfeng.bsky.social, and everyone @growai.bsky.social for all the 🔥 discussions and wisdom 🧠
140
Hokin @hokin.bsky.social · 30/06/2025
💥 Huge thanks to my amazing collaborators @williamium3000.bsky.social, Kaia Gao, Icy Wang, Tianwei Zhao, Haoran Sun, Zoey Lyu, Robert Hawkins, Nuno Vasconcelos, @talgolanneuro.bsky.social & @carrot0817.bsky.social 💐
140
Hokin @hokin.bsky.social · 30/06/2025
7) The result is astonishing 😱 No models are able to get both the manipulation tasks and the control tasks right at the same time, which seems to completely trivial to humans ... This suggests MLLMs completely lack core knowledge 🦜 and purely rely on shortcuts 🤫 ... (11/n)
130
Hokin @hokin.bsky.social · 30/06/2025
6) Last, we introduce "Concept Hacking" to reveal core knowledge deficiencies in the control experiment set-up. Concept Hacking systematically manipulates the task-relevant features while preserving all task-irrelevant conditions ... (10/n)
130