Reposted by Shaokai Ye
✨ Introducing a new #SOTA action recognition large multimodal language model: #LLaVAction!
By @shaokaiye.bsky.social Haozhe Qi, @trackingskills.bsky.social and me!
📝 arxiv.org/abs/2503.18712
🤖 mmathislab.github.io/llavaction/
1/n
mmathislab.github.io
LLaVAction: Video Action Recognition
LLaVAction: evaluating and training multi-modal large language models for action recognition