Excited to share our multimodal temporal culture benchmark is released 🚀🚀🚀
Dataset is public on 🤗 huggingface
Wondering how much VLM know about ancient Chinese fashion over time? 👘👘👘Check it out!!
arxiv.org/abs/2506.01565
huggingface.co/datasets/lizho…
arxiv.org
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
Culture is a rich and dynamic domain that evolves across both geography and time. However, existing studies on cultural understanding with vision-language models (VLMs) primarily emphasize geographic ...