Sign in

arXiv cs.CV Computer Vision and Pattern Recognition

@cscv-bot.bsky.social
263 followers 1 following 106K posts

Unofficial bot by @vele.bsky.social w/ github.com/so-okada/bXiv arxiv.org/list/cs.CV/new List bsky.app/profile/vele.bsky.social/l… ModList bsky.app/profile/vele.bsky.social/l…

PostsRepliesMedia
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Hao-Tang Tsui, Yu-Rou Tuan, Xiaoxuan Ma, Nicolas Ugrinovic, Takaaki Shiratori, Kris Kitani: Point2Part: Unified 3D Partitioning from Point Prompts arxiv.org/abs/2609.38180 arxiv.org/pdf/2609.38180 arxiv.org/html/2609.38180
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Jung, Yu, An, Han, Kim, Jeon, Shin, Moon, Tombari, Barath, Pollefeys, Kim, Hong: Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering arxiv.org/abs/2609.38177 arxiv.org/pdf/2609.38177 arxiv.org/html/2609.38177
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen: Adversarial Training for Pixel Diffusion arxiv.org/abs/2609.38170 arxiv.org/pdf/2609.38170 arxiv.org/html/2609.38170
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini: Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Ti... arxiv.org/abs/2609.38165 arxiv.org/pdf/2609.38165 arxiv.org/html/2609.38165
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Jiang, Liu, Wang, Sun, Chen, Wang, Wang, Chen, Yao, Zhao, Yuan, Su, Sui, Liu, Wang: Rethinking Representations for World-Action Modeling arxiv.org/abs/2609.38163 arxiv.org/pdf/2609.38163 arxiv.org/html/2609.38163
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Lin, Zhang, Zhou, Zheng, Liu, Yang, Lin, Yang, Nguyen: DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses arxiv.org/abs/2609.38156 arxiv.org/pdf/2609.38156 arxiv.org/html/2609.38156
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua: Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies arxiv.org/abs/2609.38155 arxiv.org/pdf/2609.38155 arxiv.org/html/2609.38155
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Yang, Wang, Huang, Chen, Zhang, Fu, Ma, Lin, Mao, Chu, Han, Chen: LongLive-Plug: Once-for-All Distillation for Video Generation arxiv.org/abs/2609.38154 arxiv.org/pdf/2609.38154 arxiv.org/html/2609.38154
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Trong-Tung Nguyen, Anand Bhattad: PowerSim: Differentiable Physics Simulation and Rendering with Power Diagrams arxiv.org/abs/2609.38153 arxiv.org/pdf/2609.38153 arxiv.org/html/2609.38153
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Trong-Tung Nguyen, Jiahan Zhang, Anand Bhattad: FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation arxiv.org/abs/2609.38152 arxiv.org/pdf/2609.38152 arxiv.org/html/2609.38152
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Ji, Wang, Xu, Li, Mao, Chen, Shan, Zhang, Hua, Xie, Cheng, Tu: LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation arxiv.org/abs/2609.38146 arxiv.org/pdf/2609.38146 arxiv.org/html/2609.38146
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang: Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE arxiv.org/abs/2609.38140 arxiv.org/pdf/2609.38140 arxiv.org/html/2609.38140
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Teng Zhou, Yunhao Chen: CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer arxiv.org/abs/2609.38136 arxiv.org/pdf/2609.38136 arxiv.org/html/2609.38136
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Ke, Pan, Tian, Wang, Huang, Wu, Fang, Liu, Qi, Wang, Yuan, Chen, Zhan, Chen, Xue, Guo: HelixWorld: A Real-time Interactive Audio-Visual World Model arxiv.org/abs/2609.38123 arxiv.org/pdf/2609.38123 arxiv.org/html/2609.38123
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Jinfa Huang, Jianming Xu, Jingyang Lin, Zhengyuan Yang, Jiebo Luo: VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents arxiv.org/abs/2609.38119 arxiv.org/pdf/2609.38119 arxiv.org/html/2609.38119
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Ahmed, Casado, Castro, Sharifipour, Kumar, L\'opez: GA-EIRFS: A Geometry-Augmented Repeat-Factor Sampling Method for Long-Tailed LiDAR 3D Object Detection arxiv.org/abs/2609.38116 arxiv.org/pdf/2609.38116 arxiv.org/html/2609.38116
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Weiqiang Wang, Zhuokun Chen, Yusheng Dai, Boying Li, Yi Zhang, Hossein Rahmani, Qiuhong Ke, Jianfei Cai: Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History arxiv.org/abs/2609.38114 arxiv.org/pdf/2609.38114 arxiv.org/html/2609.38114
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Jiang, Qian, Chen, Li, Li, Li, Liu, Sun: VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents arxiv.org/abs/2609.38086 arxiv.org/pdf/2609.38086 arxiv.org/html/2609.38086
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Ge, Qin, Xie, Jiang, Han, Zhang, Dai, Yang, Malik, Krishna, Min, Feng, Xue, Shi, Darrell, Wang: OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? arxiv.org/abs/2609.38079 arxiv.org/pdf/2609.38079 arxiv.org/html/2609.38079
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Jiaming Tan, Zhen Li, Shuwei Shi, Minggui Teng, Siqi Yang, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang: MUGEN: Interactive Panoramic World Exploration via Camera Control arxiv.org/abs/2609.38077 arxiv.org/pdf/2609.38077 arxiv.org/html/2609.38077
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li: RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA arxiv.org/abs/2609.38072 arxiv.org/pdf/2609.38072 arxiv.org/html/2609.38072
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Zhou, Wu, Li, Zheng, Zhang, Yu, Yang, Liu, Sun, Yang, Jiang, Su, Huang, Tian: EVO-WAM: Evolving World Action Models through Video-Action Verification arxiv.org/abs/2609.38057 arxiv.org/pdf/2609.38057 arxiv.org/html/2609.38057
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Bangxun Tang: Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation arxiv.org/abs/2609.38019 arxiv.org/pdf/2609.38019 arxiv.org/html/2609.38019
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Benkedadra, Saoudi, Gloesener, Mahmoudi, Mancas: From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection arxiv.org/abs/2609.38010 arxiv.org/pdf/2609.38010 arxiv.org/html/2609.38010
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Chen, Niu, Lu, Lian, Tang, Yan, Hong, Du, Liu, Chen, Shen: HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents arxiv.org/abs/2609.38008 arxiv.org/pdf/2609.38008 arxiv.org/html/2609.38008
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Xuyi Hu, Francesco Palandra, Shangzhe Wu, Daniel Cremers, Riccardo Marin, Silvia Zuffi: ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals arxiv.org/abs/2609.37986 arxiv.org/pdf/2609.37986 arxiv.org/html/2609.37986
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Liu, Ye, Xue, Li, Chen, Li, Yu, Wang, Zhang, Zhu, Han, Xie: SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video arxiv.org/abs/2609.37969 arxiv.org/pdf/2609.37969 arxiv.org/html/2609.37969
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Tan, Qi, Liu, Wen, Liu, Wang, Chen, Zheng, Wu, Zhang, Chen, Sun, Wu, Timofte, Paudel, Peng: Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark arxiv.org/abs/2609.37938 arxiv.org/pdf/2609.37938 arxiv.org/html/2609.37938
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao: Look Closer: Patch-wise Supervision for AI-Generated Image Detection arxiv.org/abs/2609.37937 arxiv.org/pdf/2609.37937 arxiv.org/html/2609.37937
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang, Weidong Zhang, Tianfan Xue: Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation arxiv.org/abs/2609.37925 arxiv.org/pdf/2609.37925 arxiv.org/html/2609.37925
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Ziyun Zeng, Hang Hua, Shaden Alshammari, Rogerio Feris, William T. Freeman, Jiebo Luo: EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory arxiv.org/abs/2609.37923 arxiv.org/pdf/2609.37923 arxiv.org/html/2609.37923
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami: SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation arxiv.org/abs/2609.37918 arxiv.org/pdf/2609.37918 arxiv.org/html/2609.37918
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Tao Hu, Zhinuo Zhou, Xialiang Tong, De-Chuan Zhan, Da-Wei Zhou: ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning arxiv.org/abs/2609.37889 arxiv.org/pdf/2609.37889 arxiv.org/html/2609.37889
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Tao Hu, Zhen-Hao Xie, Jingcai Guo, De-Chuan Zhan, Da-Wei zhou: Visual Branch is What You Need for CLIP-based Class-Incremental Learning arxiv.org/abs/2609.37888 arxiv.org/pdf/2609.37888 arxiv.org/html/2609.37888
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Jiaqi Huang, Shidong Wang, Tong Xin, Kabita Adhikari: EndoPrior-GS: Dynamic Endoscopic Reconstruction with a Joint Texture Prior arxiv.org/abs/2609.37874 arxiv.org/pdf/2609.37874 arxiv.org/html/2609.37874
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Zhixiang Hao, Shaodi You, Yu Li, Kunming Li, Feng Lu: Learning from synthetic photorealistic raindrop for single image raindrop removal arxiv.org/abs/2609.37870 arxiv.org/pdf/2609.37870 arxiv.org/html/2609.37870
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Fan Zhou, Shuairan Chen, Mengying Zhang, Yulin Wu, Sadegh Jafari, Sixing Yu, Rui Li, Ali Jannesari, Guowen Song: HandAnthro: Automated Hand Anthropometry from a Single Image arxiv.org/abs/2609.37855 arxiv.org/pdf/2609.37855 arxiv.org/html/2609.37855
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Zhiqi Li, Bo Zhu: FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators arxiv.org/abs/2609.37851 arxiv.org/pdf/2609.37851 arxiv.org/html/2609.37851
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Xijun Wang, Xin Li, Zirui Lang, Suhang Yao, Haoran Li, Zhibo Chen: RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution arxiv.org/abs/2609.37850 arxiv.org/pdf/2609.37850 arxiv.org/html/2609.37850
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala: Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study arxiv.org/abs/2609.37848 arxiv.org/pdf/2609.37848 arxiv.org/html/2609.37848
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Xijun Wang, Xin Li, Suhang Yao, Zirui Lang, Bingchen Li, Zhibo Chen: ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing arxiv.org/abs/2609.37831 arxiv.org/pdf/2609.37831 arxiv.org/html/2609.37831
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Zhongping Ji: Minkowski Attractor Networks: Closed-Form Hyperbolic Flows for Visual Representations arxiv.org/abs/2609.37817 arxiv.org/pdf/2609.37817 arxiv.org/html/2609.37817
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
No\'e Lallouet, Michael Fischer, Elie Michel: WINGS: Reference-Free Gaussian Splatting Inpainting with 3D-Native Generative Priors arxiv.org/abs/2609.37816 arxiv.org/pdf/2609.37816 arxiv.org/html/2609.37816
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Sven Ligensa, Jan Pauls, Karsten Schr\"odter, Ibrahim Fayad, Fabian Gieseke: Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study arxiv.org/abs/2609.37809 arxiv.org/pdf/2609.37809 arxiv.org/html/2609.37809
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Thomas A. O'Shea-Wheller: ByteTraX: Enhancing the ByteTrack Architecture with Optimised Thresholding arxiv.org/abs/2609.37801 arxiv.org/pdf/2609.37801 arxiv.org/html/2609.37801
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
R\'emi Kazmierczak, Johanne Cohen, Marianne Clausel: CHOQOLATE: Organizing Concept Bottleneck Latent Spaces with Choquet Integrals arxiv.org/abs/2609.37786 arxiv.org/pdf/2609.37786 arxiv.org/html/2609.37786
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Arjun Rao, Sebastian Loeschcke, Anthony Fuller, Isaac Corley, Nico Lang, Evan Shelhamer: Planetary Feature Fields are Scalable Earth Representations arxiv.org/abs/2609.37784 arxiv.org/pdf/2609.37784 arxiv.org/html/2609.37784
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Kelly McConvey, et al.: A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System arxiv.org/abs/2609.37783 arxiv.org/pdf/2609.37783 arxiv.org/html/2609.37783
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Tengfei Liu, Jing Jin, Yuan Zhou: HiRAE: Hierarchical Representation Autoencoding with Residual Budgets arxiv.org/abs/2609.37775 arxiv.org/pdf/2609.37775 arxiv.org/html/2609.37775
000
arXiv cs.CV Computer Vision and Pattern Recognition @cscv-bot.bsky.social · 13h
Aawez Mansuri, et al.: Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection arxiv.org/abs/2609.37750 arxiv.org/pdf/2609.37750 arxiv.org/html/2609.37750
000