Fernando Pérez-García @fepegar.com · 30/01/2026COLIPRI is generally superior to concurrent methods across all tasks. This is particularly clear when plugging an MLLM on top of our vision backbone. Our models are particularly stronger at clinical metrics, which are most relevant in practice. 🧵8/12 100
Fernando Pérez-García @fepegar.com · 30/01/2026To overcome this domain shift, we introduce an Opposite Sentence Loss (OSL), a simple but effective mechanism that complementes the contrastive loss and improves our metrics substantially. 🧵7/12 100
Fernando Pérez-García @fepegar.com · 30/01/2026We generate reports during training to ensure that the vision encoder extracts from the image all the information that would be needed for reporting, similar to CapPa (@mtschannen.bsky.social et al., NeurIPS 2023). 🧵5/12 100
Fernando Pérez-García @fepegar.com · 30/01/2026We resampled the volumes to 2-mm isotropic spacing using and used an input size of 160^3. We randomly shuffled and shortened sentences in the reports used for contrastive alignment. We initialised our encoder from CXR-BERT (Boecking, @naotous.bsky.social et al., ECCV 2022). 🧵4/12 110
Fernando Pérez-García @fepegar.com · 30/01/2026We first pre-train our encoder only on images (no reports) sourced from different datasets, using a 3D MAE (Wald et al., CVPR 2025). This allows us to leverage more training data, as we did for Rᴀᴅ-DINO (@fepegar.com et al., Nature Machine Intelligence 2025). 🧵2/12 100
Fernando Pérez-García @fepegar.com · 30/01/2026We are excited to release the weights of @msftresearch.bsky.social's COLIPRI, our 3D vision–language encoder for chest CT scans, on @hf.co 🤗 Model: aka.ms/colipri Demo: aka.ms/colipri-demo Paper: aka.ms/colipri-paper Why does COLIPRI matter? 🧵0/12 👇 132
Fernando Pérez-García @fepegar.com · 21/08/2025Having some fun with DINOv3 and PCA! Although I'm not happy my nose has such a low foreground probability :D 060
Fernando Pérez-García @fepegar.com · 04/12/2024Careful with leftover debugging print statements when pushing 😆 120