Sign in

Haoli Yin

@haoliyin.bsky.social
307 followers 253 following 28 posts

multimodal data curation @datologyai.com. haoliyin.me

PostsRepliesMedia
Haoli Yin @haoliyin.bsky.social · 08/12/2024
For more details about where I'll be, come visit me at @datologyai.com's booth 303 Times I’ll be there (in local time): - Tuesday Dec 10th, 12pm-4pm - Wednesday Dec 11th, 1pm-5pm - Thursday Dec 12th, 9am-12:30pm #neurips
050
Haoli Yin @haoliyin.bsky.social · 04/12/2024
I'll be at NeurIPS next week starting Tuesday! Please reach out if you want to talk anything multimodal, data curation, synthetic data, and inference optimizations. I'd love to learn more about your research area as well :))
030
Reposted by Haoli Yin
Daniel van Strien @danielvanstrien.bsky.social · 25/11/2024
I'm re-sharing some recent blog posts on using VLMs for synthetic data generation since there are no link penalties here! How to generate a dataset of queries for training and fine-tuning domain-specific ColPali models using a VLM. 🔗 danielvanstrien.xyz/posts/post-w...
danielvanstrien.xyz
Generating a dataset of queries for training and fine-tuning ColPali models on a UFO dataset – Daniel van Strien
Learn how to generate custom ColPali dataset using an open VLM for multimodal retrieval model training and fine-tuning.
0202
Haoli Yin @haoliyin.bsky.social · 25/11/2024
Working on making data curation dirt cheap btw If you're a cracked engineer we'd love to have you :)) DM me if you have any questions! jobs.ashbyhq.com/DatologyAI (also looking for enthusiastic research interns)
jobs.ashbyhq.com
DatologyAI Jobs
DatologyAI Jobs
031
Haoli Yin @haoliyin.bsky.social · 25/11/2024
The text team cooked so much 🧑‍🍳 it might be better than your Thanksgiving meal Check out this super thorough thread on what and how we achieved the best curated text dataset using public data
081
Haoli Yin @haoliyin.bsky.social · 24/11/2024
Was working on some model inference optimization research (speculative decoding) but in the multimodal setting with vision-language models (i.e. conditioned on images) blue = draft model tokens red = target model tokens yellow = bonus target model tokens #dataviz am I doing this right?
251
Haoli Yin @haoliyin.bsky.social · 21/11/2024
now using uv for any new project and trying to migrate existing projects to uv Starting a new project: uv init uv venv --python 3.xx, source .venv/bin/activate uv add (dependencies) or uv pip install -r requirements.txt ❤️ - installing torch in like 10 seconds - uv sync for fast startup
040
Reposted by Haoli Yin
Ethan @ethanrosenthal.com · 14/11/2024
Massive, impressive post on data curation strategies for producing better models with less data and compute. The best part of data curation is that it's a (relatively small) one time cost that gets amortized over all future models. Link to the technical write-up: www.datologyai.com/post/product...
datologyai.com
Technical Deep-Dive: Image-Text Data Curation at the Billion-Sample Scale
Introducing DatologyAI’s state-of-the-art data curation pipeline.
0103
Haoli Yin @haoliyin.bsky.social · 14/11/2024
Web-Scale Data Curation is a frontier challenge - I'm excited to show the progress we've made in just 6 months @datologyai tl;dr: we've pretrained the most data-efficient and best-in-class CLIP models! Read on to see how our product powers multimodal data curation 1/n 🧵
131
Haoli Yin @haoliyin.bsky.social · 11/11/2024
Pouring one out for my bluesky x data gang 🍻 Get ready to see the culmination of the web-scale multimodal data curation work we've been cooking up at DatologyAI! I'm beyond excited to share soon :)) if there's enough interest, might drop something earlier 👀
010
Haoli Yin @haoliyin.bsky.social · 29/10/2024
hello world! I was convinced by @codestar.bsky.social so lets see how much signal is here
160