Sign in

DrivenData

@drivendata.org
51 followers 38 following 70 posts

DrivenData builds AI solutions for social good through data science competitions and our expert in-house project team. Challenges - www.drivendata.org Data Consulting - drivendata.co

PostsRepliesMedia
DrivenData @drivendata.org · 16/09/2026
All three tracks are still open for the Lost in Transcription Challenge! Compete for $20k in prizes, and put your models to the test transcribing bilingual speech across English-Spanish, Spanish-Nahuatl, and Indonesian-Javanese. Join → competitions.mozilladatacollective.com/competitions...
competitions.mozilladatacollective.com
https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/
000
DrivenData @drivendata.org · 15/09/2026
Clinical NLP gets tricky fast, especially with rare concepts. This benchmark tests how well models can turn clinical notes into structured SNOMED CT data using the largest public dataset of labeled clinical notes out there. Take a look: www.drivendata.org/benchmarks/3...
drivendata.org
SNOMED CT Entity Linking Benchmark
A benchmark for linking text in medical notes to entities in SNOMED Clinical Taxonomy.
010
DrivenData @drivendata.org · 11/09/2026
The Spanish-English track of Lost in Transcription is open! Transcribe bilingual speech for a shot at $6,000, or $20,000 if you take all three tracks. Join --> competitions.mozilladatacollective.com/competitions...
competitions.mozilladatacollective.com
https://competitions.mozilladatacollective.com/competitions/2/lost-in-transcription-sp-en/
000
DrivenData @drivendata.org · 09/09/2026
There's one week left to enter the DaT Parkinson's Challenge! Develop models that predict Parkinson's disease from DaT scans and compete for €25,000 in prizes. Join --> www.drivendata.org/competitions...
drivendata.org
DaT Parkinson's Challenge
Help support early and accurate detection of parkinsonian syndromes by developing models that classify dopamine transporter (DaT) scans as normal or abnormal.
000
DrivenData @drivendata.org · 03/09/2026
Submissions are open for the Spanish-Nahuatl track of Lost in Transcription! Put your models to the test transcribing bilingual Spanish-Nahuatl speech. $6,000 in prizes here, $20,000 if you take on all three tracks. Sign up to compete → competitions.mozilladatacollective.com/competitions...
competitions.mozilladatacollective.com
https://competitions.mozilladatacollective.com/competitions/3/lost-in-transcription-sp-nh/
000
DrivenData @drivendata.org · 25/08/2026
Submissions are open for the Indonesian-Javanese track of Lost in Transcription! Put your models to the test transcribing Indonesian-Javanese bilingual speech. Sign up to compete and get notified when the next two language tracks open. Join --> competitions.mozilladatacollective.com/competitions...
competitions.mozilladatacollective.com
https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/
000
DrivenData @drivendata.org · 20/08/2026
One week left to enter Trace the Ace! Help advance how we understand and improve tutoring, with $50K in prizes. Submissions close August 27. platform.k12-ai-infrastructure.org/competitions...
platform.k12-ai-infrastructure.org
Competition - Trace the Ace
Predict the learning gains from a tutoring session measured by quiz performance in this tutoring outcomes prediction challenge
000
DrivenData @drivendata.org · 19/08/2026
DaT imaging helps diagnose Parkinsonian syndromes, but expert interpretation isn't available everywhere. Build computer vision models that can help, with a shot at €25K in prizes. www.drivendata.org/competitions... @healthdatahub-pds.bsky.social
drivendata.org
DaT Parkinson's Challenge
Help support early and accurate detection of parkinsonian syndromes by developing models that classify dopamine transporter (DaT) scans as normal or abnormal.
010
DrivenData @drivendata.org · 10/08/2026
A lot of healthcare data is stuck in free text. The SNOMED benchmark focuses on turning clinical notes into structured concepts. Explore it or try your own model: www.drivendata.org/benchmarks/3...
drivendata.org
SNOMED CT Entity Linking Benchmark
A benchmark for linking text in medical notes to entities in SNOMED Clinical Taxonomy.
000
DrivenData @drivendata.org · 06/08/2026
New challenge! 🔊 Lost in Transcription asks you to build ASR models that transcribe bilingual speech across three language pairs. Sign up now to explore the data, get notified when submissions open, and compete for $20k. Join → competitions.mozilladatacollective.com/competitions...
competitions.mozilladatacollective.com
000
DrivenData @drivendata.org · 05/08/2026
Trace the Ace is still live!  Build models that predict student learning from tutoring transcripts, and compete for a share of $50K. About 3 weeks left — submissions close August 27.  platform.k12-ai-infrastructure.org/competitions...
platform.k12-ai-infrastructure.org
Competition - Trace the Ace
Predict the learning gains from a tutoring session measured by quiz performance in this tutoring outcomes prediction challenge
010
DrivenData @drivendata.org · 29/07/2026
"Beautiful" pipelines don't happen by accident. They take intentional structure, cleanup, and iteration over time. Check out our blog post on common pipeline challenges and how to fix them: drivendata.co/blog/pipelin...
drivendata.co
5 Challenges of Creating Beautiful Data Pipelines
A look into the hidden complexity of data pipelines, and some suggestions to improve the process.
010
DrivenData @drivendata.org · 22/07/2026
We've got a new challenge! Help build models that classify DaT SPECT scans to support Parkinson's diagnosis. €25K in prizes, built on an open dataset from 10 French hospitals, with winning models going open-source. Join → www.drivendata.org/competitions... @healthdatahub-pds.bsky.social
drivendata.org
DaT Parkinson's Challenge
Help support early and accurate detection of parkinsonian syndromes by developing models that classify dopamine transporter (DaT) scans as normal or abnormal.
021
DrivenData @drivendata.org · 13/07/2026
Structuring clinical notes is one of healthcare AI's toughest problems. See how your model compares in the SNOMED Clinical Entity Linking Benchmark: www.drivendata.org/benchmarks/3...
drivendata.org
SNOMED CT Entity Linking Benchmark
A benchmark for linking text in medical notes to entities in SNOMED Clinical Taxonomy.
000
DrivenData @drivendata.org · 09/07/2026
Calling all solvers! Join the Trace the Ace challenge and help build models  that predict student learning based on tutoring session transcripts, and compete for a share of the $50k prize pool. Submissions are open now through August 27. platform.k12-ai-infrastructure.org/competitions...
platform.k12-ai-infrastructure.org
Competition - Trace the Ace
Predict the learning gains from a tutoring session measured by quiz performance in this tutoring outcomes prediction challenge
010
DrivenData @drivendata.org · 07/07/2026
How do you study conversation at scale? DrivenData and BetterUp Labs built the infrastructure to do it, resulting in the CANDOR Corpus - a 1,000+ hour dataset of real-world American English conversations. Read more: drivendata.co/case-studies/building-research-infrastructure-for-conversational-ai
drivendata.co
Building research infrastructure for conversational AI
DrivenData partnered with BetterUp Labs to design and build the data infrastructure, machine learning pipelines, and research tools behind the CANDOR Corpus—enabling cutting-edge insights into human…
000
DrivenData @drivendata.org · 01/07/2026
Welcome to the K-12 AI Infrastructure Program’s platform - the hub for researchers and developers to access curated K-12 datasets, models, and benchmarks aligned with scientific principles of teaching and learning, and free for everyone to use. Check it out: platform.k12-ai-infrastructure.org
000
DrivenData @drivendata.org · 26/06/2026
A lot of “data usability” comes down to pretty unglamorous work: cleaning, standardizing, reconciling edge cases - but that’s the part that makes everything else possible. We wrote our approach here: drivendata.co/blog/last-mi...
drivendata.co
Solving the last-mile public data problem
Using "baked" data to transform public data repositories into analysis-ready resources
031
DrivenData @drivendata.org · 22/06/2026
Results are in for the On Top of Pasketti: Children's Speech Recognition Challenge! 828+ participants tackled kids' voices - one of the hardest problems in speech recognition - and winners cut error rates in half. Meet the winners: drivendata.co/blog/on-top-...
drivendata.co
Meet the winners of the On Top of Pasketti: Children's Speech Recognition Challenge
Learn how competition winners, working with one of the largest labeled children's speech datasets assembled, cut transcription error rates in half.
010
DrivenData @drivendata.org · 17/06/2026
The $26M K-12 AI Infrastructure Program is launching our new platform! We're issuing grants and hosting prize challenges to create open benchmarks, datasets and models that advance the impact and reach of AI in K-12 education. Explore the platform: platform.k12-ai-infrastructure.org
000
DrivenData @drivendata.org · 08/06/2026
We added a benchmark focused on clinical entity linking. It uses a large set of de-identified notes annotated by medical professionals, which makes it a pretty realistic testbed. Check out how different approaches compare: www.drivendata.org/benchmarks/3...
drivendata.org
SNOMED CT Entity Linking Benchmark
A benchmark for linking text in medical notes to entities in SNOMED Clinical Taxonomy.
000
DrivenData @drivendata.org · 05/05/2026
How well can your model understand a doctor’s note? Our benchmark tests how accurately models can identify and link medical concepts using SNOMED CT on the largest public dataset of its kind. Get started: www.drivendata.org/benchmarks/3...
drivendata.org
SNOMED CT Entity Linking Benchmark
A benchmark for linking text in medical notes to entities in SNOMED Clinical Taxonomy.
100
Reposted by DrivenData
Peter Bull @peter.drivendata.org · 01/04/2026
Excited to be speaking at Good Tech Summit in DC April 7 www.goodtechtogether.org/summit We’ll share a program focused on K-12 education and talk about investing in the foundations of AI: data, models, and benchmarks. We'll explore how these shape AI development in a field. Join us!
021
DrivenData @drivendata.org · 30/03/2026
The performance gap in children’s ASR is real — and solvable! With 1 week until the April 6 deadline, we’re inviting the global ML community to help close it. Compete for $120K and contribute to better speech technology for kids. kidsasr.drivendata.org
kidsasr.drivendata.org
On Top of Pasketti: Children’s Speech Recognition Challenge
Develop cutting-edge ASR algorithms specifically for children's speech to advance early education assessments and teaching tools.
000
DrivenData @drivendata.org · 16/03/2026
Children’s speech remains one of ASR’s toughest challenges — and the leaderboard is moving! 3 weeks left to compete for $120K in the On Top of Pasketti Challenge. Submit your model: kidsasr.drivendata.org
kidsasr.drivendata.org
On Top of Pasketti: Children’s Speech Recognition Challenge
Develop cutting-edge ASR algorithms specifically for children's speech to advance early education assessments and teaching tools.
010
DrivenData @drivendata.org · 06/03/2026
10 years ago, "data science for social good" was just an idea. Today, it's a global movement. Our 10-Year Impact Report reflects on a decade of responsible, real-world AI. Take a look back with us: s3.amazonaws.com/drivendata-p...
s3.amazonaws.com
000
DrivenData @drivendata.org · 03/03/2026
Competing in the On Top of Pasketti Word Track? We've published a reference tutorial walking through how to build a model for children's speech recognition — covering data exploration, model training, and submission packaging. Get started: drivendata.co/blog/child-a...
drivendata.co
Improving Automatic Speech Recognition for Kids - On Top of Pasketti Word-Track Benchmark
Learn how to train a model to transcribe child speech for the On Top of Pasketti Challenge (Word Track)
000
DrivenData @drivendata.org · 03/03/2026
From Bogotá to Almaty, our community continues to impress! IGCPHARMA built prize-winning early dementia prediction models in PREPARE. Kirill Brodt has 23 competitions, 8 top-10s, and 6 top-3 finishes. Two new Community Spotlights: drivendata.co/blog/communi... drivendata.co/blog/communi...
000
DrivenData @drivendata.org · 02/03/2026
Automatic speech recognition struggles with children’s speech. That gap matters! The $120K Children’s Speech Recognition Challenge is driving progress toward models that truly understand kids. Join breakthrough! Submit your solution by April 6th. kidsasr.drivendata.org
kidsasr.drivendata.org
On Top of Pasketti: Children’s Speech Recognition Challenge
Develop cutting-edge ASR algorithms specifically for children's speech to advance early education assessments and teaching tools.
000
DrivenData @drivendata.org · 26/02/2026
A machine learning competition for NASA sparked something bigger. CyFi (cyanobacteria finder) is now an open-source tool using Sentinel-2 data to monitor harmful algal blooms worldwide, with local validation underway. How we got here: drivendata.co/blog/cyfi-sm...
drivendata.co
Bringing small water bodies into view: Sentinel-2 satellite monitoring of harmful algal blooms (HABs)
CyFi enhances modern HAB monitoring programs by extending their reach and informing field-based components.
010
DrivenData @drivendata.org · 25/02/2026
What happens when AI agents enter ML competitions? Spoiler: Humans 1, Agents 0.25–0.93. The top of the leaderboard — where “good” becomes “great” — still looks very human. How might that change? See the results: drivendata.co/blog/ai-agen...
drivendata.co
AI Agents in Data Science Competitions: Lessons from the Leaderboard
How good are AI agents at data science? Here's what we've learned from initial experiments about what works, what doesn't, and what the future might hold.
010
DrivenData @drivendata.org · 25/02/2026
We’re launching our first benchmark competition. The SNOMED CT Entity Linking Benchmark evaluates how well models structure clinical notes using the SNOMED CT nomenclature for a large, de-identified dataset of annotated records. Help set the baseline: www.drivendata.org/benchmarks/3...
010
DrivenData @drivendata.org · 24/02/2026
Building data pipelines is deceptively complex; Cold starts, unstable inputs, shifting requirements, and delivery trade-offs create friction at every step. To ease the pain, we examined five recurring challenges and suggest practical improvements: drivendata.co/blog/pipelin...
drivendata.co
000
DrivenData @drivendata.org · 23/02/2026
ML competitions moved fast in 2025 - from 512-GPU training runs to benchmark-style challenges. The new State of Machine Learning Competitions report explores what winning teams used. We're grateful to be part of such a dynamic ML community. Read: mlcontests.com/state-of-mac...
mlcontests.com
The State of Machine Learning Competitions
More than 390 machine learning competitions took place in 2025, across 30+ platforms, with a total prize pool of over $16m. These competitions included multi-million-dollar government-funded…
000
DrivenData @drivendata.org · 19/02/2026
The gap between public data and usable data is the “last-mile data problem.” We’ve all seen it: confusing CSVs, messy schemas, opaque data dictionaries. We’re testing a “baked data” approach to solve it, with promising results. See our recipe: drivendata.co/blog/last-mi...
000
DrivenData @drivendata.org · 18/02/2026
We gave AI agents 24 hours on the leaderboard. Claude Code (Opus 4.5) and Codex (GPT 5.2) produced dozens of submissions. Some hit the top 20. Others overfit or plateaued. We identified 9 obstacles and 6 open questions. Do they match your experience? drivendata.co/blog/ai-agen...
drivendata.co
AI Agents in Data Science Competitions: Lessons from the Leaderboard
How good are AI agents at data science? Here's what we've learned from initial experiments about what works, what doesn't, and what the future might hold.
000
DrivenData @drivendata.org · 16/02/2026
Voice-based tools have the potential to support learning, accessibility, and early literacy, but only if they work for children. In the $120k On Top of Pasketti Children’s Speech Recognition Challenge, solvers will build ASR systems that understand kids. kidsasr.drivendata.org
kidsasr.drivendata.org
On Top of Pasketti: Children’s Speech Recognition Challenge
Develop cutting-edge ASR algorithms specifically for children's speech to advance early education assessments and teaching tools.
000
DrivenData @drivendata.org · 10/02/2026
Hang out with us at #SciPy2026 this summer! Senior Data Scientist Katie Wetstone is co-chairing the Environmental, Earth, and Climate Sciences track, which you can submit to by February 25.
bsky.app
SciPy Conference 2026 (@scipyconf.bsky.social)
📣 Call for Proposals is OPEN for #SciPy2026! Have a talk, tutorial, or poster idea you’re excited to share with the scientific Python community? Now’s the time! 🚀 🗓 Submit by: February 25, 2026 🔗…
000
DrivenData @drivendata.org · 06/02/2026
We're excited to be a part of the K-12 AI Infrastructure Program, advancing open datasets, models, and benchmarks for AI in teaching & learning. The first RFP is now open ($50K–$250K) - apply now! k12-ai-infrastructure.org/rfp-due-marc...
000
DrivenData @drivendata.org · 04/02/2026
Kids learn through voice, but today's ASR tech can barely understand them. In a new data science challenge, solvers will develop models that work with children's unique speech patterns and compete for a share of the $120k prize pool. kidsasr.drivendata.org
kidsasr.drivendata.org
On Top of Pasketti: Children’s Speech Recognition Challenge
Develop cutting-edge ASR algorithms specifically for children's speech to advance early education assessments and teaching tools.
000
Reposted by DrivenData
Center for Educational Data Science & Innovation @edsi-umd.bsky.social · 04/02/2026
🚨 New opportunity: Help build open-source speech recognition AI 🎙️📚 @drivendata.org is hosting a data science competition to advance automatic speech recognition (ASR) for early education. Two tracks, real impact, and $120K in prizes. Learn more & compete: kidsasr.drivendata.org
kidsasr.drivendata.org
On Top of Pasketti: Children’s Speech Recognition Challenge
Develop cutting-edge ASR algorithms specifically for children's speech to advance early education assessments and teaching tools.
011
DrivenData @drivendata.org · 29/01/2026
DrivenData's Katie Wetstone will be co-chairing the climate sciences track at @scipyconf.bsky.social, where you can share YOUR work in environmental data science!  Submit a #SciPy2026 talk, tutorial, or poster by February 25. See you there!
bsky.app
SciPy Conference 2026 (@scipyconf.bsky.social)
📣 Call for Proposals is OPEN for #SciPy2026! Have a talk, tutorial, or poster idea you’re excited to share with the scientific Python community? Now’s the time! 🚀 🗓 Submit by: February 25, 2026 🔗…
000
DrivenData @drivendata.org · 29/12/2025
The $10k prize pool Poverty Prediction Challenge sponsored by The World Bank tackles a major challenge in development research: How do you estimate current poverty rates without recent household expenditure data? Submissions open through midnight UTC February 4, 2026.
drivendata.org
Poverty Prediction Challenge
Estimate individual and aggregate household consumption from limited survey data.
000
DrivenData @drivendata.org · 10/12/2025
In our newest machine learning challenge, solvers will use survey data and help uncover imputation methods for monitoring poverty trends. The top teams will take home a share of the $10,000 prize, provided by The World Bank. Learn more and submit predictions here: www.drivendata.org/competitions...
drivendata.org
Poverty Prediction Challenge
Estimate individual and aggregate household consumption from limited survey data.
020
DrivenData @drivendata.org · 15/09/2025
Throwback to our "Hateful Memes" challenge where teams detected harmful content combining text and images. Critical work for online safety! 🛡️💻 #ContentModeration #AI drivendata.org/competitions/64/
020
DrivenData @drivendata.org · 11/09/2025
Hot topic in #DataScience: Multi-modal learning is bridging text, images, and audio. Our blog explores practical applications beyond the hype! 🎭🔗 #MultiModal #AI drivendata.co/blog.html
drivendata.co
DrivenData Labs
DrivenData helps mission-driven organizations harness their data to work smarter and offer more impactful services using data science, machine learning, and AI.
020
DrivenData @drivendata.org · 08/09/2025
Our latest blog dives into "Causal Inference for Data Scientists" - moving beyond correlation to understand what actually drives outcomes! 🔗💡 #CausalInference #DataScience drivendata.co/blog.html
drivendata.co
DrivenData Labs
DrivenData helps mission-driven organizations harness their data to work smarter and offer more impactful services using data science, machine learning, and AI.
020
DrivenData @drivendata.org · 04/09/2025
Ever wonder how AI helps track endangered species? Our "Pri-matrix Factorization" competition used camera trap data to identify primates in the wild! 🐒📷 #WildlifeConservation #ComputerVision drivendata.org/competitions/49/
drivendata.org
Pri-matrix Factorization
Data scientists from more than 90 countries around the world drew on 300,000 video clips in a competition to build the best machine learning models for identifying wildlife from camera trap footage. …
010
DrivenData @drivendata.org · 21/08/2025
Community Spotlight: Kirill Brodt The Community Spotlight features fantastic members from our DrivenData community. Kirill Brodt, a researcher in computer graphics at the University of Montreal, talks animation, pose estimation, and data science challenges.
drivendata.co
Community Spotlight: Kirill Brodt
The Community Spotlight features fantastic members from our DrivenData community. Kirill Brodt, a researcher in computer graphics at the University of Montreal, talks animation, pose estimation, and…
000
DrivenData @drivendata.org · 30/06/2025
Want to build ML systems that actually work in production? Download our free ebook "The 10 Rules of Reliable Data Science" and learn the essentials from our years of real-world experience. Game-changing insights await! 📊🔬 #DataScience #MachineLearning
drivendata.co
DrivenData Labs
DrivenData helps mission-driven organizations harness their data to work smarter and offer more impactful services using data science, machine learning, and AI.
030