Data Frosch @datafrosch.bsky.social · 30/09/2026Things nobody tells you about working with AI, Part 3: More context ≠ better answers Have you ever dumped a 100-page PDF into a chat, thinking: now it has everything, the answers will be perfect? It's a reasonable assumption. More information should mean better answers, right? 110
Data Frosch @datafrosch.bsky.social · 29/09/2026British donation data has 21,000 "different" donor names. Many aren't different at all: Lord Harris of Peckham Lord Philip Harris of Peckham Lord Phillip Harris of Peckham (two L's) Same guy. Your pivot is already wrong. 100
Data Frosch @datafrosch.bsky.social · 25/09/2026@cghlewis.bsky.social Hi Crystal! I really liked your 3 part post on data cleaning and would like to feature it in our newsletter. Do you perhaps have any other tips for data cleaning resources? Cheers! ~ Ada 110
Data Frosch @datafrosch.bsky.social · 24/09/2026We are looking for data nerds to feature in our Pondcast series! Would you like to share cool tool or workflow with us? Or is there someone you would like to see share theirs? Get in touch & share! 💚 🐸 #ddj #data #datajournalism youtu.be/kE-uP7o6SQI?...youtu.be🧹 Pondcast #8: On data cleaning (with Jonathan Stoneman and Ada Homolova)YouTube video by Data Frosch 000
Data Frosch @datafrosch.bsky.social · 22/09/2026Come learn about the AI plugin for Open Refine from no one else than @herve.checkfirst.network ! Arrive with curiosity, leave with practical, checkable AI-assisted workflows for your next data cleaning task 🦾💪 Grab the link to the meeting here: datafrosch.fun/pond.html The session is free 💚datafrosch.funThe Pond | Data FroschA community for and by nerds in the newsroom — join the Discord, subscribe to the newsletter, and read past editions. 011
Data Frosch @datafrosch.bsky.social · 17/09/2026Have you noticed an AI model telling you that you’re absolutely right, your ideas are amazing, and your approach is unique? This is called "sycophancy": the tendency of AI models to agree with, flatter, or validate the user rather than challenge their assumptions. 110
Data Frosch @datafrosch.bsky.social · 16/09/2026Same dirty dataset, two approaches: 🧹 Jonathan takes 2.5 hours to clean 50,000 rows with OpenRefien 🤖 Ada takes 1 prompt and 2 minutes and gets a quick report on the problematic parts We did and discuss both, live: youtu.be/kE-uP7o6SQIyoutu.be🧹 Pondcast #8: On data cleaning (with Jonathan Stoneman and Ada Homolova)YouTube video by Data Frosch 010
Data Frosch @datafrosch.bsky.social · 15/09/2026What if you opened a map of what the world is searching for instead of a news website? Learn all about text embeddings with us! www.youtube.com/watch?v=_teY...youtube.com🟠 Pondcast #5: Text embeddings and how to use them (with Ada Homolova and Johan Schuijt)YouTube video by Data Frosch 020