Sign in

Data Frosch

@datafrosch.bsky.social
15 followers 39 following 20 posts

Come join us at The Pond! datafrosch.fun/pond.html

PostsRepliesMedia
Reposted by Data Frosch
The European Correspondent @eurcorrespond.bsky.social · 02/10/2026
While EU citizens can move freely and settle anywhere within the Union, they cannot vote in national elections where they live, unless they are also citizens of that country.
164
Reposted by Data Frosch
Ada Homolova @adahomolova.bsky.social · 12h
My first in-person analogue workshop on how ho get the best results from working with language models for the local knowledge sharing festival 🥰 As someone who works mostly digitally, the process of creating objects and experience in the real world was very enjoyable.
1. Don't trust the robot
2. Context is key
3. Keep it short
4. Keep it to the point
5. You might be wrongCircle of people picture with polaroid cameraRobot face in the makingA person holding an adapted robot emoji in front of her face
001
Data Frosch @datafrosch.bsky.social · 30/09/2026
So before you attach everything, try: 1. Curate, don't dump: include only what's relevant to the task at hand. 2. Put the important stuff at the start or end of your prompt, not the middle. 3. Ask for what you need first: sometimes a targeted question beats a document dump. Less is really more 🐸
000
Data Frosch @datafrosch.bsky.social · 30/09/2026
More context also means more ways for the model to get distracted: contradictory details, outdated versions of the same document, tangents you explored earlier. And a practical bonus: extra context costs you money. Every token in the window gets processed and billed, whether it helps or not.
100
Data Frosch @datafrosch.bsky.social · 30/09/2026
There's even a name for this: "lost in the middle". Models tend to focus on the beginning and the end of the context, and things in the middle get overlooked. So yes, the crucial fact on page 47 might be exactly the one the model skips.
100
Data Frosch @datafrosch.bsky.social · 30/09/2026
With AI, it often works the other way. Models have a limited attention budget. The more you stuff into the context, the harder it gets for the model to find what actually matters for your question. Your key detail can end up buried between pages of stuff that's irrelevant to the task.
110
Data Frosch @datafrosch.bsky.social · 30/09/2026
Things nobody tells you about working with AI, Part 3: More context ≠ better answers Have you ever dumped a 100-page PDF into a chat, thinking: now it has everything, the answers will be perfect? It's a reasonable assumption. More information should mean better answers, right?
110
Data Frosch @datafrosch.bsky.social · 29/09/2026
And sometimes the dirt IS the story: Jonathan found the Ministry of Sound giving tickets to MPs. Read more: datafrosch.fun/blog/how-clu... #openrefine #datacleaning
youtu.be
🧹 Pondcast #8: On data cleaning (with Jonathan Stoneman and Ada Homolova)
How do you clean a dataset without destroying it, and without spending a week on it? In this Pond session, Jonathan Stoneman (who has been cleaning data in OpenRefine since it was called Google…
000
Data Frosch @datafrosch.bsky.social · 29/09/2026
The fix, in OpenRefine: - clean a copy of the column, never the original - lowercase + trim whitespace first — remove the fake differences - then cluster, strictest first: fingerprint → n-gram → phonetic - nearest neighbor (Levenshtein) if you need looser matches
100
Data Frosch @datafrosch.bsky.social · 29/09/2026
British donation data has 21,000 "different" donor names. Many aren't different at all: Lord Harris of Peckham Lord Philip Harris of Peckham Lord Phillip Harris of Peckham (two L's) Same guy. Your pivot is already wrong.
100
Data Frosch @datafrosch.bsky.social · 26/09/2026
As the coding itself can be actually done by the model, I wonder if focusing on the code is what we should be teaching? Why not instead highlight the concepts, put them into context, and teach how to then prompt the model to make what I need?
010
Data Frosch @datafrosch.bsky.social · 26/09/2026
Thank you and thank you for sharing 💚
010
Data Frosch @datafrosch.bsky.social · 25/09/2026
@cghlewis.bsky.social Hi Crystal! I really liked your 3 part post on data cleaning and would like to feature it in our newsletter. Do you perhaps have any other tips for data cleaning resources? Cheers! ~ Ada
110
Data Frosch @datafrosch.bsky.social · 24/09/2026
We are looking for data nerds to feature in our Pondcast series! Would you like to share cool tool or workflow with us? Or is there someone you would like to see share theirs? Get in touch & share! 💚 🐸 #ddj #data #datajournalism youtu.be/kE-uP7o6SQI?...
youtu.be
🧹 Pondcast #8: On data cleaning (with Jonathan Stoneman and Ada Homolova)
YouTube video by Data Frosch
000
Data Frosch @datafrosch.bsky.social · 22/09/2026
Come learn about the AI plugin for Open Refine from no one else than @herve.checkfirst.network ! Arrive with curiosity, leave with practical, checkable AI-assisted workflows for your next data cleaning task 🦾💪 Grab the link to the meeting here: datafrosch.fun/pond.html The session is free 💚
datafrosch.fun
The Pond | Data Frosch
A community for and by nerds in the newsroom — join the Discord, subscribe to the newsletter, and read past editions.
011
Data Frosch @datafrosch.bsky.social · 17/09/2026
1. Keep your chat clean 🧹: Keep the conversation short 2. Separate brainstorming from evaluation: First ask the model to develop your idea, then start a new chat and ask it explicitly to critique the idea. 🧠👀 3. Don't turn on memory in your chat app 🙅‍♀️: make sure that every new chat is really new
000
Data Frosch @datafrosch.bsky.social · 17/09/2026
If you’re using AI for research, having the model absorb and reinforce your biases can seriously work against you. This is also one of the big reason why not to use AI in situations where you need to be challenged rather than validated: therapist, doctor, or a boyfriend for that matter.
100
Data Frosch @datafrosch.bsky.social · 17/09/2026
In practice, this can create an echo chamber: instead of helping you examine an idea from multiple angles, the assistant starts mirroring your assumptions and misconceptions back to you. So the longer your conversation, the more chance that the model will agree with anything you say.
100
Data Frosch @datafrosch.bsky.social · 17/09/2026
The models were trained to respond like this, as people apparently really like being told they’re right. This tendency can become stronger as a conversation gets longer. If a stored memory contains a bias or mistaken belief, the model may end up reinforcing it instead of questioning it.
100
Data Frosch @datafrosch.bsky.social · 17/09/2026
Have you noticed an AI model telling you that you’re absolutely right, your ideas are amazing, and your approach is unique? This is called "sycophancy": the tendency of AI models to agree with, flatter, or validate the user rather than challenge their assumptions.
110
Data Frosch @datafrosch.bsky.social · 16/09/2026
Same dirty dataset, two approaches: 🧹 Jonathan takes 2.5 hours to clean 50,000 rows with OpenRefien 🤖 Ada takes 1 prompt and 2 minutes and gets a quick report on the problematic parts We did and discuss both, live: youtu.be/kE-uP7o6SQI
youtu.be
🧹 Pondcast #8: On data cleaning (with Jonathan Stoneman and Ada Homolova)
YouTube video by Data Frosch
010
Data Frosch @datafrosch.bsky.social · 15/09/2026
What if you opened a map of what the world is searching for instead of a news website? Learn all about text embeddings with us! www.youtube.com/watch?v=_teY...
youtube.com
🟠 Pondcast #5: Text embeddings and how to use them (with Ada Homolova and Johan Schuijt)
YouTube video by Data Frosch
020