Sign in

Towards Data Science

@towardsdatascience.com
1.4K followers 103 following 3.8K posts

The world's leading publication for data science and artificial intelligence professionals. Website 🌐 towardsdatascience.com Submit an Article ✍️ contributor.insightmediagroup.io Subscribe to our Newsletter 📩 bit.ly/TDS-Newsletter

PostsRepliesMedia
Towards Data Science @towardsdatascience.com · 41m
In a new deep dive, Tigran Hayrapetyan introduces Guided Merge Sort, an optimized sorting approach that chooses the best options among ordinary and multi-way merge-sort algorithms.
towardsdatascience.com
Guided Merge Sort : An Optimized Sorting that Picks the Best from Ordinary and Multi-Way Merge Sort Algorithms
A solution where the "goto" operator becomes irreplaceable.
010
Towards Data Science @towardsdatascience.com · 3h
Follow along @taupirho.bsky.social's latest tutorial to learn how to benchmark the impact of fewer, larger files across three SQL workloads.
towardsdatascience.com
I Compacted 1,000 Apache Iceberg Files Into 6. Here’s What Happened to Query Performance.
Benchmarking the impact of fewer, larger files across three SQL workloads
000
Towards Data Science @towardsdatascience.com · 7h
In an insightful analysis, Yu Dong revisits a timely topic: the shape and extent of AI's impact on data science roles.
towardsdatascience.com
AI Made Data Scientists Faster. Now It’s Expanding the Job.
The real shift is bigger than productivity: AI is reshaping ownership, judgment, and the career path of data scientists.
000
Towards Data Science @towardsdatascience.com · 11h
"This has never sat right with me: a bounded yes/no or routing decision was often passed through the same autoregressive decoder used to generate a paragraph." Lambert Leong digs into the architectural constraints that are steering the field towards "decision" models like Jev.
towardsdatascience.com
When All You Have Are Decoders, Every Decision Looks Like Generation
Not every decision needs a decoder, generation Is not always a decision
000
Towards Data Science @towardsdatascience.com · 16h
Getting started in designing safe and trustworthy agentic systems? Thuwarakesh Murallie presents a comprehensive guide to the architectural guardrails you need to be fluent in.
towardsdatascience.com
How to Design Architectural Guardrails Around AI Agents
Agent design patterns every data engineer must know
000
Towards Data Science @towardsdatascience.com · 17h
We're thrilled to welcome Vasileios Vonikakis to TDS! You should explore his debut article, which unpacks the process of building fair evaluation sets when dealing with imbalanced data.
towardsdatascience.com
Building Fair Evaluation Sets Is a Combinatorial Problem
Here’s how to solve it exactly.
000
Towards Data Science @towardsdatascience.com · 22h
What if you could transform an open-source LLM into your own, Jev-like single-pass text classifier? Anubhab Banerjee outlines an approach that can accomplish this result.
towardsdatascience.com
How to Make Your Own JEV Model from an Open LLM
Turn a small open-source Qwen LLM into a fast, single-pass text classifier by swapping its language-modeling head
020
Towards Data Science @towardsdatascience.com · 29/09/2026
"At some point, with no new data and nothing visibly different about the setup, the model's performance on the unseen questions jumped from barely better than random to nearly perfect." Utkarsh Mangal unpacks the phenomenon of grokking in machine learning.
towardsdatascience.com
The AI That Learned to Understand Long After It Stopped Trying
A small, strange discovery in machine learning called grokking
010
Towards Data Science @towardsdatascience.com · 29/09/2026
Learn how you can detect hidden shifts in feature relationships to avoid data drift — Benjamin Nweke leverages adversarial validation and scikit-learn in his latest tutorial.
towardsdatascience.com
How to Catch Data Drift When Every Feature Looks Normal
Detect hidden shifts in feature relationships with adversarial validation and scikit-learn
000
Towards Data Science @towardsdatascience.com · 29/09/2026
Can you leverage the power of Jev in the context of scalable knowledge graphs? Partha Sarkar explains what that looks like in practice in his excellent hands-on guide.
towardsdatascience.com
GraphRAG with TypeSafe Jev: A System One Approach to Scalable Knowledge Graphs
How calibrated decision models can handle high-frequency graph decisions while LLMs remain focused on reasoning, synthesis, and open-ended generation.
000
Towards Data Science @towardsdatascience.com · 29/09/2026
"Cursor, Claude Code and Copilot all sit in front of the same gap, whatever each of them indexes, because the information is not in any repository they can see." Yonatan Sason zooms in on a structural issue affecting AI-assisted software development.
towardsdatascience.com
Good Architecture Deletes the Signals Your Agent Depends On
Every boundary you draw removes a signal your tooling was relying on. That is a structure problem, not a search problem.
101
Towards Data Science @towardsdatascience.com · 29/09/2026
AI-generated content has entered training data — and at a massive rate. Abdullahi Dattijo presents his findings from testing three approaches to spot it.
towardsdatascience.com
AI Slop Is Already in Your Training Dataset. I Tested Three Ways to Spot It.
My AI detectors flagged many genuine reviews, and filtering them made the sentiment model less accurate.
000
Towards Data Science @towardsdatascience.com · 29/09/2026
How do paragraphs "work" as a meaningful unit of content in the context of LLMs? Shuyang Xiang offers an accessible explainer on an often-overlooked topic.
towardsdatascience.com
Your LLM Has a Curved Space of Paragraphs
Inside a transformer, token index is a coordinate. Paragraph structure is what turns it into a metric.
011
Towards Data Science @towardsdatascience.com · 28/09/2026
Faster code generation changes where engineering teams spend their time. Eivind Kjosbakken explains how Claude Code can be integrated into more efficient deployment workflows.
towardsdatascience.com
How to Effectively Deploy Code With Claude Code | Towards Data Science
Learn how to optimize your CI/CD pipeline for coding agents
000
Towards Data Science @towardsdatascience.com · 28/09/2026
The human checkpoint that once protected enterprise data is disappearing. Shafeeq Ur Rahaman examines how data warehouses need to evolve when AI agents become active decision-makers.
towardsdatascience.com
Building an Agent-Ready Data Warehouse: What Traditional Architectures Do Wrong | Towards Data Science
Giving an AI agent access to a data warehouse doesn't automatically make it agent-ready. The real challenge lies in teaching the agent what the data means and when it's reliable enough to use.
000
Towards Data Science @towardsdatascience.com · 28/09/2026
Structured output is a key piece of building dependable LLM applications. Shuai Guo explores how to implement it with local models while keeping sensitive data on your own infrastructure.
towardsdatascience.com
How to Implement Structured Output with Local LLMs | Towards Data Science
Why use it? How to implement it? What can we do when it fails?
000
Towards Data Science @towardsdatascience.com · 28/09/2026
Running an LLM locally solves the privacy problem, but not the integration problem. Shuai Guo breaks down how structured outputs make local models easier to connect to real-world workflows.
towardsdatascience.com
How to Implement Structured Output with Local LLMs | Towards Data Science
Why use it? How to implement it? What can we do when it fails?
000
Towards Data Science @towardsdatascience.com · 28/09/2026
When AI writes the code, deployment becomes the next engineering bottleneck. Eivind Kjosbakken explores how to optimize CI/CD workflows for coding agents.
towardsdatascience.com
How to Effectively Deploy Code With Claude Code | Towards Data Science
Learn how to optimize your CI/CD pipeline for coding agents
110
Towards Data Science @towardsdatascience.com · 28/09/2026
Private AI workflows can now work with more than text. Shuai Guo explores how to build multimodal applications with a local LLM using Gemma 4 and Ollama.
towardsdatascience.com
Building Multimodal Workflows with a Local LLM | Towards Data Science
Image inputs and structured outputs with Gemma 4 and Ollama
011
Towards Data Science @towardsdatascience.com · 28/09/2026
Choosing between LangChain and LangGraph depends on the complexity of the workflow you need to build. Soner Yıldırım explains where each framework fits and when one makes more sense than the other.
towardsdatascience.com
LangChain vs LangGraph: 4 Key Differences and When to Use Each | Towards Data Science
A practical guide to choose the proper tool for your agentic workflows and systems
010
Towards Data Science @towardsdatascience.com · 27/09/2026
Giving an AI agent access to your warehouse is only the beginning. Shafeeq Ur Rahaman explores why data meaning, reliability, and context are essential for building agent-ready architectures.
towardsdatascience.com
Building an Agent-Ready Data Warehouse: What Traditional Architectures Do Wrong | Towards Data Science
Giving an AI agent access to a data warehouse doesn't automatically make it agent-ready. The real challenge lies in teaching the agent what the data means and when it's reliable enough to use.
010
Towards Data Science @towardsdatascience.com · 27/09/2026
LangChain and LangGraph solve different problems in modern agentic workflows. Soner Yıldırım breaks down four key differences to help you choose the right tool for your system.
towardsdatascience.com
LangChain vs LangGraph: 4 Key Differences and When to Use Each | Towards Data Science
A practical guide to choose the proper tool for your agentic workflows and systems
000
Towards Data Science @towardsdatascience.com · 27/09/2026
Getting noticed in a competitive ML job market starts with presenting your experience effectively. Egor Howell explains the process used with Claude to craft a resume that helped secure a $200k+ offer.
towardsdatascience.com
How Claude Helped Me Build My $200k+ ML Resume | Towards Data Science
How to use Claude to craft an outstanding resume that lands offers
000
Towards Data Science @towardsdatascience.com · 27/09/2026
Local AI does not need to be complicated or expensive. Mauro Di Pietro walks through creating a free CLI agent with Python and Ollama from the ground up.
towardsdatascience.com
How to Build CLI Agents with Python & Ollama | Towards Data Science
Create a local CLI Agent from scratch completely for free
000
Towards Data Science @towardsdatascience.com · 27/09/2026
Data pipelines should be easy to trust, debug, and maintain as they scale. @taupirho.bsky.social breaks down how the Medallion architecture can make those goals more achievable.
towardsdatascience.com
The Medallion Data Architecture: An Introduction | Towards Data Science
A practical guide to Bronze, Silver and Gold, with a working Python and DuckDB example
110
Towards Data Science @towardsdatascience.com · 27/09/2026
The real potential of coding agents may extend well beyond the code editor. Eivind Kjosbakken explains how to use them for practical tasks that have nothing to do with programming.
towardsdatascience.com
How to Apply Coding Agents to Non-Programming Tasks | Towards Data Science
Perform non-programming tasks with coding agents
032
Towards Data Science @towardsdatascience.com · 27/09/2026
AI agents become more valuable when they can handle the messy details of real customer interactions. Soner Yıldırım shows how a LangGraph agent can collect information, manage state, and streamline a booking workflow.
towardsdatascience.com
I Replaced a 15-Minute Booking Process with a LangGraph AI Agent | Towards Data Science
A step-by-step guide to building, running, and monitoring a stateful customer support agent using Python, LangGraph, and Langfuse.
231
Towards Data Science @towardsdatascience.com · 26/09/2026
A strong technical resume can make a major difference in landing competitive ML roles. Egor Howell shares how Claude helped shape a resume that led to a $200k+ Machine Learning Engineer offer.
towardsdatascience.com
How Claude Helped Me Build My $200k+ ML Resume | Towards Data Science
How to use Claude to craft an outstanding resume that lands offers
010
Towards Data Science @towardsdatascience.com · 26/09/2026
Running an AI agent directly from your terminal changes how you interact with local models. Mauro Di Pietro shows how to build a CLI agent from scratch with Python and Ollama.
towardsdatascience.com
How to Build CLI Agents with Python & Ollama | Towards Data Science
Create a local CLI Agent from scratch completely for free
100
Towards Data Science @towardsdatascience.com · 26/09/2026
As data pipelines grow, keeping them reliable and maintainable becomes increasingly difficult. @taupirho.bsky.social explains how the Medallion architecture uses Bronze, Silver, and Gold layers to bring structure to data workflows.
towardsdatascience.com
The Medallion Data Architecture: An Introduction | Towards Data Science
A practical guide to Bronze, Silver and Gold, with a working Python and DuckDB example
011
Towards Data Science @towardsdatascience.com · 26/09/2026
What technologies (that aren't AI-focused) should data professionals learn about? Rashi Desai zooms in on 10 areas shaping our future.
towardsdatascience.com
10 Things I’m Learning Beyond AI to Become More Technologically Fluent
Part 1: Understanding the technologies shaping our future
000
Towards Data Science @towardsdatascience.com · 25/09/2026
In the second deep dive from his series on probabilistic forecasting for physical signals, Waleed Esmail explains how to ensure that a forecast preserves honest uncertainty.
towardsdatascience.com
Your Model's MSE Is Lying to You: Part II
Autoregressive rollout and uncertainty propagation. Second in a series on probabilistic forecasting for physical signals.
000
Towards Data Science @towardsdatascience.com · 25/09/2026
What's the best way to connect a RAG system to an agentic workflow? Emmimal P Alexander walks us through the connective layer she built for this purpose.
towardsdatascience.com
RAG Isn't an Agent — I Built the Layer Between Retrieval and Action
RAG retrieves. Agents act. I built both separately, connected them explicitly, and ran the same nine tasks through all three systems.
000
Towards Data Science @towardsdatascience.com · 25/09/2026
"I compared Jev with Qwen on 3,080 bank messages, looking not only at which model made the right decision, but also at speed, confidence, and what happened when the available answers did not quite fit the input." Nhu Hoang shares a detailed comparison of LLMs with new "System One" model, Jev.
towardsdatascience.com
Jev vs. LLMs: When AI Moves from Generation to Decision-Making
I tested TypeSafe AI’s Jev on 3,080 classification tasks to see how its accuracy, latency, calibration, and confidence compare with LLMs — and whether it works as a practical decision layer for AI…
000
Towards Data Science @towardsdatascience.com · 25/09/2026
"I had to figure out a way to make my coding and subscription plans last longer while impacting quality minimally." Eivind Kjosbakken shares actionable tips for getting the most out of your AI coding assistants without depleting the credits you purchased in a day.
towardsdatascience.com
How to Maximize Your Coding Agent Subscriptions
Get more out of your coding agent subscriptions
000
Towards Data Science @towardsdatascience.com · 25/09/2026
"The problem is that both the code and the tests come from the same reading of the same ambiguous sentences, which means the tests cannot disagree with the implementation." Gal Arav tackles a core limitation of coding agents head-on in a new series.
towardsdatascience.com
Towards Spec-Driven Test Automation: Part 1
Why a green test suite can mean nothing
000
Towards Data Science @towardsdatascience.com · 25/09/2026
For his debut TDS article, Hubert García Gordon presents an insightful answer to a thorny question: what does your pipeline output when the correct answer to a query is "nothing"?
towardsdatascience.com
When the Correct Answer Is Nothing, What Does Your Pipeline Return?
The reliability mechanisms we add to LLM pipelines are often the ones that make them confidently wrong.
000
Towards Data Science @towardsdatascience.com · 25/09/2026
Coding agents are capable of much more than writing and debugging software. Eivind Kjosbakken explores how they can handle a wide range of everyday computer tasks.
towardsdatascience.com
How to Apply Coding Agents to Non-Programming Tasks | Towards Data Science
Perform non-programming tasks with coding agents
000
Towards Data Science @towardsdatascience.com · 25/09/2026
A 15-minute booking process can become an automated workflow with the right agent architecture. Soner Yıldırım walks through building a stateful customer support agent with Python, LangGraph, and Langfuse.
towardsdatascience.com
I Replaced a 15-Minute Booking Process with a LangGraph AI Agent | Towards Data Science
A step-by-step guide to building, running, and monitoring a stateful customer support agent using Python, LangGraph, and Langfuse.
000
Towards Data Science @towardsdatascience.com · 24/09/2026
Backpropagation can seem intimidating until you understand the idea that makes it work. Nikhil Dasari breaks down the core concept behind backpropagation in a beginner-friendly way.
towardsdatascience.com
Backpropagation Explained for Beginners (Part 2): There Has to Be a Better Way | Towards Data Science
The idea that makes backpropagation possible.
010
Towards Data Science @towardsdatascience.com · 24/09/2026
Complex mathematical conjectures are easier to understand when the abstract ideas have something concrete to anchor them. James O’Brien uses geometry and algebra to make the Jacobian conjecture more approachable.
towardsdatascience.com
A Simplified View of the Jacobian Conjecture | Towards Data Science
The full conjecture is stated over abstract fields, but the counterexample is a concrete 3D function that we can explain and visualize using familiar geometric ideas and a little algebra.
000
Towards Data Science @towardsdatascience.com · 24/09/2026
A company brain is not built by simply connecting an LLM to a pile of documents. Tomer Mesika explores the context layer needed to turn scattered organizational knowledge into something AI can reliably use.
towardsdatascience.com
How to Build a Context Layer and a Company Brain | Towards Data Science
What it actually takes to turn a company's scattered knowledge into something an LLM can reliably use — and why the demo is 5% of the work.
000
Towards Data Science @towardsdatascience.com · 24/09/2026
We're thrilled to welcome Utkarsh Mangal to TDS with a fascinating debut article, in which he walks us through his attempt to reproduce Anthropic's "Toy Models of Superposition" from scratch in NumPy.
towardsdatascience.com
I Trained a Tiny Network to Compress Data. It Drew a Pentagon.
Reproducing Anthropic's "Toy Models of Superposition" from scratch in NumPy, with hand-derived gradients and no borrowed numbers.
000
Towards Data Science @towardsdatascience.com · 24/09/2026
How do words turn into vectors? Nikhil Dasari's latest explainer explores TF-IDF, vector space, and text classification.
towardsdatascience.com
From Words to Vectors: What Happens in Between?
A Journey through TF-IDF, vector space, and text classification
000
Towards Data Science @towardsdatascience.com · 24/09/2026
"What interests me the most about Group Relative Policy Optimization, or GRPO as most of us know it, is how little feedback the basic setup needs." Benjamin Nweke explains how GRPO trains small language models with verifiable rewards.
towardsdatascience.com
How GRPO Trains Small Language Models with Verifiable Rewards
The mechanics behind local reasoning experiments with Unsloth and why the reward function matters as much as the model.
000
Towards Data Science @towardsdatascience.com · 24/09/2026
Curious about world models — what they are, how they work, and how you can create them? Anubhab Banerjee's deep dive patiently walks us through the process of building a CartPole-focused one in Python.
towardsdatascience.com
How to Make Your First World Model from Scratch
A beginner-friendly guide to building a world model in Python, letting it daydream its way through CartPole, and accurately measuring when the illusion collapses.
001
Towards Data Science @towardsdatascience.com · 24/09/2026
AI agents become far more useful when they can operate beyond APIs and text. Shuai Guo walks through building a browser-use agent with Playwright MCP and the OpenAI Agents SDK.
towardsdatascience.com
How to Give an LLM Agent a Browser | Towards Data Science
Building a browser-use agent with OpenAI Agents SDK and Playwright MCP
010
Towards Data Science @towardsdatascience.com · 23/09/2026
Running an LLM locally still comes with a real energy bill. Justin Stewart measures the actual wall-socket cost of running five models on Apple Silicon and reveals how throughput shapes cost.
towardsdatascience.com
How Much Does a Local LLM Actually Cost to Run? I Measured Every Watt on Apple Silicon | Towards Data Science
Five models, sustained generation, real wall-socket energy at $0.31/kWh — and the surprise the RTX-3090 numbers predicted, only bigger.
110
Towards Data Science @towardsdatascience.com · 23/09/2026
What makes backpropagation possible in the first place? Nikhil Dasari explores the key idea that turns a complex learning process into something easier to understand.
towardsdatascience.com
Backpropagation Explained for Beginners (Part 2): There Has to Be a Better Way | Towards Data Science
The idea that makes backpropagation possible.
000
Towards Data Science @towardsdatascience.com · 23/09/2026
If you'd like to truly stress-test your RAG pipeline before it goes into production, go through the adevrsarial test set that Sara Nobrega outlines in her new guide.
towardsdatascience.com
Break Your Own RAG Pipeline Before Users Do
A small adversarial test set that catches the retrieval failures your evaluation set never will
000