Sign in

Data Elixir

@dataelixir.com
738 followers 829 following 110 posts

Data Elixir is a weekly newsletter with curated data science picks from around the web. Subscribe at dataelixir.com and follow us here for selections between issues. Covering machine learning, data visualization, analytics, and strategy.

PostsRepliesMedia
Data Elixir @dataelixir.com · 07/07/2026
Nice opportunity for researchers and applied folks working with text data: Text as Data 2026 is accepting presentation proposals. The focus is on computational analysis of documents across politics, society, culture, social science and the humanities. One-page proposals are due Aug 1. tada2026.org
tada2026.org
Text as Data 2026
New Directions in Analyzing Text as Data (TADA) — Monday, October 5, 2026 at UC Berkeley. One-page submissions due August 1.
031
Data Elixir @dataelixir.com · 20/02/2026
"You need a vector database" is the default advice but complexity means there's new infrastructure, new formats, and data coordination overhead. What if tuning Parquet page sizes and adding lightweight footer metadata gave you vector search without any of that? blog.xiangpeng.systems/posts/vector...
blog.xiangpeng.systems
Vector search using only Parquet and DataFusion – Xiangpeng’s blog
Just tune your existing stack
020
Data Elixir @dataelixir.com · 13/02/2026
Most data bugs aren't algorithm failures, they're assumption violations upstream. Pointblank makes validation a first-class citizen in R/Python workflows instead of an afterthought scattered in assert statements.
posit.co
Fifty Releases of Pointblank: a Year of Building Data Quality Tooling - Posit
The Pointblank Python library's journey has resulted in a premier, data-agnostic toolkit for tabular data validation.
041
Data Elixir @dataelixir.com · 13/02/2026
Plot twist: You can have Normal residuals even when Y isn't Normal. And non-Normal residuals when Y actually is Normal. Model misspecification shows up everywhere except where you expect it.
robertkubinec.com
All Your Data R Fixed – Homepage
In all conventional regression models, the data are fixed–they don’t have a probability distribution. This means transforming your data with logs or other functions won’t change the probability…
000
Data Elixir @dataelixir.com · 06/02/2026
Here's what they don't tell you about topic modeling: preprocessing determines everything. Corpus homogeneity, document length, and context windows matter more than algorithm choice. BERTopic won't fix bad inputs. No model will.
css-polytechnique.github.io
The General Inquirer in the time of LLMs: a BERTopic tutorial – CSS@IPP — Tutorials and resources
A. Morin
093
Data Elixir @dataelixir.com · 05/02/2026
Indexes speed up reads but slow writes. The real insight: indexes only help when you're returning <15-20% of rows. Beyond that, sequential scans win. Space and memory costs matter too. btree indexes often exceed table size.
dlt.github.io
Introduction to PostgreSQL Indexes
Who’s this for Basics How data is stored in disk How indexes speedup access to data Costs associated with indexes Disk Space Write operations Query planner Memory usage Types of Indexes Btree Hash…
010
Data Elixir @dataelixir.com · 05/02/2026
What if a model’s prediction was literally an eigenvalue? This post kicks off a thoughtful series exploring spectrum-based models as a middle ground between linear models and neural nets, with interpretability and robustness baked in. A fun, slightly off-the-beaten-path ML read.
alexshtf.github.io
Behold the power of the spectrum!
Eigenvalues as neurons: represent nonlinear models as the k-th eigenvalue of a learned symmetric matrix pencil. Explore monotonicity/convexity properties and train simple spectral models.
010
Data Elixir @dataelixir.com · 04/02/2026
Data To Art is a new curated gallery transforming datasets into visual storytelling. Latest addition: Alisa Singer's Environmental Graphiti turns climate science into vibrant data-driven art. The line between science and abstraction is thinner than you think.
data-to-art.com
Data to Art
A curated gallery celebrating data visualization as art. Discover innovative artworks from international artists who transform data into emotionally compelling visual experiences.
020
Data Elixir @dataelixir.com · 10/01/2026
ClickHouse solved Advent of Code 2025 puzzles using single ClickHouse queries. No UDFs, no temp tables, no preprocessing. Just pure SQL doing things SQL probably shouldn't do. Impressive and slightly cursed in the best way.
clickhouse.com
Solving the "Impossible" in ClickHouse: Advent of Code 2025
At ClickHouse, we don't like the word "impossible." We believe that with the right tools, everything is a data problem. To prove it, we decided to complete the 2025 Advent of Code unconventionally:…
000
Data Elixir @dataelixir.com · 09/01/2026
If you're wondering where database tech is really headed in 2026, Andy Pavlo's annual review cuts through the noise. Sharding, Postgres evolution, and hot takes you won't find in vendor blogs.
cs.cmu.edu
Databases in 2025: A Year in Review
The world tried to kill Andy off but he had to stay alive to to talk about what happened with databases in 2025.
010
Data Elixir @dataelixir.com · 06/01/2026
MarkItDown (Microsoft) converts a variety file types to Markdown optimized for LLMs. Handles structure, not just text. Pairs well with Kreuzberg for extraction (supports 50+ file formats!). If you're building RAG systems, these belong in your stack.
tryolabs.com
Top Python libraries of 2025
Explore our 11th annual Top Python Libraries roundup, featuring two curated Top 10 lists for General Use and AI / ML / Data tools that matter today.
030
Data Elixir @dataelixir.com · 03/01/2026
jax-js compiles NumPy-style array code into WebAssembly and WebGPU kernels that run entirely client-side. No server, no dependencies, just JAX's programming model in your browser. This changes what's possible for interactive ML demos.
ss.ekzhang.com
jax-js: an ML library for the web
JAX in pure JavaScript, as a flexible machine learning library and compiler.
000
Data Elixir @dataelixir.com · 02/01/2026
Jan Van Haaren's 2025 soccer analytics review is a solid reference if you're doing applied ML in sports or want to see how domain experts handle sequential data. Covers spatio-temporal models, graph methods, Bayesian forecasting, and tracking data metrics. janvanhaaren.be/posts/soccer...
janvanhaaren.be
Soccer Analytics 2025 Review – Jan Van Haaren
Collection of the soccer analytics content that I liked the most in 2025!
010
Data Elixir @dataelixir.com · 20/12/2025
ML progress isn't driven by elegant theory. It's benchmarks, leaderboards, and engineering culture. In this post, Ben Recht explores why empirical testing beats clean math in practice and why that tension defines the field.
argmin.net
Benchmark Studies
It is impossible to disentangle technical innovation from technical debt
020
Data Elixir @dataelixir.com · 19/12/2025
The paradox: show a complex graph and execs check their phones. Show a simple one and they demand endless breakdowns. The solution isn't more or less data. It's recentering discussions on the actual decision at hand. methodmatters.github.io/true-stories...
methodmatters.github.io
True Stories from the (Data) Battlefield – Part 1: Communicating About Data
A blog about data science, statistics, and data analysis with open-source software.
000
Data Elixir @dataelixir.com · 18/12/2025
Your browser can run Python (Pyodide), execute OCR on PDFs, crop videos, and call LLM APIs—all without uploading anything to a server. The localStorage + CORS pattern makes surprisingly powerful tools possible with zero backend infrastructure.
simonwillison.net
Useful patterns for building HTML tools
I’ve started using the term HTML tools to refer to HTML applications that I’ve been building which combine HTML, JavaScript, and CSS in a single file and use them to …
000
Data Elixir @dataelixir.com · 17/12/2025
Tired of charts that hide the story? Density plots + percentile intervals reveal what averages can't: full variability, extremes, historical context. Nine examples comparing Spain's temperature data show how geometry choice changes what readers understand.
dominicroye.github.io
Broken Chart: discover 9 visualization alternatives
Researcher in climate science at MBG-CSIC
010
Data Elixir @dataelixir.com · 16/12/2025
Fisher arbitrarily chose p<0.05 a century ago and we've just... kept it. The problem: calling it "arbitrary" only works if you can suggest something less arbitrary. No one has, so here we are circling p-values close to 0.05 like it means something. vilgot-huhn.github.io/mywebsite/po...
000
Data Elixir @dataelixir.com · 15/12/2025
Haskell for data science? dataHaskell adds dataframes, NSE-style column operations, and compiler optimizations that turn chained operations into single-pass computations. Immutability + strong types + functional composition might be the combo we've been missing. jcarroll.com.au/2025/12/05/h...
jcarroll.com.au
Haskell IS a Great Language for Data Science
I’ve been learning Haskell for a few years now and I am really liking a lot of the features, not least the strong typing and functional approach. I thought it was lacking some of the things I missed…
060
Data Elixir @dataelixir.com · 05/12/2025
Learning SQL is like learning a foreign language: you need to read more variations than you'll actually write. Learn disciplined canonical syntax for your own queries, but understand the messy dialects others use.
kb.databasedesignbook.com
A modern guide to SQL JOINs
There are many SQL JOINs guides and tutorials, but this one takes a different approach. We try to avoid misleading wording and imagery, and we structure the material in a different way. The goal of…
110
Data Elixir @dataelixir.com · 01/12/2025
Side projects still open more doors than traditional applications in data roles. The trick: build small, interesting things that signal skills and make your work discoverable. Especially relevant given the current job market.
presentofcoding.substack.com
Make Things, Tell People
On side projects and finding work
083
Data Elixir @dataelixir.com · 29/11/2025
Most time-to-event metrics are broken. Amazon thought customer support wait times were under 1 min. Bezos called in: 10+ minutes. Why? They only measured customers who stayed on hold long enough to be served. The ones who hung up? Never counted. www.counting-stuff.com/why-you-are-...
counting-stuff.com
Why You Are (Probably) Measuring Time Wrong: Why do we need to use Survival Analysis more
Author: Michał Chorowski [Hey everyone! A guest post from Michał this week! This newsletter is always willing to host/share data-related content, so if you've created anything you'd like to share,…
030
Data Elixir @dataelixir.com · 29/11/2025
Hot take: Your analysts doing "some basic data engineering" is killing your analytics function. The MTA hired 5 dedicated data engineers and it unlocked everything else. Stop asking data scientists to maintain pipelines. www.mta.info/article/less...
mta.info
Lessons learned in starting a central data team
Learn how the MTA succeeded in setting up a central data team and a general purpose, cloud-based platform for data analytics.
073
Data Elixir @dataelixir.com · 24/11/2025
Sometimes the best polars pattern is knowing when to exit the DataFrame. partitionby() splits data into a dict of frames, letting you process with list comprehensions. Cleaner than forcing everything through mapgroups() when further wrangling isn't needed.
emilyriederer.com
Python Rgonomics: User-defined functions in polars | Emily Riederer
Polars provides a consistent API for conducting transformations against a DataFrame. But what do you do when you need to apply a user-defined function beyond the native API? This post surveys the…
010
Data Elixir @dataelixir.com · 23/11/2025
Most devs treat AI coding agents like infinite context machines. Reality: a 200k token window fills fast. The /compact feature is a trap. Better approach: /clear + document state in markdown, then resume. Treat context like disk space. You need a cleanup strategy.
blog.sshh.io
How I Use Every Claude Code Feature
A brain dump of all the ways I've been using Claude Code.
110
Data Elixir @dataelixir.com · 22/11/2025
Most modern dimensionality reduction (t-SNE, UMAP, Isomap) shares a pattern: represent data as a graph capturing local similarity, then embed to preserve that structure. It's graphs all the way down.
alechelbling.com
A Visual Introduction to Dimensionality Reduction with Isomap
"To deal with hyper-planes in a 14-dimensional space, visualize a 3D space and say 'fourteen' to yourself very loudly. Everyone does it." - Geoffrey Hinton
010
Reposted by Data Elixir
Maria Antoniak @mariaa.bsky.social · 14/11/2025
I curated some readings for class on "data tensions" and the list felt worth sharing. Come on a tour of datasets, books, the web, and AI with me... We'll start with this piece on the Google Books project: the hopes, dreams, disasters, and aftermath of building a public library on the internet. 1/n
theatlantic.com
Torching the Modern-Day Library of Alexandria
“Somewhere at Google there is a database containing 25 million books and nobody is allowed to read them.”
612042
Data Elixir @dataelixir.com · 14/11/2025
Everyone's rushing to pgvector for "simple" vector search in Postgres. This reality check shows what actually happens at scale: indexing nightmares and performance walls. Simple isn't always sustainable in production.
alex-jacobs.com
The Case Against pgvector | Alex Jacobs
What happens when you try to run pgvector in production and discover all the things the blog posts conveniently forgot to mention
000
Data Elixir @dataelixir.com · 13/11/2025
When healthcare becomes algorithmic, what gets optimized out? This Guardian essay asks the hard question about AI spreading through diagnostics and therapy: are we trading care quality for efficiency without realizing the cost?
theguardian.com
What we lose when we surrender care to algorithms | Eric Reinhart
A dangerous faith in AI is sweeping American healthcare – with consequences for the basis of society itself
010
Data Elixir @dataelixir.com · 08/11/2025
Thinking Machines Lab solved a problem everyone accepted as unsolvable: LLM nondeterminism at temperature 0. Same prompt, same model, 1000 runs → 80 different outputs. With batch-invariant kernels? Bitwise identical every time. Open sourced. www.distributedthoughts.org/will-i-make-...
distributedthoughts.org
Will I Make It To The Restaurant Before The Soup Dumplings Get Cold? (And Other Problems In Machine Learning)
I'm chronically late. Not because I want to be rude - I feel terrible about it every single time - but because I'm catastrophically bad at predicting how long it takes to get anywhere. Turns out…
010
Data Elixir @dataelixir.com · 06/11/2025
Most marketplaces have SKUs. Etsy has 100M+ unique items with no standard attributes. How do you build filters when one listing is a "porcelain sculpture that looks like a t-shirt" and dimensions live in random photo text? www.etsy.com/codeascraft/...
etsy.com
020
Data Elixir @dataelixir.com · 31/10/2025
GeoUtil converts between GeoJSON, TopoJSON, Shapefile, KML, WKT, and CSV without touching a server. TopoJSON compression alone cuts file sizes 80%+ while preserving topology. All free, all browser-based. geoutil.com
geoutil.com
GeoUtil — Free Online Map & Geography Tools
All-in-one online geography toolkit. Measure distance & area, convert GeoJSON, TopoJSON, JSON, merge or minify files, and more — fast, free, and browser-based.
000
Data Elixir @dataelixir.com · 30/10/2025
Debugging constraint problems is backwards: remove constraints until something works, then figure out what broke. No stack traces, just an "unsatisfiable." Forces you to think differently about what you're actually asking the system to solve. www.righto.com/2025/10/solv...
righto.com
Solving the NYTimes Pips puzzle with a constraint solver
The New York Times recently introduced a new daily puzzle called Pips . You place a set of dominoes on a grid, satisfying various condition...
000
Data Elixir @dataelixir.com · 29/10/2025
Most interpretable models sacrifice accuracy. Most accurate models are black boxes. TRUST breaks this trade-off by combining parametric and non-parametric approaches and offers full prediction explanations without losing performance. Published at PRICAI 2025.
pypi.org
trust-free
Transparent, Robust & Ultra-Sparse Trees (TRUST™) - Free Version
000
Data Elixir @dataelixir.com · 29/10/2025
Data products shouldn't live forever by default. Netflix treats outdated metrics like deprecated software and actively sunset them. The cost of maintaining zombie datasets? Lost trust and accumulated technical debt that blocks innovation.
netflixtechblog.medium.com
Data as a Product: Applying a Product Mindset to Data at Netflix
Introduction: What if we treated data with the same care and intentionality as a consumer-facing product? Adopting a “data as a product”…
000
Reposted by Data Elixir
Cadence @cadence-gis.bsky.social · 24/10/2025
What will you be using this #30DayMapChallenge? Cadence is offering its users £2,500 worth of prizes this #30daymapchallenge. And every user (new or existing) gets a free Professional upgrade for November! Learn more: cadence.cityscience.com/blog/30-day-...
cadence.cityscience.com
30 Day Map Challenge
The 30 Day Map Challenge with Cadence – November 2025  This November, Cadence is proud to support the 30 Day Map Challenge – a global celebration of maps, creativity and storytelling. Whether …
011
Data Elixir @dataelixir.com · 24/10/2025
Why does neural network training almost never fail? Pure combinatorics. A 6K parameter network contains 10^1089 possible sparse subnetworks. That's 10^900 solutions per atom in the universe. We're not smart, we're just brute forcing.
embedding-space.github.io
Sparse Networks and Lottery Winners
Embedding Space is a blog about machine learning and artificial intelligence.
040
Data Elixir @dataelixir.com · 23/10/2025
The best part about #30DayMapChallenge is that it's tool-agnostic. Whether you're using QGIS, Python's geopandas, R's rayshader, or even Blender for 3D visualizations, the focus is on creativity over tech stack. No programming required.
30daymapchallenge.com
30DayMapChallenge
Daily mapping challenge happening every November!
110
Data Elixir @dataelixir.com · 23/10/2025
Stop fitting separate binomial models to compositional data! Your predictions that 130% of respondents chose option A reveal fundamental model misspecification. Dirichlet regression with Gaussian processes respects constraints.
ecogambler.netlify.app
Compositional modeling of plant communities with Dirichlet regression | GAMbler
Compositional data appears everywhere in scientific research, yet many analysts fall back on problematic approaches that ignore fundamental mathematical constraints. I demonstrate how Dirichlet…
010
Data Elixir @dataelixir.com · 17/10/2025
AI systems are now generating, testing, and validating their own hypotheses. DeepMind's Co-Scientist and Stanford's Virtual Lab represent something new: AI as actual scientific collaborator, not just a fancy search engine.
stateof.ai
State of AI Report 2025
The State of AI Report analyses the most interesting developments in AI. Read and download here.
000
Data Elixir @dataelixir.com · 17/10/2025
Black and white, hand-drawn data viz countering visual noise at one of NYC's busiest hubs. Smart choice. Sometimes the most effective data art isn't about adding more complexity but finding clarity in the chaos through intentional restraint.
pentagram.com
‘A Data Love Letter to the Subway’
A data-driven animation for Fulton Center commissioned by MTA Arts & Design for its 40th anniversary.
020
Data Elixir @dataelixir.com · 15/10/2025
The metrics that matter most are the hardest to measure. Bookings take weeks to materialize, but clicks happen instantly. The trap: optimizing for clicks can actually decrease bookings. Correlation isn't causation, especially in A/B tests.
booking.ai
How to estimate correlation between metrics from past A/B tests
Authors: Miha Gazvoda, Christina Katsimerou
000
Data Elixir @dataelixir.com · 15/10/2025
Psychology's dirty secret: "data available upon request" usually means data not available at all. Researchers found systematic patterns in why data disappears over time. Open-washing is real and it's undermining reproducibility.
open.lnu.se
LnuOpen | Meta-Psychology
Many journals now require data sharing and require articles to include a Data Availability Statement. However, several studies over the past two decades have shown that promissory notes about data…
011
Reposted by Data Elixir
Carl Zimmer @carlzimmer.com · 09/10/2025
Today my @nytimes.com colleagues and I are launching a new series called Lost Science. We interview US scientists who can no longer discover something new about our world, thanks to this year‘s cuts. Here is my first interview with a scientist who studied bees and fires. Gift link: nyti.ms/3IWXbiE
nyti.ms
14147041812
Data Elixir @dataelixir.com · 13/10/2025
Parquet is showing its age. CMU researchers built F3 - a columnar format that embeds WebAssembly decoders directly in files. Universal compatibility without the usual compatibility hell. Smart approach for modern ML workloads. db.cs.cmu.edu/papers/2025/...
db.cs.cmu.edu
021
Data Elixir @dataelixir.com · 10/10/2025
Mathematicians feared nuclear winter would freeze Earth, but it turns out CO2 might do it instead. The math behind climate tipping points is fascinating and terrifyingly unpredictable. Sometimes the thing you're not watching is the real threat.
quantamagazine.org
The Math of Climate Change Tipping Points | Quanta Magazine
Tipping points in our climate predictions are both wildly dramatic and wildly uncertain. Can mathematicians make them useful?
000
Reposted by Data Elixir
Crystal Lewis @cghlewis.bsky.social · 08/10/2025
Data dictionary template: osf.io/ynqcu Project summary template: osf.io/q6g8d Dataset level README template: osf.io/tk4cb
0274
Data Elixir @dataelixir.com · 09/10/2025
"Silicon samples" - using LLMs to generate fake survey responses instead of recruiting humans. Sounds efficient until you realize small model tweaks completely flip your results. Shortcuts in research usually aren't.
arxiv.org
The threat of analytic flexibility in using large language models to simulate human data: A call to attention
Social scientists are now using large language models to create "silicon samples" - synthetic datasets intended to stand in for human respondents, aimed at revolutionising human subjects research.…
082
Data Elixir @dataelixir.com · 08/10/2025
Your ggplot2 charts work fine, but are they memorable? Real color engineering: brightness first (strongest differentiator), then hue, finally saturation. Most people get this backwards and wonder why their viz falls flat. www.chartography.net/p/color-engi...
chartography.net
Color Engineering
The tool that snaps mercurial design into mechanical focus.
0189
Data Elixir @dataelixir.com · 30/09/2025
bsky.app/profile/libb...
000