Sign in

David Gasquez

@davidgasquez.com
5.4K followers 1.4K following 1.2K posts

Data @ Protocol Labs. Open Data, Open Source, Open Protocols. Walks taker. Progressive Metal enjoyer. davidgasquez.com

PostsRepliesMedia
Reposted by David Gasquez
danielroe @danielroe.dev · 19/09/2026
don't be nice.
roe.dev
Don't be nice
Sometimes, it's more important to be good than it is to be nice.
41734245
David Gasquez @davidgasquez.com · 18/09/2026
Some of these are really close to reality. 😅 youtu.be/FG8sUgjBGXs
youtu.be
Interview with Big Data engineer in 2026.
YouTube video by Kai Lentit
100
David Gasquez @davidgasquez.com · 01/09/2026
Used @ap.brid.gy to request Nik to bridge his Mastodon posts into Bluesky! Smooth experience and happy to have the posts here. 😊 bsky.app/profile/ludi...
bsky.app
010
Reposted by David Gasquez
Doug Turnbull @softwaredoug.bsky.social · 10/08/2026
This pattern of asking a cheap LLM to generate hypothetical / fake classifications that we then later resolve to real ones is one of the dominant way I'm doing query understanding these days softwaredoug.com/blog/2026/08...
softwaredoug.com
The hallucinating classifier pattern
A useful pattern for LLM classification at scale is to let it hallucinate plausible, fake entities. Then resolve to real ones.
3476
David Gasquez @davidgasquez.com · 26/08/2026
This quote is even more relevant these days with all the "Company Brain" talk going around. Everyone is talking about Retrieval Augmented Generation, but most companies don't actually have any internal documentation worth retrieving. Fix. Your. Shit.
100
Reposted by David Gasquez
Hannes Mühleisen @hannes.muehleisen.org · 26/08/2026
Some big news: @ducklabs.com is joining AWS, with the move expected to complete by early September. Our team will stay together in Amsterdam, with more resources to take DuckDB, DuckLake and Quack further - while the open-source Duck Stack remains free under MIT license ducklabs.com/news/2026/08...
ducklabs.com
DuckLabs to Join AWS, Projects to Remain Open Source
DuckLabs provides services around the DuckDB in-process OLAP data management system directly from its main developers.
106219
David Gasquez @davidgasquez.com · 26/08/2026
Every now and then, I think about this awesome post and imagine how the author must feel with all the increased AI insanity going on right now on every corner of most companies. ludic.mataroa.blog/blog/i-will-...
ludic.mataroa.blog
I Will Fucking Piledrive You If You Mention AI Again — Ludicity
220
Reposted by David Gasquez
Hazel Weakly @hazelweakly.me · 24/12/2025
The two hardest problems in Computer Science are 1. Human communication 2. Getting people in tech to believe that human communication is important
251073265
David Gasquez @davidgasquez.com · 07/08/2026
Wrote about why context engineering is a data problem. Organizations need to own the process that turns raw sources into useful artifacts for their agents. Much of this is work data teams already know how to do! davidgasquez.com/context-engi...
Companies want your data and context because [your context is their moat](https://x.com/samzliu/status/2080210797465379147).
Products that own your context can force you to use their agent... and, I've never seen a good hosted agent!

Organizations need to own their context, and the way to do so effectively already exists and is well understood.

## Context as Data Infrastructure

Have you noticed the [shape of the diagrams](https://storage.googleapis.com/gweb-cloudblog-publish/images/WorkspaceIntelligence.max-2200x2200.png) that every company is adding to their "intelligence" products? I see that and it seems like I'm looking at a Fivetran or dbt product page in 2018. That is because they're the same thing!

Most organizations will need to build and maintain a model-agnostic knowledge base ([aka ontology, company brain, ...](https://x.com/DBredvick/status/2078150905078206789)) for the same reasons they maintain a data warehouse. A curated and normalized layer helps both humans and agents make sense of all the structure that outlives any particular model, harness, tool, or product. Context engineering is [that same work](https://x.com/JoshARosen/status/2084693306722705629): extracting, filtering, curating, modeling, and publishing artifacts to help the organization make better decisions.

For now, it seems there isn't a great set of tools (or [modern context stack](https://x.com/davidgasquez/status/2082895547677954466)) built for this purpose. Something like Fivetran and dbt optimized for the new sources (Slack, Drive) and transformations (transcription, summarization, text extraction from a slide deck, ...).
391
Reposted by David Gasquez
Ronen Tamari @ronentk.me · 07/08/2026
Great post! There is a clear parallel to how the @atproto.science ecosystem can disrupt the extractive science publishing industry & bring on millions of researchers with a new cooperative research stack. The system is ripe for disruption - see also @edhagen.net's post in the thread linked below >
1245
David Gasquez @davidgasquez.com · 01/08/2026
One of the best things I did recently was to archive all @gordon.bsky.social's Substack posts into my personal digital library. I index, embed and query it with github.com/tobi/qmd. When my agents now $consult-library they stumble with many great Gordonisms, making them better for my taste!
 Screenshot of a YAML configuration in a dark-themed code editor. It
 defines a squishy-computer source with a local archive path, Markdown
 glob pattern, commented Substack export command, and context describing
 Gordon Brander’s essays on tools for thought, cybernetics, complex
 systems, decentralized protocols, network society, and AI agents.
280
Reposted by David Gasquez
Simon Späti 🏔️ @ssp.sh · 31/07/2026
I still hope AT Proto will win in the social media race. It has such great ideas and architecture behind it, and it's just there, for everyone to query. Great talk with @danabra.mov, as always.
youtu.be
ATPROTO isn’t just for Bluesky weiners
YouTube video by Syntax
2121
David Gasquez @davidgasquez.com · 31/07/2026
@atproto.com (open data networks in general) allows users to build all kind of personal features or views apps themselves may never have an incentive to build. An app doesn’t have to decide that a certain use case is large/profitable enough to support/build it.
1100
David Gasquez @davidgasquez.com · 31/07/2026
Wanted to take a peek at @semble.so and didn't knew where to start. I built this as a static explorer of all the links in there. sites.wisp.place/davidgasquez... Love AT Protocol for letting me build the UI I want! 😍
Interactive “Semantic Field Map” for the Semble Link Atlas, showing
 31,367 links grouped into 28 color-coded topic clusters. Thousands of
 colored dots form a central map labeled with themes such as Experimental
 Web Projects, AI Industry Criticism, Language Model Engineering, Books
 and Literature, and Open Source Infrastructure; a searchable topic
 legend with link counts appears on the right.
5627
David Gasquez @davidgasquez.com · 16/07/2026
"All happy families are alike; each unhappy family is unhappy in its own way"
020
Reposted by David Gasquez
David Gasquez @davidgasquez.com · 15/07/2026
Write more ad-hoc ephemeral workbenches! They help you learn and a higher bandwith way to communicate with your agent. davidgasquez.com/ephemeral-ag...
When I work with agents, my first instinct is usually to “just prompt” through the task. That is fine for many coding-adjacent tasks, but not everything I do can be easily codified, and I bet the same thing happens to you!

I’ve explored how to make agents do subjective work better, and I think there is another underrated approach to using agents for subjective tasks, or tasks where the goals are fuzzy and you’d like to keep the final call: build local and ephemeral workbenches.

The idea is to have agents produce small pages/apps for the task at hand. For me, the clearest recent example is a grant review workbench I built. Instead of asking a model “which application is better?”, I got it to build me a workbench to help me better understand the task.

In the workbench, I could see and sort every application, add random tags or notes as I read through them, visualize them on a 2D map to spot duplicates or core applications, etc.

The useful part was the handoff. I could click a button to copy the state and ask my agent to do things like complete the labeling from my manual labels, recluster, add a new field to applications, suggest due diligence questions, etc.

This gave both of us a high-bandwidth way to share state and made the feedback loop faster than chat. Chat alone is a bad interface for this kind of work because too much state stays implicit, or even hidden from you. With the workbench, I could inspect both the applications and the changes the agent wanted to make.
011
David Gasquez @davidgasquez.com · 16/07/2026
Wonderful post by @summerscope.bsky.social! Hat tip to @mariozechner.at. The Human-in-the-Loop is Tired. pydantic.dev/articles/the...
pydantic.dev
The Human-in-the-Loop is Tired
On reward functions, dopamine, and what it actually feels like when the code starts writing itself
1346
David Gasquez @davidgasquez.com · 15/07/2026
Write more ad-hoc ephemeral workbenches! They help you learn and a higher bandwith way to communicate with your agent. davidgasquez.com/ephemeral-ag...
When I work with agents, my first instinct is usually to “just prompt” through the task. That is fine for many coding-adjacent tasks, but not everything I do can be easily codified, and I bet the same thing happens to you!

I’ve explored how to make agents do subjective work better, and I think there is another underrated approach to using agents for subjective tasks, or tasks where the goals are fuzzy and you’d like to keep the final call: build local and ephemeral workbenches.

The idea is to have agents produce small pages/apps for the task at hand. For me, the clearest recent example is a grant review workbench I built. Instead of asking a model “which application is better?”, I got it to build me a workbench to help me better understand the task.

In the workbench, I could see and sort every application, add random tags or notes as I read through them, visualize them on a 2D map to spot duplicates or core applications, etc.

The useful part was the handoff. I could click a button to copy the state and ask my agent to do things like complete the labeling from my manual labels, recluster, add a new field to applications, suggest due diligence questions, etc.

This gave both of us a high-bandwidth way to share state and made the feedback loop faster than chat. Chat alone is a bad interface for this kind of work because too much state stays implicit, or even hidden from you. With the workbench, I could inspect both the applications and the changes the agent wanted to make.
011
David Gasquez @davidgasquez.com · 13/07/2026
Complexity is the mind killer. github.com/PhilipK/arti...
Complexity is the mind killer
TL;DR: When faced with a choice, always pick the simplest thing that solves your immediate problem. Then make sure you can change your mind later. By the simplest thing I mean the solution that is easiest to reason about.

When we always pick the simplest solution we avoid the analysis paralysis of finding out which solution might be the best in all the future scenarios we can imagine. We simply pick the simplest thing, and move on, knowing that it is easy to change our mind later. After we implement the simple solution, we have a much better understanding of our problem and we can ask "is my immediate problem solved?" If not; repeat.

Complexity is the mind killer. Our mind gets overloaded when we try to change a complex system, and this causes us to make mistakes. Systems must be changeable, it is what keeps them from becoming legacy and it keeps us agile and adaptable.

A lot of complexity gets added in the name of changeability. We create microservice systems so we can change our programming language and service implementation, but we add the complexity of network, service discovery, message queues, container orchestration and so on. It also adds a resource overhead that forces us to scale horizontally earlier, which is complex. Could we simply have hidden the implementation behind an interface or a function? Will we actually change our programming language any time soon? Does it actually solve our immediate problem?
0100
Reposted by David Gasquez
DuckDB @duckdb.org · 09/07/2026
Check out the Awesome DuckDB repo – curated by @davidgasquez.com – to explore the growing list of handy tools and cool projects built around #DuckDB, DuckLake and more. 😎 🦆 Notice a great tool that’s missing? Submit a PR to get it added. Start exploring: davidgasquez.github.io/awesome-duck...
1272
David Gasquez @davidgasquez.com · 01/07/2026
The primitive is the product. Great read on the power of extensibility and thinking in terms of abstractions and capabilities. www.amplifypartners.com/blog-posts/t...
amplifypartners.com
The primitive is the product | Amplify Partners
Why every software company is now a dev tools company.
011
David Gasquez @davidgasquez.com · 01/07/2026
Came across this old XKCD. Not only works for sports commentaty these days! 😅 xkcd.com/904
 Two stick figures talk. Caption: “A weighted random number generator just
 produced a new batch of numbers.” One says, “Let’s use them to build
 narratives!” Bottom caption: “All sports commentary.”
2110
David Gasquez @davidgasquez.com · 30/06/2026
Integration complete! My posts are now on atproto too thanks to @standard.site. davidgasquez.com
davidgasquez.com
davidgasquez.com
David Gasquez's blog
040
Reposted by David Gasquez
dan @danabra.mov · 30/06/2026
i want to spread the word about atproto. if you’re running a podcast, i’d be happy to hop in to talk about it. i understand it decently well and can explain it without much jargon. it’s one of the most interesting developments in the last 30 years of web, and i can defend that. please dm to set up!
3745455
David Gasquez @davidgasquez.com · 26/06/2026
Good example! UMAP rules. bsky.app/profile/gran...
010
Reposted by David Gasquez
ADRI 👨🏻‍💻 @adrimaqueda.com · 24/06/2026
🌡️ Cada vez que llega una ola de calor me desespera lo difícil que es saber los récords de temperatura. Así que he montado un mapa que rastrea los récords de temperatura de todas las estaciones de AEMET en España y muestra los días que han pasado desde el último récord
182
Reposted by David Gasquez
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 24/06/2026
Proposed norm: you're allowed to respond to an unedited LLM document with an unedited LLM document that is twice as long
717117
David Gasquez @davidgasquez.com · 23/06/2026
Lovely piece analyzing people's relationshipts over time. pudding.cool/2026/06/love... Great visuals and style!
pudding.cool
A love story
Tracking 1,000+ people through the ups and downs of their relationships
000
David Gasquez @davidgasquez.com · 19/06/2026
Got 3.3M links (follows of everyone I follow on Bluesky) into a cute map. These are the bubbles of my feed. Data people, AI people, ATProto builders, OSS, ... sites.wisp.place/davidgasquez...
A web visualization titled “davidgasquez.com follow map”.

 It shows a large network map of followed accounts, represented by circular
 profile images. Accounts are grouped into colored clusters with labels such
 as:

 - Software engineering writers
 - Data Science AI Practitioners
 - R statistics community
 - Tech internet figures
 - Spanish data journalism
 - Data visualization cartography
 - Data Engineering and Analytics
 - Distributed data systems
 - Python data science OSS
 - Open source data tools
 - AI research community
 - AI and ATProto builders
 - Decentralized protocol builders

 On the right, a legend lists each cluster with a color and count. The largest
 groups appear to be AI and ATProto builders, Data Engineering and Analytics,
 Python data science OSS, and AI research community.

 There is also a size scale in the bottom right indicating follower count,
 from “fewer” up to about 202,723 followers. A Reset View button appears near
 the top left.
1151
Reposted by David Gasquez
David Gasquez @davidgasquez.com · 16/06/2026
The future is now! - Dial keys, not IPs, thanks to Iroh - Fetch hashes, not URLs, thanks to DASL - Build on shared records, not siloed databases, thanks to the AT Protocol
0306
David Gasquez @davidgasquez.com · 17/06/2026
Agents are everywhere now but is worth remembering that there are still so many cool things we can do with good ol matrix factorization and embeddings. No LLMs required.
1132
David Gasquez @davidgasquez.com · 17/06/2026
How community notes reduce viral misinformation. Great talk/conversation about community notes internals and impact. www.ted.com/talks/keith_...
ted.com
How Community Notes reduce viral misinformation
Community Notes on X started with a wild idea: Instead of tech companies deciding what's true, what if you let people fact-check each other? Keith Coleman and Jay Baxter, who helped build the crowdsou...
040
David Gasquez @davidgasquez.com · 16/06/2026
The future is now! - Dial keys, not IPs, thanks to Iroh - Fetch hashes, not URLs, thanks to DASL - Build on shared records, not siloed databases, thanks to the AT Protocol
0306
David Gasquez @davidgasquez.com · 12/06/2026
Not sure how well am doing it (feedback appreciated) but from now on my posts are also published to the Atmosphere via @standard.site! 🌱 Tiny sync script in case you want to steal it. 👇 github.com/davidgasquez....
all my posts as they appear on davidgasquez.com
031
David Gasquez @davidgasquez.com · 11/06/2026
Cool to see @hypercerts.org on the What's Hot list of community lexicons! 🤩
1
mu newsFeedPrefs🔥
social.mu.newsFeedPrefs

221 active +Infinity%

2
certified profile🔥
app.certified.actor.profile

151 active +Infinity%

3
hypercerts attachment🔥
org.hypercerts.context.attachment

133 active +Infinity%

4
hypercerts collection🔥
org.hypercerts.collection

130 active +Infinity%

5
certified organization🔥
app.certified.actor.organization

126 active +Infinity%

6
hypercerts activity🔥
org.hypercerts.claim.activity

116 active +Infinity%
153
Reposted by David Gasquez
Martin Kleppmann @martin.kleppmann.com · 09/06/2026
Video and transcript from my talk at QCon London are now online. Thanks @qconferences.com! www.infoq.com/presentation...
infoq.com
Mitigating Geopolitical Risks with Local-First Software and atproto
Martin Kleppmann discusses the urgent need for technological sovereignty in modern infrastructure. Exploring the shifting landscape of global tech dependencies, he shares how engineering leaders can l...
16814
David Gasquez @davidgasquez.com · 03/06/2026
I keep using this qmd setup for research and it is really great! The workflow is simple: put a bunch of resources in a YAML file and qmd will pick them up and index/embed them. Answers are more grounded than just web search and I can control more what the models see!
Screenshot of a dark-themed code editor with a YAML index file open
 beside a Markdown spec titled “Atmospheric Data Portals.”
040
David Gasquez @davidgasquez.com · 02/06/2026
Huh, interesting! I wonder if something like this could work to make it more trustless. ATProto for identity/catalog/social portability, move payment settlement + entitlement state to Ethereum/L2s. The goal would be to make brokers an optional UX infrastructure, not a trusted party.
100
David Gasquez @davidgasquez.com · 01/06/2026
Who is building Strava on the AT Protocol?
2100
Reposted by David Gasquez
Mario Zechner @mariozechner.at · 31/05/2026
I wrote up how I built the shitty robot so you can too. This was a fun project that will keep on giving. Thanks to all the open weights folks out there, without whom this would not have been possible. mariozechner.at/posts/2026-0...
mariozechner.at
How to build a shitty robot
A work in progress post about building a shitty robot.
4709
David Gasquez @davidgasquez.com · 29/05/2026
Content defined chunking and contend addressed storage rocks indeed!
Screenshot of a tweet by Julien Chaumond saying Hugging Face is becoming bullish on data infrastructure. The tweet says he cloned 68 TB to a Hugging Face training bucket in under two minutes using Xet deduplication and infrastructure optimizations. Below the text is a video still of a Hugging Face “Copy to bucket” dialog showing a copy in progress from jasperai/monet to julien-c/my-training-bucket/monet-copy, with 178+ files, total size 68.2 TB, and an estimated time of less than one minute.
161
David Gasquez @davidgasquez.com · 29/05/2026
TIL about Off Protocol! A podcast from the Bluesky DevRel team with coversations around the AT Protocol and the open social web.
110
Reposted by David Gasquez
Emily Hunt @emily.space · 28/05/2026
Floating point numbers would be AMAZING to have on the atproto for scientific data - a huge amount of data is floating point, and not being able to have them here adds a lot of complexity to posting things. This is a great proposal from @vmx.cx that demonstrates how they could be added:
1225
David Gasquez @davidgasquez.com · 27/05/2026
Really enjoyed this talk from @bcantrill.bsky.social on why trust _is_ infrastructure. youtu.be/WF7J7qtZ8TA Recommended watching!
youtu.be
Trust as Infrastructure | Bryan Cantrill | Monktoberfest 2025
YouTube video by RedMonk Tech Events
093
David Gasquez @davidgasquez.com · 23/05/2026
Wrote a small post advocating for more mindful sharing of raw AI outputs. I'm feeling it more and more (specially at work) and would love to reiterate it once more in writting.✌️ davidgasquez.com/keep-your-sl...
You’ve probably heard this already but, I keep coming across this pattern and I wanted to add another post to the cause.

Generating a wall of text is now free while reading, verifying, and distilling still cost time and effort on the recipient. That asymmetry is what makes you sharing that document consisting of raw unrequested AI output rude. As this vibecoded website says, stop sloppypasta!

The overall principle I keep in mind is what you send should take you more effort to produce than it takes me to read. This applies to chats, but also to code, bug reports, PRs, emails, docs, and now, it seems slides too.

This is also especially bad when the point of the writing is to demonstrate your thinking. Rewriting that with an LLM changes meaning subtly, blurs authorship, and erodes voice. And people can tell.

There is also “good uses of slop” though! I think it is ok to:

Send a draft you’ve read, edited, and would defend as yours.
Share the output as a side artifact. I like to see other folks’ raw prompts and sessions (I’ve learned a lot from reading Simon Willison prompts) so I can learn from it or fork it and tweak it.
So, two small asks:

Keep slop to yourself and share drafts once you’ve read it, verified it, and distilled them down to what actually matters.
If you do share it, disclose it. Ideally send the prompt and a link to the chat rather than the output. Sessions are more interesting than transcripts, and the recipient can tweak and run their own models on top of it (I don’t usually trust people’s LLM context management skills). If your prompt is too embarrassing to share, that’s a signal worth listening to.
These new manners are still emerging and I’m probably also offending folks with this sometimes (slopiness is an spectrum). But, let’s at least try not to outsource our thinking onto each other. Otherwise, we all end up drowning in slop or delegating to agents to get more slop!
1120
Reposted by David Gasquez
Vicki @vickiboykis.com · 19/05/2026
New post: Tagging my blog with BERTopic and LLMs. I suspect as token costs increase, we'll see more blended traditional ML/LLM systems. And, in general, they work really well! vickiboykis.com/2026/05/18/t...
vickiboykis.com
Tagging my blog posts with BERTopic and LLMs
LLMs mean you still need a human in the loop, but in a different part
4549
Reposted by David Gasquez
⿻ Audrey Tang 唐鳳 @audreyt.org · 18/05/2026
You are more than welcome, @antirez.bsky.social.🙏 Creating opportunities & expanding horizons is our collective responsibility in the Age of AI.🌌 ▶️ pi.audreyt.org Let’s #WagePeace🕊️ & #FreeTheFuture — together!🫂 #LLAP🖖
pi.audreyt.org
pi-ds4 · Audrey Tang
Run a frontier model on your own machine with stable, contestable decision traces. Full install, steering, reproducibility, and tuning guide.
0392
David Gasquez @davidgasquez.com · 18/05/2026
Wrote about why the AT Protocol is the ultimate API for you and your agents! davidgasquez.com/atproto-agents We need more people building on the atmosphere!
My atmosphere data lives at at://davidgasquez.com. You can explore it all without API keys, auth, or arbitrary HTML to parse. Your agent can browse it, query it, and link to it. Getting someone latest posts is one click/curl away.

That’s the pitch. The rest of this post is why it works, and why it might be the substrate your agents have been waiting for.

Hostile Platforms
You’ve probably experienced more than once recently, platforms being hostile to you (or your agent) getting the data out: throttled or limited APIs, bans for scraping, no identity persistence, no real-time access, walled gardens. Every integration is a deal that can be revoked when the CEO wakes up in a bad mood.

The AT Protocol is an amazing technology that I think is very underrated for agents! Turns out, the properties Bluesky (the biggest AT Protocol application as of today) needed for humans (portable identity, open data, federated infrastructure, structured schemas) happen to be exactly what you’d want for agents too.

As a quick and dirty intro (you should read the official ones, or this great walkthrough by mackuba), in the atproto world, every user is a personal, signed JSON repository. Every record (a post, a like, a follow, a photo, a blog post, anything really) is addressable by an at:// URI, has a public schema (a Lexicon), and is broadcast in real time on a global event stream (the firehose). Dan Abramov explains all of this in more detail and clarity in “A Social Filesystem”.
512515
David Gasquez @davidgasquez.com · 16/05/2026
I dont want to interact with your agent, I want my data!
020
David Gasquez @davidgasquez.com · 14/05/2026
What an awesome line-up of humans working together! 😍 github.com/antirez/ds4/...
github.com
feat(server): add /v1/responses (OpenAI Responses API) for Codex CLI by audreyt · Pull Request #91 · antirez/ds4
Implements the Responses API endpoint that Codex CLI (and other modern OpenAI tooling) speaks instead of /v1/chat/completions. The wire format is documented in OpenAI's Responses API; this impl...
020