Sign in

David Gasquez

@davidgasquez.com
5.4K followers 1.4K following 1.2K posts

Data @ Protocol Labs. Open Data, Open Source, Open Protocols. Walks taker. Progressive Metal enjoyer. davidgasquez.com

PostsRepliesMedia
Reposted by David Gasquez
danielroe @danielroe.dev · 19/09/2026
don't be nice.
roe.dev
Don't be nice
Sometimes, it's more important to be good than it is to be nice.
41734246
David Gasquez @davidgasquez.com · 18/09/2026
Some of these are really close to reality. 😅 youtu.be/FG8sUgjBGXs
youtu.be
Interview with Big Data engineer in 2026.
YouTube video by Kai Lentit
100
David Gasquez @davidgasquez.com · 16/09/2026
I should've replied in here actually but I saw it first on X. 😅 Thanks for re-sharing it here Simon!
020
David Gasquez @davidgasquez.com · 16/09/2026
Really cool! So many things coming up around open data in the atmosphere. Need to do another pass at davidgasquez.com/atmospheric-... and github.com/orgs/datonic...!
020
David Gasquez @davidgasquez.com · 01/09/2026
Used @ap.brid.gy to request Nik to bridge his Mastodon posts into Bluesky! Smooth experience and happy to have the posts here. 😊 bsky.app/profile/ludi...
bsky.app
010
David Gasquez @davidgasquez.com · 31/08/2026
Creo que aún no hay gente suficiente que lo conozca, menos aún en España, pero promete bastante!
000
David Gasquez @davidgasquez.com · 31/08/2026
Yo me siento igual! La verdad es que ahora veo muchas aplicaciones y pienso... en atproto estaría mejor! Es el mejor enfoque en descentralización usable que he visto, y es una pasada ver todo lo que está haciendo la gente encima del ecosistema.
111
Reposted by David Gasquez
Doug Turnbull @softwaredoug.bsky.social · 10/08/2026
This pattern of asking a cheap LLM to generate hypothetical / fake classifications that we then later resolve to real ones is one of the dominant way I'm doing query understanding these days softwaredoug.com/blog/2026/08...
softwaredoug.com
The hallucinating classifier pattern
A useful pattern for LLM classification at scale is to let it hallucinate plausible, fake entities. Then resolve to real ones.
3476
David Gasquez @davidgasquez.com · 26/08/2026
This quote is even more relevant these days with all the "Company Brain" talk going around. Everyone is talking about Retrieval Augmented Generation, but most companies don't actually have any internal documentation worth retrieving. Fix. Your. Shit.
100
Reposted by David Gasquez
Hannes Mühleisen @hannes.muehleisen.org · 26/08/2026
Some big news: @ducklabs.com is joining AWS, with the move expected to complete by early September. Our team will stay together in Amsterdam, with more resources to take DuckDB, DuckLake and Quack further - while the open-source Duck Stack remains free under MIT license ducklabs.com/news/2026/08...
ducklabs.com
DuckLabs to Join AWS, Projects to Remain Open Source
DuckLabs provides services around the DuckDB in-process OLAP data management system directly from its main developers.
106219
David Gasquez @davidgasquez.com · 26/08/2026
Every now and then, I think about this awesome post and imagine how the author must feel with all the increased AI insanity going on right now on every corner of most companies. ludic.mataroa.blog/blog/i-will-...
ludic.mataroa.blog
I Will Fucking Piledrive You If You Mention AI Again — Ludicity
220
Reposted by David Gasquez
Hazel Weakly @hazelweakly.me · 24/12/2025
The two hardest problems in Computer Science are 1. Human communication 2. Getting people in tech to believe that human communication is important
251073265
David Gasquez @davidgasquez.com · 21/08/2026
Thanks! Excited for the public release. Been working and thinking a lot about these kind of agents / knowledge context repositories and am trying to gather examples of real world companies. davidgasquez.com/context-engi...
davidgasquez.com
Context Engineering Is a Data Problem
Companies want your data and context because your context is their moat. Products that own your context can force you to use their agent... and, I've never seen a good hosted agent! Organizations need to own their context, and the way to do so effectively already exists and is...
000
David Gasquez @davidgasquez.com · 20/08/2026
Any resources/links to the Bluesky company knowledge base agent? Sounds cool and would love to learn more about it!
160
David Gasquez @davidgasquez.com · 08/08/2026
So cool! I tried with "a laptop" and got this awesome picture.
Screenshot of a British Library image viewer showing an 1873 illustration of a tilted book or stone slab surrounded by an arch of flowers. A dark metadata panel lists the French title, publication date, collection, scan link, and file path.
010
David Gasquez @davidgasquez.com · 07/08/2026
Totally! 💯 Got convinced when I realized I was spending more time on my github.com/tobi/qmd setup than on the actual classical data engineering side of my job.
010
David Gasquez @davidgasquez.com · 07/08/2026
Wrote about why context engineering is a data problem. Organizations need to own the process that turns raw sources into useful artifacts for their agents. Much of this is work data teams already know how to do! davidgasquez.com/context-engi...
Companies want your data and context because [your context is their moat](https://x.com/samzliu/status/2080210797465379147).
Products that own your context can force you to use their agent... and, I've never seen a good hosted agent!

Organizations need to own their context, and the way to do so effectively already exists and is well understood.

## Context as Data Infrastructure

Have you noticed the [shape of the diagrams](https://storage.googleapis.com/gweb-cloudblog-publish/images/WorkspaceIntelligence.max-2200x2200.png) that every company is adding to their "intelligence" products? I see that and it seems like I'm looking at a Fivetran or dbt product page in 2018. That is because they're the same thing!

Most organizations will need to build and maintain a model-agnostic knowledge base ([aka ontology, company brain, ...](https://x.com/DBredvick/status/2078150905078206789)) for the same reasons they maintain a data warehouse. A curated and normalized layer helps both humans and agents make sense of all the structure that outlives any particular model, harness, tool, or product. Context engineering is [that same work](https://x.com/JoshARosen/status/2084693306722705629): extracting, filtering, curating, modeling, and publishing artifacts to help the organization make better decisions.

For now, it seems there isn't a great set of tools (or [modern context stack](https://x.com/davidgasquez/status/2082895547677954466)) built for this purpose. Something like Fivetran and dbt optimized for the new sources (Slack, Drive) and transformations (transcription, summarization, text extraction from a slide deck, ...).
391
David Gasquez @davidgasquez.com · 07/08/2026
Yes! Offer an opinionated alphabet and a bunch of simple rules.
020
Reposted by David Gasquez
Ronen Tamari @ronentk.me · 07/08/2026
Great post! There is a clear parallel to how the @atproto.science ecosystem can disrupt the extractive science publishing industry & bring on millions of researchers with a new cooperative research stack. The system is ripe for disruption - see also @edhagen.net's post in the thread linked below >
1245
David Gasquez @davidgasquez.com · 05/08/2026
They come from a clustering algorithm and, yes, they're more or less arbitrary! I better approach is to do what @zzstoatzz.io did here and have multiple resolutions for the clusters. Something for the next one! pub-search.waow.tech/atlas
pub-search.waow.tech
pub search / atlas
2d semantic map of atproto publishing platforms
120
David Gasquez @davidgasquez.com · 02/08/2026
Yep! I totally understand the trade-offs there as someone in both ends (wants deletions to be respected and is a data engineer that loves open datasets).
020
David Gasquez @davidgasquez.com · 02/08/2026
This is a really sweet approach for bulk open data sharing. I've been using it for a while for datasets on the ~400gb range and so far so good! Wrote about it a while back! davidgasquez.com/modern-open-... And this is a great post too. www.robinlinacre.com/parquet_api/
020
David Gasquez @davidgasquez.com · 02/08/2026
Will dig a bit more on Pagefind, thanks for sharing! I think I had a different "federated search" concept. More like a search engine over community curated resources.
020
David Gasquez @davidgasquez.com · 02/08/2026
You could even run SQL queries from your browser (DuckDB WASM) ang get only the data you're interested on of, say, r2://data.microcosm.com/2026/05/22/data.parquet
120
David Gasquez @davidgasquez.com · 02/08/2026
Sure! It would be awesome to have daily compressed files of all the data you're gathering for analytical/batch purposes. Something like a partitioned folder of compressed ND-JSON or Parquet files on R2 would be cheap and very powerful!
220
David Gasquez @davidgasquez.com · 02/08/2026
Would you consider serving daily partitioned files over an object storage?
200
David Gasquez @davidgasquez.com · 02/08/2026
I wrote about a cool pattern I've been using for community level indexing of interest resources a while back. davidgasquez.com/sharing-a-qm... I have an old draft I should get out with the complete picture but it has been working very well (and very cheap) for Filecoin. github.com/davidgasquez...
davidgasquez.com
Indexing and Sharing Organizational Context with qmd
Data engineering is mostly context gathering. It involves tons of time spent going through code, docs, specs, issues, and conversations to figure out how stakeholders want active users to be counted. qmd turns that pile of documents into something you can actually query withou...
110
David Gasquez @davidgasquez.com · 02/08/2026
Awesome! I was going to ask for this actually after dealing with Substack for the export! bsky.app/profile/davi...
010
David Gasquez @davidgasquez.com · 01/08/2026
One of the best things I did recently was to archive all @gordon.bsky.social's Substack posts into my personal digital library. I index, embed and query it with github.com/tobi/qmd. When my agents now $consult-library they stumble with many great Gordonisms, making them better for my taste!
 Screenshot of a YAML configuration in a dark-themed code editor. It
 defines a squishy-computer source with a local archive path, Markdown
 glob pattern, commented Substack export command, and context describing
 Gordon Brander’s essays on tools for thought, cybernetics, complex
 systems, decentralized protocols, network society, and AI agents.
280
David Gasquez @davidgasquez.com · 01/08/2026
Do you use any atproto backed podcast app?
210
David Gasquez @davidgasquez.com · 01/08/2026
I see voyage-4! Wanted to use it but ended up with a Qwen model. Have you used both or have any thoughts on Voyage? Love the labels at different levels too. Great idea I might steal in future atlases. 😀
010
David Gasquez @davidgasquez.com · 01/08/2026
So cool! Much better UX too. Which models/approach did you use?
110
Reposted by David Gasquez
Simon Späti 🏔️ @ssp.sh · 31/07/2026
I still hope AT Proto will win in the social media race. It has such great ideas and architecture behind it, and it's just there, for everyone to query. Great talk with @danabra.mov, as always.
youtu.be
ATPROTO isn’t just for Bluesky weiners
YouTube video by Syntax
2121
David Gasquez @davidgasquez.com · 31/07/2026
It is amazing all the things you can do when platforms don't lock you out of your data! Even more these days of agents. I was curious about Semble this morning and was able to build a custom UI with a couple of prompts. bsky.app/profile/davi...
110
David Gasquez @davidgasquez.com · 31/07/2026
I have a couple of long flights soon and I know what I'm going to watch!
030
David Gasquez @davidgasquez.com · 31/07/2026
I should've shared the @tangled.org repository instead of the GitHub one. Shame on me. tangled.org/davidgasquez...
tangled.org
davidgasquez.com/sembleverse
🌌 Explore Semble links as a semantic map
020
David Gasquez @davidgasquez.com · 31/07/2026
Yes! Pushed the highly vibe-coded scripts to github.com/davidgasquez.... There are a few things that could improve the final output (serialization as Markdown, building a FAISS index, using Qwen, ...) so take it as it is, a 10m PoC!
130
David Gasquez @davidgasquez.com · 31/07/2026
I've built personal recommendation systems, custom UIs, and LLM classifiers, .... All using powerful models that woulnd't scale to the entire Bluesky network.
010
David Gasquez @davidgasquez.com · 31/07/2026
@atproto.com (open data networks in general) allows users to build all kind of personal features or views apps themselves may never have an incentive to build. An app doesn’t have to decide that a certain use case is large/profitable enough to support/build it.
1100
David Gasquez @davidgasquez.com · 31/07/2026
Also fun to see clusters like Bone Inmune Biology, Creative Culture Projects, and others that are not only Bluesky adjacent!
Zoomed-in semantic map showing interconnected clusters of colored links, including Alternative Internet Research, Bluesky User Profiles and Apps, AT Protocol resources, Tangled Code Repositories, Experimental Web Projects, Software Code Repositories, and Open Source Infrastructure.
260
David Gasquez @davidgasquez.com · 31/07/2026
Wanted to take a peek at @semble.so and didn't knew where to start. I built this as a static explorer of all the links in there. sites.wisp.place/davidgasquez... Love AT Protocol for letting me build the UI I want! 😍
Interactive “Semantic Field Map” for the Semble Link Atlas, showing
 31,367 links grouped into 28 color-coded topic clusters. Thousands of
 colored dots form a central map labeled with themes such as Experimental
 Web Projects, AI Industry Criticism, Language Model Engineering, Books
 and Literature, and Open Source Infrastructure; a searchable topic
 legend with link counts appears on the right.
5627
David Gasquez @davidgasquez.com · 28/07/2026
Love it! I might steal this idea too. Thanks for sharing!
011
David Gasquez @davidgasquez.com · 24/07/2026
Any chance this gets open sourced? I'd love to learn more about how it works! 👀 Would love to also have wathever SKILL/MCP/tools Attie is using in the open so I can use my agent instead!
620
David Gasquez @davidgasquez.com · 23/07/2026
@daviddao.org might be the one you're looking for! Certified is a component of @hypercerts.org! docs.hypercerts.org They've been running some rounds on atproto, like the recent maearth.com one.
120
David Gasquez @davidgasquez.com · 17/07/2026
I can imagine! Perhaps someday!
020
David Gasquez @davidgasquez.com · 16/07/2026
"All happy families are alike; each unhappy family is unhappy in its own way"
020
Reposted by David Gasquez
David Gasquez @davidgasquez.com · 15/07/2026
Write more ad-hoc ephemeral workbenches! They help you learn and a higher bandwith way to communicate with your agent. davidgasquez.com/ephemeral-ag...
When I work with agents, my first instinct is usually to “just prompt” through the task. That is fine for many coding-adjacent tasks, but not everything I do can be easily codified, and I bet the same thing happens to you!

I’ve explored how to make agents do subjective work better, and I think there is another underrated approach to using agents for subjective tasks, or tasks where the goals are fuzzy and you’d like to keep the final call: build local and ephemeral workbenches.

The idea is to have agents produce small pages/apps for the task at hand. For me, the clearest recent example is a grant review workbench I built. Instead of asking a model “which application is better?”, I got it to build me a workbench to help me better understand the task.

In the workbench, I could see and sort every application, add random tags or notes as I read through them, visualize them on a 2D map to spot duplicates or core applications, etc.

The useful part was the handoff. I could click a button to copy the state and ask my agent to do things like complete the labeling from my manual labels, recluster, add a new field to applications, suggest due diligence questions, etc.

This gave both of us a high-bandwidth way to share state and made the feedback loop faster than chat. Chat alone is a bad interface for this kind of work because too much state stays implicit, or even hidden from you. With the workbench, I could inspect both the applications and the changes the agent wanted to make.
011
David Gasquez @davidgasquez.com · 16/07/2026
Wonderful post by @summerscope.bsky.social! Hat tip to @mariozechner.at. The Human-in-the-Loop is Tired. pydantic.dev/articles/the...
pydantic.dev
The Human-in-the-Loop is Tired
On reward functions, dopamine, and what it actually feels like when the code starts writing itself
1346
David Gasquez @davidgasquez.com · 15/07/2026
Write more ad-hoc ephemeral workbenches! They help you learn and a higher bandwith way to communicate with your agent. davidgasquez.com/ephemeral-ag...
When I work with agents, my first instinct is usually to “just prompt” through the task. That is fine for many coding-adjacent tasks, but not everything I do can be easily codified, and I bet the same thing happens to you!

I’ve explored how to make agents do subjective work better, and I think there is another underrated approach to using agents for subjective tasks, or tasks where the goals are fuzzy and you’d like to keep the final call: build local and ephemeral workbenches.

The idea is to have agents produce small pages/apps for the task at hand. For me, the clearest recent example is a grant review workbench I built. Instead of asking a model “which application is better?”, I got it to build me a workbench to help me better understand the task.

In the workbench, I could see and sort every application, add random tags or notes as I read through them, visualize them on a 2D map to spot duplicates or core applications, etc.

The useful part was the handoff. I could click a button to copy the state and ask my agent to do things like complete the labeling from my manual labels, recluster, add a new field to applications, suggest due diligence questions, etc.

This gave both of us a high-bandwidth way to share state and made the feedback loop faster than chat. Chat alone is a bad interface for this kind of work because too much state stays implicit, or even hidden from you. With the workbench, I could inspect both the applications and the changes the agent wanted to make.
011
David Gasquez @davidgasquez.com · 14/07/2026
Would be also awesome to trick @koaning.bsky.social into making some atproto content in his incredibly fun educational style! 😜
110