Sign in

David Gasquez

@davidgasquez.com
5.4K followers 1.4K following 1.2K posts

Data @ Protocol Labs. Open Data, Open Source, Open Protocols. Walks taker. Progressive Metal enjoyer. davidgasquez.com

PostsRepliesMedia
David Gasquez @davidgasquez.com · 08/08/2026
So cool! I tried with "a laptop" and got this awesome picture.
Screenshot of a British Library image viewer showing an 1873 illustration of a tilted book or stone slab surrounded by an arch of flowers. A dark metadata panel lists the French title, publication date, collection, scan link, and file path.
010
David Gasquez @davidgasquez.com · 07/08/2026
Wrote about why context engineering is a data problem. Organizations need to own the process that turns raw sources into useful artifacts for their agents. Much of this is work data teams already know how to do! davidgasquez.com/context-engi...
Companies want your data and context because [your context is their moat](https://x.com/samzliu/status/2080210797465379147).
Products that own your context can force you to use their agent... and, I've never seen a good hosted agent!

Organizations need to own their context, and the way to do so effectively already exists and is well understood.

## Context as Data Infrastructure

Have you noticed the [shape of the diagrams](https://storage.googleapis.com/gweb-cloudblog-publish/images/WorkspaceIntelligence.max-2200x2200.png) that every company is adding to their "intelligence" products? I see that and it seems like I'm looking at a Fivetran or dbt product page in 2018. That is because they're the same thing!

Most organizations will need to build and maintain a model-agnostic knowledge base ([aka ontology, company brain, ...](https://x.com/DBredvick/status/2078150905078206789)) for the same reasons they maintain a data warehouse. A curated and normalized layer helps both humans and agents make sense of all the structure that outlives any particular model, harness, tool, or product. Context engineering is [that same work](https://x.com/JoshARosen/status/2084693306722705629): extracting, filtering, curating, modeling, and publishing artifacts to help the organization make better decisions.

For now, it seems there isn't a great set of tools (or [modern context stack](https://x.com/davidgasquez/status/2082895547677954466)) built for this purpose. Something like Fivetran and dbt optimized for the new sources (Slack, Drive) and transformations (transcription, summarization, text extraction from a slide deck, ...).
391
David Gasquez @davidgasquez.com · 01/08/2026
One of the best things I did recently was to archive all @gordon.bsky.social's Substack posts into my personal digital library. I index, embed and query it with github.com/tobi/qmd. When my agents now $consult-library they stumble with many great Gordonisms, making them better for my taste!
 Screenshot of a YAML configuration in a dark-themed code editor. It
 defines a squishy-computer source with a local archive path, Markdown
 glob pattern, commented Substack export command, and context describing
 Gordon Brander’s essays on tools for thought, cybernetics, complex
 systems, decentralized protocols, network society, and AI agents.
280
David Gasquez @davidgasquez.com · 31/07/2026
Also fun to see clusters like Bone Inmune Biology, Creative Culture Projects, and others that are not only Bluesky adjacent!
Zoomed-in semantic map showing interconnected clusters of colored links, including Alternative Internet Research, Bluesky User Profiles and Apps, AT Protocol resources, Tangled Code Repositories, Experimental Web Projects, Software Code Repositories, and Open Source Infrastructure.
260
David Gasquez @davidgasquez.com · 31/07/2026
Wanted to take a peek at @semble.so and didn't knew where to start. I built this as a static explorer of all the links in there. sites.wisp.place/davidgasquez... Love AT Protocol for letting me build the UI I want! 😍
Interactive “Semantic Field Map” for the Semble Link Atlas, showing
 31,367 links grouped into 28 color-coded topic clusters. Thousands of
 colored dots form a central map labeled with themes such as Experimental
 Web Projects, AI Industry Criticism, Language Model Engineering, Books
 and Literature, and Open Source Infrastructure; a searchable topic
 legend with link counts appears on the right.
5627
David Gasquez @davidgasquez.com · 15/07/2026
Write more ad-hoc ephemeral workbenches! They help you learn and a higher bandwith way to communicate with your agent. davidgasquez.com/ephemeral-ag...
When I work with agents, my first instinct is usually to “just prompt” through the task. That is fine for many coding-adjacent tasks, but not everything I do can be easily codified, and I bet the same thing happens to you!

I’ve explored how to make agents do subjective work better, and I think there is another underrated approach to using agents for subjective tasks, or tasks where the goals are fuzzy and you’d like to keep the final call: build local and ephemeral workbenches.

The idea is to have agents produce small pages/apps for the task at hand. For me, the clearest recent example is a grant review workbench I built. Instead of asking a model “which application is better?”, I got it to build me a workbench to help me better understand the task.

In the workbench, I could see and sort every application, add random tags or notes as I read through them, visualize them on a 2D map to spot duplicates or core applications, etc.

The useful part was the handoff. I could click a button to copy the state and ask my agent to do things like complete the labeling from my manual labels, recluster, add a new field to applications, suggest due diligence questions, etc.

This gave both of us a high-bandwidth way to share state and made the feedback loop faster than chat. Chat alone is a bad interface for this kind of work because too much state stays implicit, or even hidden from you. With the workbench, I could inspect both the applications and the changes the agent wanted to make.
011
David Gasquez @davidgasquez.com · 13/07/2026
Complexity is the mind killer. github.com/PhilipK/arti...
Complexity is the mind killer
TL;DR: When faced with a choice, always pick the simplest thing that solves your immediate problem. Then make sure you can change your mind later. By the simplest thing I mean the solution that is easiest to reason about.

When we always pick the simplest solution we avoid the analysis paralysis of finding out which solution might be the best in all the future scenarios we can imagine. We simply pick the simplest thing, and move on, knowing that it is easy to change our mind later. After we implement the simple solution, we have a much better understanding of our problem and we can ask "is my immediate problem solved?" If not; repeat.

Complexity is the mind killer. Our mind gets overloaded when we try to change a complex system, and this causes us to make mistakes. Systems must be changeable, it is what keeps them from becoming legacy and it keeps us agile and adaptable.

A lot of complexity gets added in the name of changeability. We create microservice systems so we can change our programming language and service implementation, but we add the complexity of network, service discovery, message queues, container orchestration and so on. It also adds a resource overhead that forces us to scale horizontally earlier, which is complex. Could we simply have hidden the implementation behind an interface or a function? Will we actually change our programming language any time soon? Does it actually solve our immediate problem?
0100
David Gasquez @davidgasquez.com · 01/07/2026
Came across this old XKCD. Not only works for sports commentaty these days! 😅 xkcd.com/904
 Two stick figures talk. Caption: “A weighted random number generator just
 produced a new batch of numbers.” One says, “Let’s use them to build
 narratives!” Bottom caption: “All sports commentary.”
2110
David Gasquez @davidgasquez.com · 19/06/2026
Got 3.3M links (follows of everyone I follow on Bluesky) into a cute map. These are the bubbles of my feed. Data people, AI people, ATProto builders, OSS, ... sites.wisp.place/davidgasquez...
A web visualization titled “davidgasquez.com follow map”.

 It shows a large network map of followed accounts, represented by circular
 profile images. Accounts are grouped into colored clusters with labels such
 as:

 - Software engineering writers
 - Data Science AI Practitioners
 - R statistics community
 - Tech internet figures
 - Spanish data journalism
 - Data visualization cartography
 - Data Engineering and Analytics
 - Distributed data systems
 - Python data science OSS
 - Open source data tools
 - AI research community
 - AI and ATProto builders
 - Decentralized protocol builders

 On the right, a legend lists each cluster with a color and count. The largest
 groups appear to be AI and ATProto builders, Data Engineering and Analytics,
 Python data science OSS, and AI research community.

 There is also a size scale in the bottom right indicating follower count,
 from “fewer” up to about 202,723 followers. A Reset View button appears near
 the top left.
1151
David Gasquez @davidgasquez.com · 12/06/2026
Not sure how well am doing it (feedback appreciated) but from now on my posts are also published to the Atmosphere via @standard.site! 🌱 Tiny sync script in case you want to steal it. 👇 github.com/davidgasquez....
all my posts as they appear on davidgasquez.com
031
David Gasquez @davidgasquez.com · 11/06/2026
Cool to see @hypercerts.org on the What's Hot list of community lexicons! 🤩
1
mu newsFeedPrefs🔥
social.mu.newsFeedPrefs

221 active +Infinity%

2
certified profile🔥
app.certified.actor.profile

151 active +Infinity%

3
hypercerts attachment🔥
org.hypercerts.context.attachment

133 active +Infinity%

4
hypercerts collection🔥
org.hypercerts.collection

130 active +Infinity%

5
certified organization🔥
app.certified.actor.organization

126 active +Infinity%

6
hypercerts activity🔥
org.hypercerts.claim.activity

116 active +Infinity%
153
David Gasquez @davidgasquez.com · 03/06/2026
I keep using this qmd setup for research and it is really great! The workflow is simple: put a bunch of resources in a YAML file and qmd will pick them up and index/embed them. Answers are more grounded than just web search and I can control more what the models see!
Screenshot of a dark-themed code editor with a YAML index file open
 beside a Markdown spec titled “Atmospheric Data Portals.”
040
David Gasquez @davidgasquez.com · 29/05/2026
Content defined chunking and contend addressed storage rocks indeed!
Screenshot of a tweet by Julien Chaumond saying Hugging Face is becoming bullish on data infrastructure. The tweet says he cloned 68 TB to a Hugging Face training bucket in under two minutes using Xet deduplication and infrastructure optimizations. Below the text is a video still of a Hugging Face “Copy to bucket” dialog showing a copy in progress from jasperai/monet to julien-c/my-training-bucket/monet-copy, with 178+ files, total size 68.2 TB, and an estimated time of less than one minute.
161
David Gasquez @davidgasquez.com · 23/05/2026
Wrote a small post advocating for more mindful sharing of raw AI outputs. I'm feeling it more and more (specially at work) and would love to reiterate it once more in writting.✌️ davidgasquez.com/keep-your-sl...
You’ve probably heard this already but, I keep coming across this pattern and I wanted to add another post to the cause.

Generating a wall of text is now free while reading, verifying, and distilling still cost time and effort on the recipient. That asymmetry is what makes you sharing that document consisting of raw unrequested AI output rude. As this vibecoded website says, stop sloppypasta!

The overall principle I keep in mind is what you send should take you more effort to produce than it takes me to read. This applies to chats, but also to code, bug reports, PRs, emails, docs, and now, it seems slides too.

This is also especially bad when the point of the writing is to demonstrate your thinking. Rewriting that with an LLM changes meaning subtly, blurs authorship, and erodes voice. And people can tell.

There is also “good uses of slop” though! I think it is ok to:

Send a draft you’ve read, edited, and would defend as yours.
Share the output as a side artifact. I like to see other folks’ raw prompts and sessions (I’ve learned a lot from reading Simon Willison prompts) so I can learn from it or fork it and tweak it.
So, two small asks:

Keep slop to yourself and share drafts once you’ve read it, verified it, and distilled them down to what actually matters.
If you do share it, disclose it. Ideally send the prompt and a link to the chat rather than the output. Sessions are more interesting than transcripts, and the recipient can tweak and run their own models on top of it (I don’t usually trust people’s LLM context management skills). If your prompt is too embarrassing to share, that’s a signal worth listening to.
These new manners are still emerging and I’m probably also offending folks with this sometimes (slopiness is an spectrum). But, let’s at least try not to outsource our thinking onto each other. Otherwise, we all end up drowning in slop or delegating to agents to get more slop!
1120
David Gasquez @davidgasquez.com · 18/05/2026
Wrote about why the AT Protocol is the ultimate API for you and your agents! davidgasquez.com/atproto-agents We need more people building on the atmosphere!
My atmosphere data lives at at://davidgasquez.com. You can explore it all without API keys, auth, or arbitrary HTML to parse. Your agent can browse it, query it, and link to it. Getting someone latest posts is one click/curl away.

That’s the pitch. The rest of this post is why it works, and why it might be the substrate your agents have been waiting for.

Hostile Platforms
You’ve probably experienced more than once recently, platforms being hostile to you (or your agent) getting the data out: throttled or limited APIs, bans for scraping, no identity persistence, no real-time access, walled gardens. Every integration is a deal that can be revoked when the CEO wakes up in a bad mood.

The AT Protocol is an amazing technology that I think is very underrated for agents! Turns out, the properties Bluesky (the biggest AT Protocol application as of today) needed for humans (portable identity, open data, federated infrastructure, structured schemas) happen to be exactly what you’d want for agents too.

As a quick and dirty intro (you should read the official ones, or this great walkthrough by mackuba), in the atproto world, every user is a personal, signed JSON repository. Every record (a post, a like, a follow, a photo, a blog post, anything really) is addressable by an at:// URI, has a public schema (a Lexicon), and is broadcast in real time on a global event stream (the firehose). Dan Abramov explains all of this in more detail and clarity in “A Social Filesystem”.
512515
David Gasquez @davidgasquez.com · 23/04/2026
Two things I've done am very happy with: - Publishing a SKILL.md file with references to datasets docs, code, and other useful context (how to plot) - Index anything a data engineer could need. Core business logic docs, external repositories, APIs, ...
SKILL file
020
David Gasquez @davidgasquez.com · 23/04/2026
Two months later, happy to report the barefoot data platforms approach has been working quite well! Shared some learnings on a small post. davidgasquez.com/growing-my-o...
It has been a couple of months since I wrote about Barefoot Data Platforms. Since then, I’ve been applying those ideas into a real data platform I’ve been rebuilding from scratch. Learned a lot about what makes these platforms work well, especially with agents.

Building It
As a quick overview, the entire orchestration engine (fdp) is ~700 lines of Python. Assets are plain .sql or .py files with ugly but useful metadata in comments. No decorators, no registration, no config files. You can drop a file in assets/, and the next run picks it up, resolves dependencies, and materializes it! Delete the file, and it’s gone too.

This minimalist, low-abstraction and filesystem based conventions data platform is great fit for agents:
1110
David Gasquez @davidgasquez.com · 24/03/2026
In earlier eras of computing unreliable hardware forced system design discipline (checksums, retries, modularity). There is a lot of that we are rediscovering in the "Agentic Engineering" era! Wrote a small post about this idea. davidgasquez.com/reliable-unr...
Agentic engineering is mostly about building reliable systems around unreliable components (like your friendly coding agent).

A good analogy I like is how early computers were powerful, but not trustworthy enough to be used “raw”. Hardware failed, bits flipped, and storage was noisy. Engineers had to build systems around the machine to make things more predictable. And, they came up with a bunch of interesting ideas!

Error correction codes
Redundancy
Checksums
Validation layers
Retry logic
With these, even though reliability was not a property of the machine/computer, it became a property of the system.

This is something close to where we are with “Agentic Engineering”.

As you’ve probably experienced, a coding agent can be fluent, useful, and very wrong at the same time! The key is to, like the early programmers did, treat agents as noisy components instead of trying to get the perfect/bigger one.
2207
David Gasquez @davidgasquez.com · 19/03/2026
Made a silly CLI to run queries over markdown files and their frontmatter inspired by Obsidian Bases. github.com/davidgasquez...
Terminal running through the tool examples.
271
David Gasquez @davidgasquez.com · 18/03/2026
New wallpaper perhaps?
A person in flowing white, sheer fabric stands on a stone balcony inside a large circular architectural space, their gauzy garments billowing dramatically as if caught by wind. Behind them, a tall narrow doorway opens into darkness, framed by concentric textured patterns that emphasize the scale and symmetry of the structure.
040
David Gasquez @davidgasquez.com · 12/03/2026
Built a tiny pi.dev extension so I can launch task-specific agents. Each profile is a small YAML file: model, thinking level, system prompt, and allowed skills. Simple, contained, and surprisingly useful! davidgasquez.com/specializing...
I’ve written about specializing Codex and Claude Code in the past. Here is how to do something similar for Pi!

I made a small profile extension so I can launch somewhat contained agents with pi --profile <name>.1

A Pi profile in this case is just a small YAML file with some keys like model, thinking level, system prompt, and an allowlist of skills.

model: openai-codex/gpt-5.4
thinking: medium
system: You are SAM, a pragmatic assistant. Concise and useful.
skills:
  - todoist-cli
  - agent-browser
When Pi starts, the extension loads that profile, switches to the configured model, replaces the system prompt for the turn, limits which /skill:* commands are allowed, and keeps sessions isolated per profile.
391
David Gasquez @davidgasquez.com · 20/02/2026
Cleaned and updated my notes on Agentic Engineering. davidgasquez.com/handbook/age...
An agent runs tools in a loop to achieve a goal. Agents are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks. Agents help decouple programming, the craft of physically typing code, from engineering, the architecture of your system, the goals, the “why” of what you’re building.

Using LLMs for coding is difficult and unintuitive, requiring significant effort to master.
Don’t delegate thinking, delegate work.
Before coding, make the plan with the model.
You can use the same or a different model to critique the plan and iterate. If you are unsure, ask to “give a few options before making changes”.
Redoing work is extremely cheap. Prioritize exploration over execution (at first). Iterate towards precision during the brainstorming phase. Start fresh once you know what and how to build it.
Failed attempts are cheap. If the plan fails and the result is bad, just delete everything and try again.
Divide the problem into smaller problems (functions, classes, …) and solve them one by one. Keep sessions short.
Use Progressive Disclosure to ensure that the agent only sees tasks or project-specific instructions when it needs them.
1101
David Gasquez @davidgasquez.com · 11/02/2026
Half baked idea and proof of concept but at least it is out there. I wrote about Tributary, a credibly neutral mechanism to incentivize and elicit useful datasets. davidgasquez.com/tributary-da... I worked on it 6 months ago at a research retreat (IERR 2025) but never wrote about it.
Text of the blog post
161
David Gasquez @davidgasquez.com · 10/02/2026
Also made a cute embedding of all their products! datania.github.io/mercadona-ca...
embedding of all mercadonas
162
David Gasquez @davidgasquez.com · 10/02/2026
Discovered Mercadona's API and wrote a quick scrapper that will run weekly and dumps everything into JSONs in HuggingFace. Useful to track prices over time, check new arrivals, ... Data. huggingface.co/datasets/dat... Code. github.com/datania/merc...
fiiles
060
David Gasquez @davidgasquez.com · 09/02/2026
Spotted this interesting design/implementation of a self-hosted context store for agents. - Turn DAG - Blob Deduplication / Content Addresses Storage - BLAKE3 - Dynamic types, visual debugging, ... github.com/strongdm/cxdb
CXDB - AI Context Store
CXDB is an AI Context Store for agents and LLMs, providing fast, branch-friendly storage for conversation histories and tool outputs with content-addressed deduplication.

Built on a Turn DAG + Blob CAS architecture, CXDB gives you:

Branch-from-any-turn: Fork conversations at any point without copying history
Fast append: Optimized for the 99% case - appending new turns
Content deduplication: Identical payloads stored once via BLAKE3 hashing
Type-safe projections: Msgpack storage with typed JSON views for UIs
Built-in UI: React frontend with turn visualization and custom renderers
4122
David Gasquez @davidgasquez.com · 04/02/2026
Pasted an screenshot of the conversation and got this done! I don't know if you get used to this feeling. 😅 davidgasquez.github.io/barefoot-dat...
-- asset.name = filtered_numbers
-- asset.schema = raw
-- asset.description = Filtered values from transformed numbers
-- asset.depends = raw.transformed_numbers
select
    value,
    square,
    label,
    double,
    parity
from raw.transformed_numbers
where value >= 3
200
David Gasquez @davidgasquez.com · 03/02/2026
Wrote about "Barefoot Data Platforms", my take on minimal and agent friendly data platforms! davidgasquez.com/barefoot-dat...
So, how does a minimal and opinionated data platform look like? Minimalism for me here means low abstraction and no frameworks. You write your scripts. These scripts are assets, and these assets depend on each other 1. That’s it. The opinion side of things is the interesting part. Here are the key ideas of the Barefoot Data Platform I built:

Classic functional, declarative, independent, composable, and idempotent transformations
Co-located assets, metadata, and documentation so assets have everything needed to be understood and reproduced in the same file
One storage layer. DuckDB or SQLite, Parquet files, or whatever you want as long as you define a materialize function
Many automated checks (ruff with all the rules enabled, ty, data tests, prek hooks, custom bdp check, …) for the agents
2262
David Gasquez @davidgasquez.com · 24/01/2026
TIL you can install Python tools without uv or Python with a single curl command! uvx.sh Useful for Docker, CI pipelines, or anywhere you don't want to mess with the Python installation or venvs!
uvx.sh
Install Python tools with a single command. Powered by uv.

For example, to install ruff:
050
David Gasquez @davidgasquez.com · 19/12/2025
Updated the design of my personal website and I love it! davidgasquez.com/handbook
Company Knowledge Management
All Organizations produce some kind of knowledge. If not properly managed, it’ll lie on your ex-employees, oral history and tribal knowledge.

If we think about a company as an organism, then a knowledge management system is essentially the (collective) brain that keeps that organism alive and running. A corporate knowledge management system, ideally, contains every single bit of codifiable information within the company resulting in a library of all projects, processes and procedures.

Managing an organization’s knowledge is mainly a people problem, not a technology problem. If the Culture is not properly set up, no one will do it and no amount of technology is going to help you magic it out of the ether.

---

All monospace.
0110
David Gasquez @davidgasquez.com · 22/10/2025
I've been using this pattern to "specialize" Codex for vaguely defined tasks like classification, filtering, soft sorting, ... davidgasquez.com/specializing... Made more than 10,000 invocations so far (reusing my ChatGPT subsciption) and am really happy with the pattern!
’ve been using Claude Code to take care of scrappy data cleaning tasks for a while. These days though, I’m using Codex as my coding agent. Similar to what I did with Claude, I’ve been “fine-tuning” Codex CLI to work on a few different vaguely defined tasks like classification, voting, filtering, or ranking.

The pattern in this post works surprisingly well when you have the following conditions:

Loosely defined open-ended tasks. e.g., tagging tweets with a set of predefined labels, extracting structured information from a GitHub issue, …
Powerful agentic capabilities. Doing the task requires something more than a simple llm call or PydanticAI script. e.g., using gh api CLI to get the number of stars of a repository.
Structured outputs. You need a response in a certain shape! This is something codex exec can do that claude couldn’t and is really powerful. e.g., return exactly True or False and nothing else.
Save money! Unlike llm or other tools/libraries that require an OPENAI_API_KEY, Codex can use your ChatGPT subscription, making things “free”.
0101
David Gasquez @davidgasquez.com · 23/07/2025
Asked Claude to move my handbook notes from Obsdian to my Astro website. Very happy with the results! davidgasquez.com/handbook/dat... I'll keep the Obsidian Publish version around while I work on making it look as good.
Data Culture
The data team needs to be focus on delivering insights and supporting decisions. The outcome of the data team are decisions and a shared context across the organization that makes coordination easier.
Your goal as a data professional is to facilitate decision making and help surface/investigate the performance of a business (e.g. operational).
Learning to drive decisions quickly, a bias to action, is a critical competency for an analyst. Every skill you learn – communication, writing, experimentation, metric design – supports this.
If analysis is not actionable, it does not really matter. Analysis must drive to action. Clear results won’t spur action themselves. The organization needs to be ready to pivot when something isn’t working.
Data doesn’t make decisions, people do.
Data’s impact is tough to measure — it doesn’t always translate to value
The value of “insights” is often unknown.
The Data Team should be building and iterating the Data Product.
Notebooks are a workshop. Production systems are the factory. Not everything needs to be put into production. Not everything should be a notebook. You need both. Lean in to the strength of each.
Data is fundamentally a collaborative design process rather than a tool, an analysis, or even a product. Data works best when the entire feedback loop from idea to production is an iterative process.
To get buy in, explain how the business could benefit from better data (e.g: more and better insights). Start small and show value.
Run Purpose Meetings or Business Metrics Review.
Purpose Meetings are 30 min meetings in which stakeholders, engineers and data align on the goal of a release and what is the best way to evaluate the impact and understand its success. Align on the goal, commit on metrics and design the data.Data Culture 
The data team needs to be focus on delivering insights and supporting decisions. The outcome of the data team are decisions and a shared context across the organization that makes coordination easier.
Your goal as a data professional is to facilitate decision making and help surface/investigate the performance of a business (e.g. operational).
Learning to drive decisions quickly, a bias to action, is a critical competency for an analyst. Every skill you learn – communication, writing, experimentation, metric design – supports this.
140
David Gasquez @davidgasquez.com · 21/07/2025
The end dataset also serves as a much smoother API! - No key required - Batch export friendly - Explorable in the browser
❯ curl -s "https://huggingface.co/datasets/datania/aemet/raw/main/valores-climatologicos/2025/07/16.json" | jq '.[] | select(.nombre == "GRANADA AEROPUERTO")'
{
  "fecha": "2025-07-16",
  "indicativo": "5530E",
  "nombre": "GRANADA AEROPUERTO",
  "provincia": "GRANADA",
  "altitud": "560",
  "tmed": "30,3",
  "tmin": "18,8",
  "horatmin": "05:08",
  "tmax": "41,8",
  "horatmax": "16:09",
  "dir": "11",
  "velmedia": "2,2",
  "racha": "10,3",
  "horaracha": "16:30",
  "sol": "13,6",
  "presMax": "955,3",
  "horaPresMax": "Varias",
  "presMin": "951,0",
  "horaPresMin": "16",
  "hrMedia": "29",
  "hrMax": "60",
  "horaHrMax": "Varias",
  "hrMin": "13",
  "horaHrMin": "16:02"
}
~
020
David Gasquez @davidgasquez.com · 16/07/2025
Updated linkweaver (custom notes & research helper) with a couple of new tricks. 1. Saves the URL content locally 2. Accepts any command to preprocess the generated Markdown. E.g: `llm "Clean the following Markdown"`. github.com/davidgasquez... Want to try it? `uvx linkweaver` and enjoy! 😀
Screenshot of linkweaver running. It is downloading a few urls❯ uvx linkweaver --help
usage: linkweaver [-h] [--list-links] [--xml-cat] [-x COMMAND] [-v] [--retries N] [-q] [--dry-run] [--no-color] [-f] [input_files ...]

Extract URLs from markdown files and save each link as individual markdown files in resource folders.
By default, existing files are skipped to save time and bandwidth.

positional arguments:
  input_files         One or more markdown files to process

options:
  -h, --help          show this help message and exit
  --list-links, -l    List all unique URLs found in the files (no downloading)
  --xml-cat           Concatenate files with their resource folders in XML structure
  -x, --exec COMMAND  Execute a shell command on the markdown output before saving (e.g., 'llm -t mdclean')
  -v, --verbose       Enable verbose output with detailed progress information
  --retries N         Number of retry attempts for failed URL fetches (default: 3)
  -q, --quiet         Minimize output (opposite of verbose, but still show errors/warnings)
  --dry-run           Show what would be done without actually doing it
  --no-color          Disable colored output
  -f, --force         Force redownload of files even if they already exist

Examples:
  linkweaver notes.md                      # Process file, skip existing resources
  linkweaver --force notes.md              # Force redownload all resources
  linkweaver --list-links notes.md         # List all URLs found
  linkweaver --xml-cat notes.md            # Concatenate with resources
  linkweaver -x 'llm -t clean' notes.md    # Process with custom command
  linkweaver --retries 5 notes.md          # Use 5 retry attempts for failed URLs
  linkweaver --quiet --dry-run notes.md    # Preview actions quietly
  linkweaver -q -v notes.md                # Quiet mode overrides verbose
  linkweaver --no-color notes.md           # Disable colored output

Common workflows:
  linkweaver --list-links *.md | head -10  # Preview first 10 URLs
160
David Gasquez @davidgasquez.com · 07/07/2025
Now making the agents in the tool chose the final name of the tool using the tool itself. 😜 github.com/davidgasquez...
010
David Gasquez @davidgasquez.com · 05/07/2025
Finally had some time to publish a vibecoded tool I (and Claude Code) built to explore PydanticAI. Replicates a contest where jurors evaluate candidates in pairwise comparisons that get turned into a leaderboard. github.com/davidgasquez...
a contest...

❯ uv run examples/energy.py
2025-07-05 17:15:20 - arbitron.agent - INFO - Initialized agent safety_specialist
2025-07-05 17:15:20 - arbitron.agent - INFO - Initialized agent environmental_specialist
2025-07-05 17:15:20 - arbitron.agent - INFO - Initialized agent economic_specialist
2025-07-05 17:15:20 - arbitron.contest - INFO - Starting competition 'Arbitron Competition' with 5 items and 3 agents
2025-07-05 17:15:20 - arbitron.contest - INFO - Agent safety_specialist will perform 10 comparisons
2025-07-05 17:15:21 - arbitron.agent - INFO - Agent safety_specialist chose nuclear_fission over wind_turbines
2025-07-05 17:15:23 - arbitron.agent - INFO - Agent safety_specialist environmental_specialist chose wind_turbines over nuclear_fission
2025-07-05 17:15:40 - arbitron.agent - INFO - Agent environmental_specialist chose nuclear_fission over hydroelectric_dams
2025-07-05 17:15:41 - arbitron.agent - INFO - Agent environmental_specialist chose geothermal_power over nuclear_fission
2025-07-05 17:15:43 - arbitron.agent - INFO - Agent environmental_specialist chose geothermal_power over hydroelectric_dams
2025-07-05 17:15:44 - arbitron.agent - INFO - Agent environmental_specialist chose solar_photovoltaic over nuclear_fission
2025-07-05 17:15:46 - arbitron.agent - INFO - Agent environmental_specialist chose wind_turbines over geothermal_power
And you get a nice ranking!

1. solar_photovoltaic (score: 1.227)
2. wind_turbines (score: 1.127)
3. nuclear_fission (score: 1.127)
4. hydroelectric_dams (score: 0.801)
5. geothermal_power (score: 0.801)
150
David Gasquez @davidgasquez.com · 04/07/2025
Spent some time going through @cameron.pfiffer.org 's Comind project and @void.comind.network public code and had a lot of fun learning about them! - comind.stream - tangled.sh/@cameron.pfi... Love Comind's components idea! Clean and simple.
Core Components 
Blips: Atomic units of information formatted as ATProto records, such as thoughts, concepts, and emotions
Links: Typed relations between blips with semantic context
Cominds: Specialized AI agents that process network activity
Spheres: Workspaces defined by core directives that focus agent cognition
Melds: Interaction protocol for accessing sphere knowledge
360
David Gasquez @davidgasquez.com · 02/07/2025
Welcome to the dark side! We have the AUR and, eventually, a few interesting system updates! 😅 This is how my current setup (github.com/davidgasquez...) looks like after 9 years of tinkering.
A screenshot with 3 windows. VSCode, btop, and Claude Code
110
David Gasquez @davidgasquez.com · 16/06/2025
I'm also publishing a Parquet file joining all the datasets metadata. Use it from your favorite tool! 💃 sql-workbench.com#queries=v0,S...
020
David Gasquez @davidgasquez.com · 02/06/2025
Alright, this is a job for a "research" template + fragments + chain. Is the `llm prompt` doing the same thing as model.chain() in Python?
import httpx
import llm


def read_url(url: str) -> str:
    """
    Use Jina Reader to convert a URL to Markdown text.

    Example usage:
      llm -f 'reader:https://simonwillison.net/tags/jina/' ...
    """
    url = "https://r.jina.ai/" + url
    response = httpx.get(url)
    if response.status_code != 200:
        raise ValueError(f"Failed to load fragment from {url}: {response.status_code}")
    return response.text


source_note = llm.Fragment(
    content="""
    # Documentation

    - If your product/tool/process documentation is not good enough, people will not use it.
    - Documentation changes should be reviewed in the same way as your code.
    - [If someone's having to read your docs, it's not "simple"](https://justsimply.dev/). Remove filler words.
    - [Principles to keep in mind when writing documentation](https://mkaz.blog/misc/notes-on-technical-writing/).
    - The purpose of technical writing is to help users accomplish tasks as quickly and effectively as possible.
    """
)

model = llm.get_model("gpt-4.1-mini")
chain_response = model.chain(
    """
    Where is "You’re keen to share your excitement" mentioned in the sources?
    """,
    tools=[read_url],
    # after_call=print,
    fragments=[source_note],
    system="You are a helpful research assistant. ALWAYS read ALL the links with `read_url`. Once you've readed all the links in the note, help the user with their questions.",
)
for chunk in chain_response:
    print(chunk, end="", flush=True)
000
David Gasquez @davidgasquez.com · 02/06/2025
Ended up writting a small CLI that you can point to a Markdown note and get an "expanded" view into it. It reads all the links, makes them Markdown and returns all of them under different XML tags for your favorite LLM to consume! github.com/davidgasquez...
❯ lw Emergence.md | llm "Summary on one sentence"
Emergence describes the unpredictable, complex outcomes that arise from simple interactions, which, while not top-down plannable, can be provoked by designing systems with adaptable building blocks, multiple pathways to goals, a low barrier to entry, and multi-agent participation.
160
David Gasquez @davidgasquez.com · 24/05/2025
Wrote a post about some things I've been doing to make my projects more LLM friendly. Spoiler alert: it makes the projects more human friendly too! davidgasquez.com/llm-friendly... Any other ideas or suggestions?
Giving them proper context.
Using simple and specific prompts.
Describing small problems with clarity and boundaries.
Making things easy to run and test (Makefile, Docker, …).
Using well-known frameworks that appear in their training data.
With these basic ideas in mind, let’s see what we can do to make the most out of the current LLMs’ capabilities.Since I’ve been doing Machine Learning projects recently (Kaggle-style competitions), I’ve developed a few extra things I do 1 on those projects.

Document the features you’re currently using and potential features you’d like to explore.
Keep track of open questions and EDA tasks you’d like to explore.
Maintain a DATA_DICTIONARY.md or similar with the terms you’re using. Context is king.
Rely on a functional pipeline style and modular feature engineering. Use “micro pipeline” scripts to generate features. If your dataset has granularity at the “user” level, create different scripts for different kinds of features that generate the feature_name.csv files. You can join them all later. This is one of the things I use a lot. Almost all the features I’ve added to my projects start like “Write a script that computes the user’s […]. Save as feature_name.csv with user, feature_name as headers.”
Track experiments. I’m bad at this one as I haven’t figured out a simple enough way to do this reliably. Ideally, keep track of input features, model parameters, preprocessing, postprocessing, local evaluation, and remote evaluation.
Have a temporary folder (listed in .gitignore) for the LLM to write small scripts and experiments. Here is where you’d have things like inspect_csv.py or other useful scripts that can give the LLMs context on the actual data (e.g., check columns, stats, …)
3142
David Gasquez @davidgasquez.com · 24/05/2025
It is! Found this plugin that does exactly that. github.com/jmdaly/llm-g... Found it with this GitHub query: github.com/search?q=%22... Seems they're adding a custom prompt though!
❯ llm -m github-copilot/claude-sonnet-4 "Hello :)"
Hello! 👋 Great to meet you! I'm GitHub Copilot, and I'm here to help you with any programming questions, code reviews, debugging, or development tasks you might have.

Whether you're working on:
- Writing new code
- Debugging existing code
- Learning a new programming language
- Discussing software architecture
- Code optimization
- Or anything else development-related

I'm ready to assist! What are you working on today?
010
David Gasquez @davidgasquez.com · 20/03/2025
Some INE tables are large! If you want to explore the "Travels, overnight stays and average stay by main features of the trips" dataset, you'd have to download a 37GB CSV! The compressed Parquet file is less than 10MB and can be queried from your browser! 💃 shell.duckdb.org#queries=v0,s...
DuckDB WASM shell.

duckdb> select count(*) from 'https://huggingface.co/datasets/davidgasquez/ine/resolve/main/tablas/15797/datos.parquet';
141
David Gasquez @davidgasquez.com · 18/03/2025
Exactly! Will try as soon as it's available. 😀
Xet is an experimental feature that allows you to store large files efficiently. Xet is currently disabled.
110
David Gasquez @davidgasquez.com · 18/03/2025
Small update! Learned some things about INE's servers. - Don't really support Range Requests. - Don't officially support gzip Accept-Encodings, but they do. Shrinking the download sizes ~8 times. That means we can loop through the compressed datasets and save them as Parquet files on GH Actions!
Github Action showing a DuckDB runThis code.


import duckdb

c = duckdb.connect()
c.sql("""
CREATE SECRET http (TYPE http, EXTRA_HTTP_HEADERS MAP {'Accept-Encoding': 'gzip'});
""")

failed_tables = []
total_tables = len(tablas)

for i, table in enumerate(tablas):
    # Si el directorio no existe, lo creamos
    table_dir = ine_dir / "tablas" / str(table["Id"])
    table_dir.mkdir(exist_ok=True, parents=True)

    # Copiamos el CSV a Parquet
    try:
        c.sql(f"""
            copy (
                from read_csv(
                    'https://www.ine.es/jaxiT3/files/t/en/csv_bdsc/{table["Id"]}.csv',
                    delim=';',
                    ignore_errors=true,
                    normalize_names=true,
                    null_padding=true,
                    parallel=true,
                    strict_mode=false,
                    compression='gzip'
                )
            )
            to 'ine/tablas/{table["Id"]}/datos.parquet' (
                format parquet,
                compression 'zstd',
                parquet_version v2,
                row_group_size 1048576
            );
        """)
    except Exception as e:
        print(f"Error processing table {table['Id']}: {str(e)}")
        failed_tables.append(table["Id"])
130
David Gasquez @davidgasquez.com · 18/03/2025
Not sure why, it opened for a brief second and then went white. The URL still is sql-workbench.com#config={%22p... so it doesn't seems to be truncating anything.
White screen on Brave with the Network tab openWhite screen on Brave with the Console tab open. A couple of errors appear. They are syntax errors.
100
David Gasquez @davidgasquez.com · 21/02/2025
Wrote a post after spending one week using @zed.dev. davidgasquez.com/trying-zed-e... TLDR: Would love to switch once they work on a few extra features (Notebooks, Agent Mode, ...)
2172
David Gasquez @davidgasquez.com · 13/02/2025
What a great definition of Data Engineering! 😅 The only thing I'd tweak is removing the "large amounts of data" requirement. Moving small/medium sized dataset can be as difficult/interesting! ludic.mataroa.blog/blog/brainwa...
Screenshot of a section of the text with the following content hightlighted.

my specialty is building systems that move large amounts of data through companies, organize them in a way that is at least marginally less of a horrific clusterfuck than what random people without specific training will do when left to their own devices, and sometimes assist with statistics
0101
David Gasquez @davidgasquez.com · 30/01/2025
A Python browser sandbox, by the @pydantic.dev folks! github.com/pydantic/pyd... Great to share small snippets of code (e.g: pydantic.run/store/c544b1...).
pydantic.run homepage wiht a simple polars example
2135