Sign in

Pete Bachant

@petebachant.me
311 followers 991 following 240 posts

RSE @caltech.edu Bicycles, fluid dynamics, Python, open source, open science, reproducibility. petebachant.me | calkit.org

PostsRepliesMedia
Pete Bachant @petebachant.me · 30/09/2026
Calkit 0.47.11 greatly improves math handling in LaTeX --> Word --> LaTeX round-tripping, including support for LibreOffice. github.com/calkit/calki... #opensource #openscience #reproducibility
A screenshot showing Calkit's LaTeX to Word to LaTeX conversion results.
031
Pete Bachant @petebachant.me · 29/09/2026
Writing code, like manufacturing parts for machines, is technician work, not engineering. Engineering is simply predicting if something is going to work. Both can be done by the same person, but we shouldn't confuse the two.
110
Pete Bachant @petebachant.me · 24/09/2026
How to prevent this problem? Instrumentation stamping artifacts with cryptographically signed attestation for provenance? That could be useful for more than just fraud detection. #metascience
130
Pete Bachant @petebachant.me · 10/09/2026
Maybe problems solved with AI should be attributed to every author whose work is present in the training data, not the company that trained the model.
000
Pete Bachant @petebachant.me · 09/09/2026
Emailing Word documents back and forth, despite its many drawbacks, is still one path of least resistance for collaborative writing.
100
Pete Bachant @petebachant.me · 05/09/2026
How do you force an AI agent to work reproducibly? Trying to figure that out myself as I have Claude experiment with CUDA kernel optimization. Not perfect, but having it work with a pipeline with definitive staleness detection is helping: calkit.io/petebachant/... #agenticai #reproducibility
calkit.io
Calkit
010
Pete Bachant @petebachant.me · 29/08/2026
Many scientific papers are tightly-coupled to the code, data, and environments that produced them, and therefore the paper itself fails to provide much standalone value. Failure to reproduce or replicate is a direct consequence of this coupling. #openscience #replicationcrisis #metascience
110
Pete Bachant @petebachant.me · 28/08/2026
Calkit now has online figure and notebook editors that run in the browser with WASM/Pyodide, but produce code and environments integrated into a local-first reproducible pipeline. #opensource #openscience #reproducibility
The Calkit figure editor.
120
Pete Bachant @petebachant.me · 18/08/2026
Thinking of revamping the Calkit onboarding a bit. What use case is more valuable to you: 1. You have a research project in progress that's a bit of a disaster and you'd like to get it under control 2. You're starting a new project and want to follow best practices from day one #reproducibility
001
Pete Bachant @petebachant.me · 14/08/2026
"The version of the plotting library that the authors used contained a bug." This is why data, code, and **environment lock files** should be shipped with all papers. #reproducibility #openscience #metascience phys.org/news/2026-08...
phys.org
Claim that science is becoming less disruptive challenged after years-long publication delay
A widely cited Nature cover article from 2023 claimed that scientific innovation has significantly declined in recent decades. This finding was based on a citation analysis of large sets of metadata c...
092
Pete Bachant @petebachant.me · 13/08/2026
We are definitely going to need some social norms around asking people to read things generated with AI. It at least needs to be explicit, and it probably should be okay to respond with: can you explain in your own words?
100
Pete Bachant @petebachant.me · 11/08/2026
Calkit Chrome extension just dropped. Collect references and notes directly into a BibTeX file in a GitHub repo from journal websites and arXiv, sync figures and results into Overleaf, and review DVC-tracked files and LaTeX diffs directly on GitHub. #opensource #openscience #reproducibility
Calkit Chrome extension in arXiv.Calkit Chrome extension showing a LaTeX diff on GitHub.Calkit Chrome extension on Overleaf.Calkit Chrome extension on a GitHub PR.
130
Pete Bachant @petebachant.me · 04/08/2026
Doing open science gets you more citations (arxiv.org/pdf/2607.00546) However, does it also result in less individualistic innovation (less notoriety, more scooping, worse tenure outcomes)? #openscience #publishorperish
arxiv.org
110
Pete Bachant @petebachant.me · 04/08/2026
Leaders are responsible for vision and strategy, i.e., prioritizing problems and giving insights into avenues for solutions. If you're handing off fully-baked solution specs to your engineers you're wasting their talent and holding them back from growing.
120
Pete Bachant @petebachant.me · 03/08/2026
There's an interesting distinction between "this can run again" and "I can prove that I ran this and here's the output" and the latter is becoming more important with increasing AI usage #agenticai #aiforscience #reproducibility
020
Pete Bachant @petebachant.me · 01/08/2026
How many more times can I hear "agentic AI" this week before I quit using computers for the rest of my life challenge
010
Pete Bachant @petebachant.me · 29/07/2026
AI agents, like humans, should be creating pipelines. They should not be parts of pipelines.
010
Pete Bachant @petebachant.me · 28/07/2026
Moving steadily towards a free, open-source, vertically-integrated research platform, calkit.io now has a project-based reference manager that can sync bidirectionally with Zotero. #openscience #opensource #reproducibility
A screenshot of the Calkit reference manager.
020
Pete Bachant @petebachant.me · 27/07/2026
New feature on calkit.io: attach evidence to answers to research questions and quickly verify that they're still up-to-date given their provenance. Example project: calkit.io/calkit/examp... #reproducibility #openscience #opensource
A screenshot of Calkit's questions section.
041
Pete Bachant @petebachant.me · 26/07/2026
Unifying development and operations in the same team/repository was a boon for software. Would science see similar gains from moving its various stages/functions closer together? #devops #openscience #opendata
020
Pete Bachant @petebachant.me · 18/07/2026
"utils" modules are bad because they show you value DRY (don't repeat yourself) more than modularity, and you end up with tight coupling and low cohesion #softwareengineering #swe #softwarearchitecture
110
Pete Bachant @petebachant.me · 17/07/2026
There are lots of great computational tools out there for doing different stages of research, but they're poorly integrated with each other, which is a source of errors and inefficiency.
110
Pete Bachant @petebachant.me · 14/07/2026
How well does failure to reproduce correlate with failure to replicate? #metascience #reproducibility
100
Pete Bachant @petebachant.me · 11/07/2026
If your collaborators don't want to learn Git/GitHub but you don't want to silo paper writing away from your code and data, calkit.io now has a LaTeX editor and WASM compiler :) #opensource #openscience #reproducibility
A screenshot of Calkit's in-browser LaTeX editor.
040
Pete Bachant @petebachant.me · 08/07/2026
If you're building a writing tool for scientists, it's pretty important to make it easy to connect to visualization/analysis tools. That is, if you make users manually upload their data or figures, that's bad.
000
Reposted by Pete Bachant
Konrad Hinsen @khinsen.net · 06/07/2026
New blog post: "Conviviality in computational science" blog.khinsen.net/posts/2026/0... "Conviviality matters for science ... if you want to derive knowledge from your work, you need to know exactly what you are doing, and that includes a detailed understanding of your tools." 🧪 #metascience
blog.khinsen.net
Konrad Hinsen's blog
34416
Pete Bachant @petebachant.me · 02/07/2026
Instead of compiling a replication package or repro pack at journal article submission time, imagine if you worked inside of a repro pack the whole project... #reproducibility #openscience
010
Pete Bachant @petebachant.me · 01/07/2026
@plos.org's concept of the "knowledge stack" in science is interesting from an architectural perspective. Is any layer truly valuable on its own? I'd argue the full stack is the most valuable unit, which means you should be shipping the whole thing with every study. #openscience #opensource
PLOS's knowledge stack diagram.
010
Pete Bachant @petebachant.me · 29/06/2026
If you know what individual to ask for help in a separate team, you are probably tightly-coupled, and maybe should be part of the same team
000
Pete Bachant @petebachant.me · 28/06/2026
The Calkit VS Code extension now allows viewing interactive @plotly.com figures, and includes a gallery view.
000
Pete Bachant @petebachant.me · 24/06/2026
If a human is supposed to read it, it shouldn't be written by AI
120
Pete Bachant @petebachant.me · 23/06/2026
For those who manage research groups: Do you tend to be more individual-focused, with each member getting their own independent project, or team-focused, where individuals need to find their niche to contribute to a larger collective goal?
010
Pete Bachant @petebachant.me · 15/06/2026
Been spending some time on the Calkit VS Code extension. Here's a screenshot using it with Claude to profile CUDA kernels in SLURM jobs and keep track of Nsight reports for each code change: #aiagents #cuda #gpu #claude
A screenshot of the Calkit VS Code extension.
130
Pete Bachant @petebachant.me · 07/06/2026
AI agents will happily create figures for you. They'll even export some source code. But can you trust that they did what they said they did? Here's a tutorial on how to get your AI agents to work reproducibly: docs.calkit.org/tutorials/ai...
docs.calkit.org
Declarative, structured prompting to make AI agents work reproducibly - Calkit
030
Pete Bachant @petebachant.me · 02/06/2026
I'm not a fan of splitting up code into a bunch of tiny functions as a default practice. It's like creating a bunch of minor characters in a story. It should only be done if those characters really add something, because each one will take away some attention. #swe #softwareengineering #dev
000
Pete Bachant @petebachant.me · 29/05/2026
Made a little poster to "sell" my free software #opensource #openscience #reproducibility
030
Pete Bachant @petebachant.me · 26/05/2026
The golden rule of architecture: Avoid moving horizontal slices away from each other, either into different repos or different teams
000
Pete Bachant @petebachant.me · 05/05/2026
Maybe we can build more human interfaces for our computing tools without making them nondeterministic and extremely computationally expensive
010
Pete Bachant @petebachant.me · 02/05/2026
On one side, you have literate programming, which views code and prose as belonging to a monolithic artifact. On the other, you have total fragmentation with code in one silo, data in another, writing in another, and no real interface between them. I think we need modularity with real interfaces.
020
Pete Bachant @petebachant.me · 01/05/2026
Tightly coupled code should not be split across multiple repos #softwareengineering #swe #softwaredesign
010
Pete Bachant @petebachant.me · 30/04/2026
IMO, the most important rule for using AI agents to do scientific research: Don't allow them to create artifacts like figures or numerical results on their own. Have them create, save, and run pipelines that create artifacts so provenance is preserved and traceable. #openscience #agentic
docs.calkit.org
Use with AI tools - Calkit
020
Pete Bachant @petebachant.me · 29/04/2026
Success boils down to properly defining how to measure it then doing as many iterations as possible.
000
Pete Bachant @petebachant.me · 20/04/2026
New feature on calkit.io: Comment on publication PDFs and create GitHub issues from them. This was inspired by the process I used to get my PhD thesis done on time. Every comment from my advisor became a GitHub issue and I calculated the necessary daily "velocity" working backward from the deadline.
010
Pete Bachant @petebachant.me · 08/04/2026
Super helpful, Google. Thanks. #nodejs
000
Pete Bachant @petebachant.me · 06/04/2026
This may be extremely niche, but if you need to run a Jupyter kernel in a SLURM job, e.g., to reserve a GPU, and connect it to a notebook in VS Code, here's a solution: docs.calkit.org/tutorials/vs... #opensource #jupyter #hpc #cuda
docs.calkit.org
Connect a Jupyter Notebook to a kernel in a SLURM environment in VS Code - Calkit
040
Pete Bachant @petebachant.me · 31/03/2026
The best time to automate your project is 10 months ago when you started. The second best time is now.
010
Pete Bachant @petebachant.me · 30/03/2026
One of the limitations of DVC is that it's relatively slow/inefficient working with large folders consisting of many small files. Here's a solution: docs.calkit.org/version-cont... #dataversioncontrol #datascience #python
docs.calkit.org
Version control - Calkit
020
Pete Bachant @petebachant.me · 18/03/2026
Is there a parallel to Conway's Law for cognitive load, i.e., that software complexity will increase to hit the cognitive load limit of the team that owns it regardless of the inherent complexity of the problem? #softwareengineering
110
Reposted by Pete Bachant
PLOS Biology @plosbiology.org · 27/02/2026
In support of #OpenScience, we routinely ask authors to openly share their #research #code before publication. We are now formalizing this practice with a mandatory #CodeSharing policy and clarifying what we mean by code sharing.
plos.io
Formalizing our commitment to code sharing
In support of open science, PLOS Biology routinely asks authors to openly share their research code before publication. We are now formalizing this practice with a mandatory code sharing policy and…
05125
Pete Bachant @petebachant.me · 27/02/2026
Quick demo of something I've been working on: Execute some Python, R, Julia, MATLAB, and LaTeX scripts/notebooks/docs and they automatically become a DVC pipeline (DAG). #opensource #reproducibility #datascience
020