Paddy Mullen @paddymullen.bsky.social · 20/02/2026Ran out of tokens at 430 on Friday. Reset at 6, time to head home 030
Paddy Mullen @paddymullen.bsky.social · 11/12/202519 million rows. 450 MB parquet file. Fully interactive in Jupyter. No servers. No subscriptions. Still open source. Just Unix processes doing the heavy lifting in the background while you keep exploring. That's LazyBuckaroo - interactive data exploration on your laptop, powered by Polars. 040
Paddy Mullen @paddymullen.bsky.social · 13/11/2025If you want to achieve less energy consumption, wouldn’t you want higher energy prices? That’s supply demand 101 010
Paddy Mullen @paddymullen.bsky.social · 15/10/2025Next fun polars extension I’m going to write - Crash_on_demand. It will crash whenever called or at specified conditions. Why? Well polars and analytical workflows fail occasionally, I’m writing a resilient workflow engine, and I want deterministic crashes for testing. 030
Paddy Mullen @paddymullen.bsky.social · 05/10/2025@marcogorelli.bsky.social Did excellent work on the rust plugin tutorial. The cookiecutter worked and came with an impressive CI setup that runs against MacOS, Windows, Linux, and multiple python versions. 020
Paddy Mullen @paddymullen.bsky.social · 05/10/2025I wrote my first rust code and polars extension - pl_series_hash. It runs xx_hash over an entire series to get a single hash u64. Works on most all Polars data types including nested structs, it's fast. WIll be very useful for caching summary stats. github.com/paddymul/pl_...github.comGitHub - paddymul/pl_series_hash: a polars plugin to performantly cache seriesa polars plugin to performantly cache series. Contribute to paddymul/pl_series_hash development by creating an account on GitHub. 120
Paddy Mullen @paddymullen.bsky.social · 28/08/2025Providence RI still has a large jewelry making industry. Lots of small shops around there. Also, look up CNC tool dealers like Method Machine tools, give them a call and ask who they would recommend from their customers. 010
Paddy Mullen @paddymullen.bsky.social · 21/08/2025BTW I'm mortified by the preview image that PyData or youtube chose. It was a live screen recording and I had to tab between multiple windows where buckaroo runs (VSCode, Jupyter, Google Colab, Marimo). 000
Paddy Mullen @paddymullen.bsky.social · 21/08/2025Looking at the data with Buckaroo - My talk from PyData Boston is now up on youtube. www.youtube.com/watch?v=Htah...youtube.comPaddy Mullen - Looking at the Data with Buckaroo (PyData Boston 2025)YouTube video by PyData 110
Paddy Mullen @paddymullen.bsky.social · 11/08/2025Getting the script plumbed into consult-mode was a bear. customizing consult-mode requires returning a builder function, that returns another function, that is called by consult. None of the args are documented. #emacs 010
Paddy Mullen @paddymullen.bsky.social · 11/08/2025This is really useful because I want to find whatever string i'm looking for, but I'd rather find it in my project's source first. github.com/minad/consul...github.comHow I plumbed in a custom find-grep scripts · minad consult · Discussion #1250I wanted to write a custom find-grep script that did the following: returned matches from the preferred file extensions (for python -> py,js, ts, jsx, tsx) and an exhaustive exclude list then retur... 100
Paddy Mullen @paddymullen.bsky.social · 11/08/2025After a bunch of work with prot, I got my custom find-grep script plumbed into emacs. It searches a python project for matches in preferred files first (not site-packages, node_modules, py, js, tx extensions preferred). Then it runs a much more comprehensive search of those excluded directories 100
Paddy Mullen @paddymullen.bsky.social · 11/08/2025Life pro tip. Next time you have to assemble IKEA furniture, buy yourself a 4mm hex t-handle. and a 4mm wera (or wiha) allen key. So much easier. www.kleintools.com/catalog/t-ha...kleintools.com4 mm Hex Key, Journeyman™ T-Handle, 6-Inch - JTH6M4 | Klein ToolsKlein hex-keys are the tools professionals cannot afford to be without. Heat-treated and tempered for superior strength and durability. Klein hex-keys are designed for a precise fit in sockets, preven... 020
Paddy Mullen @paddymullen.bsky.social · 09/08/2025Nope, some how my new enum_dataframe 3-5xd the memory usage in python. Moral of the story is that polars and parquet have some seriously impressive engineering behind them. 000
Paddy Mullen @paddymullen.bsky.social · 09/08/2025After some more work, I outputted the entire dataframe, original columns + sparse where necessary. Then I looked at the size, a couple megs more than the original parquet... No big deal, my python memory usage should be less wihtout all of those strings right? 100
Paddy Mullen @paddymullen.bsky.social · 09/08/2025sparse values. When I changed to an enum per column, the file went down to 50 MB. This was suspiciously low. I realized my new dataframe only included sparse columns, not regular columns + sparse columns. 100
Paddy Mullen @paddymullen.bsky.social · 09/08/2025UGH, I thought I was so slick. I had a nice function that used polars to convert a csv to parquet with enum columns for sparsely populated columns. I thought it saved about 30% on a 700 meg parquet file (10G csv). Then I dug some more and found out that I was encoding all possible sparse value 100
Paddy Mullen @paddymullen.bsky.social · 01/08/2025Some content is on Youtube and few other places (heavy equipment, dirt bikes...). But I realized generally I have been searching youtube for more topics because I trust the results more than SEO spam... And this is after I switched to ddg. Obviously there are shills on yt, but beats listicles. 000
Paddy Mullen @paddymullen.bsky.social · 05/07/2025I need to write more. I have basically no online following. Inhale gotten better distribution through medium. I do try to set my articles to not be behind a paywall 000
Paddy Mullen @paddymullen.bsky.social · 17/06/2025There is some good stuff on facebook. Fun post from the mainframers group about an IBM Sytem 370 card for the IBM PS/2. www.facebook.com/share/p/17tx... 000
Paddy Mullen @paddymullen.bsky.social · 13/06/2025On the toughest most annoying problems I need to try to put in a half hour to an hour a day, and let ideas percolate. But actually do that work, it's much more effective than hours of grinding. I'm thinking about devops stuff especially 000
Paddy Mullen @paddymullen.bsky.social · 13/06/2025You hear those stories about how garbage collection on lisp machines in the 80s didn't really work and engineers just restarted the machine once a day. Then I realized I do the same thing with firefox/chrome. Need to upgrade to a 32+GB laptop 000
Paddy Mullen @paddymullen.bsky.social · 12/06/2025I want larger wheels because I think they’ll help with surface roughness a lot. Something about rollerblading bothers my knees like skiing, and ice skating don’t. 000
Paddy Mullen @paddymullen.bsky.social · 11/06/2025There are constructions that are very useful that I know I would have avoided because they make the type checkers life difficult. This is a very bad habit to get into. I don’t know how to generally fix it. Getting more comfortable with typing will help. 000
Paddy Mullen @paddymullen.bsky.social · 11/06/2025I think the typing helps, a bit, but I really worry about the code that I would write if I started writing typed python. There are things I’m doing which are complex to type. The urge and incentives are to make it typed first and useful second. 100
Paddy Mullen @paddymullen.bsky.social · 11/06/2025It’s more a matter of seeing the syntax highlighting and wanting to fix it then wanting my code typed. 100
Paddy Mullen @paddymullen.bsky.social · 11/06/2025I have started heavily typing my python code. I’m late to the party on this one, I know, but here are some thoughts. I was prompted by finally getting eglot mode working with basedpyright in eMacs. 100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025All of this will be part of my article "so you want to serialize a dataframe to JS". Which is a subsection (probably the largest) of "so you want to write a table viewer" 000
Paddy Mullen @paddymullen.bsky.social · 09/06/2025The final cool idea though is to leverage multi indexes and polars structs. It would work like this. Say you want to color a column based on the diff with an original column. Put them into a struct, then render, coloring based on the struct column you don't display. 100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025I am currently kicking around two ideas. no, maybe 3, writing helps. First would be a conversion step in python, plumbing this in is a mess. Second is different renderers, that do the same thing in the frontend. This is acutally a bit cleaner 100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025But how to deal with custom coded apps that build color_maps via column comparison. By the time this gets to the frontend, those column names won't exist in the data. 100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025The other key change is converting every column name from the original to a letter based encoding sequential encoding (a-z, aa, ab,). and then having each config for columns explicitly state header_name and field. 100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025Rathrer than specialcasing pandas parquet. I am doing the following. Explicit `first_col_config` for table configuration. no more including "index" in summary stats (I don't currently have a way to display it either. 100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025This works until you want a column named "index", or you serialize a dataframe to parquet with multi-index columns. Parquet changes "index" to the string "('index', '')". 100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025for JSON, buckaroo serializes dataframes as a list of dicts, `to_json(orient='records')` or `polars.with_row_index()`. This always adds a column of "index" to the records. Parquet behaves similarly, but column oriented. 100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025Hit a snag with the multi-index refactoring. The short of it is that dataframe serialization is really tricky and relying on assumption and default behavors will bite you, always 100
Paddy Mullen @paddymullen.bsky.social · 08/06/2025And I now have pandas MultiIndex columns rendering nested locally. 100
Paddy Mullen @paddymullen.bsky.social · 08/06/2025Bahhhh. Traitlets doesn't like typing. I'm trying to add types to my project, and I have written most of them, but my project heavily uses traits, and they don't get along with python typing 000
Paddy Mullen @paddymullen.bsky.social · 07/06/2025Yep, lot's of advice for microshift from my reddit thread about this. 010
Paddy Mullen @paddymullen.bsky.social · 07/06/2025All of this reminds me to finish and publish my article "So you want to write a dataframe table". It details most of the little pitfalls that I remember in building Buckaroo. it's 2k words right now 000
Paddy Mullen @paddymullen.bsky.social · 07/06/2025Finally I remembered while looking this up that pandas can have multi-indexes for rows too. I can do that. www.ag-grid.com/javascript-d... ... but down the line. The next major bit of assumption jank to pull out of Buckaroo after multi-indexes is pandas indexes. Lot's of implicit stuff around it.ag-grid.comJavaScript Grid: Row Spanning | AG GridA single cell can be used to represent multiple contiguous leaf rows with equal values. Download AG Grid v33.3.2 today: The best JavaScript Table & JavaScript Data Grid in the world. 100
Paddy Mullen @paddymullen.bsky.social · 07/06/2025Also, this should work really well with polars tuple types. 100
Paddy Mullen @paddymullen.bsky.social · 07/06/2025I needed to use multi indexes to deal with extracting data from panderas validations last week. The UX of passing a dataframe to buckaroo and seeing only a stacktrace was really bad. A little reading of AG-Grid docs and here I am. 100
Paddy Mullen @paddymullen.bsky.social · 07/06/2025I'm kneedeep in adding proper multi index-column support to Buckaroo. MultiIndexes are kind of a corner case of pandas, super powerful but often forgotten. Many tables don't display them properly because they are tricky. They don't serialize to JSON natively. #python #pandas #datascience #pydata 110
Paddy Mullen @paddymullen.bsky.social · 05/06/2025So, apparently SRAM NX 1x12 isn't a good fit on a 20 inch wheel. The RD hits the tire in the 3rd largest gear. #cargobike #bike 211
Paddy Mullen @paddymullen.bsky.social · 01/06/2025I have been playing with beartype. It's a library for fast runtime type verification in python. The maintainer is one of the most entertaining writers I have seen. Period. His response to my feature request/(trying to understand his library) has brought so much joy. github.com/beartype/bea...github.com[Feature Request] New `BeartypeInferHintConf(infer_hint_dict_kind)` configuration option for fine-grained dictionary type hint inference 🧘 · Issue #529 · beartype/beartypeIs it within the scope of beartype to do the following: def takes_a_dict(foo): temp = foo['a'] + len(foo['some_str']) b = "asdf" return dict(temp=temp, b=b) infer_hint(takes_a_dict) --- collections... 000
Paddy Mullen @paddymullen.bsky.social · 20/05/2025Also everytime I open it on osx, it starts updating, I tab over to a different app. It finishes updating, restarts and grabs focus again. Does a longer update cycle, finally it’s ready. Bring back VBulletin 000