Sign in

Paddy Mullen

@paddymullen.bsky.social
96 followers 550 following 118 posts

Boston/Newport. Python/PyData/Jupyter dev. Building the Buckaroo widgetto enhance the DataFrame viewing experience in Jupyter github.com/paddymul/buckaroo

PostsRepliesMedia
Paddy Mullen @paddymullen.bsky.social · 20/02/2026
Ran out of tokens at 430 on Friday. Reset at 6, time to head home
030
Paddy Mullen @paddymullen.bsky.social · 11/12/2025
19 million rows. 450 MB parquet file. Fully interactive in Jupyter. No servers. No subscriptions. Still open source. Just Unix processes doing the heavy lifting in the background while you keep exploring. That's LazyBuckaroo - interactive data exploration on your laptop, powered by Polars.
040
Paddy Mullen @paddymullen.bsky.social · 13/11/2025
If you want to achieve less energy consumption, wouldn’t you want higher energy prices? That’s supply demand 101
010
Paddy Mullen @paddymullen.bsky.social · 15/10/2025
Next fun polars extension I’m going to write - Crash_on_demand. It will crash whenever called or at specified conditions. Why? Well polars and analytical workflows fail occasionally, I’m writing a resilient workflow engine, and I want deterministic crashes for testing.
030
Paddy Mullen @paddymullen.bsky.social · 05/10/2025
@marcogorelli.bsky.social Did excellent work on the rust plugin tutorial. The cookiecutter worked and came with an impressive CI setup that runs against MacOS, Windows, Linux, and multiple python versions.
020
Paddy Mullen @paddymullen.bsky.social · 05/10/2025
I wrote my first rust code and polars extension - pl_series_hash. It runs xx_hash over an entire series to get a single hash u64. Works on most all Polars data types including nested structs, it's fast. WIll be very useful for caching summary stats. github.com/paddymul/pl_...
github.com
GitHub - paddymul/pl_series_hash: a polars plugin to performantly cache series
a polars plugin to performantly cache series. Contribute to paddymul/pl_series_hash development by creating an account on GitHub.
120
Paddy Mullen @paddymullen.bsky.social · 28/08/2025
Providence RI still has a large jewelry making industry. Lots of small shops around there. Also, look up CNC tool dealers like Method Machine tools, give them a call and ask who they would recommend from their customers.
010
Paddy Mullen @paddymullen.bsky.social · 21/08/2025
BTW I'm mortified by the preview image that PyData or youtube chose. It was a live screen recording and I had to tab between multiple windows where buckaroo runs (VSCode, Jupyter, Google Colab, Marimo).
000
Paddy Mullen @paddymullen.bsky.social · 21/08/2025
Looking at the data with Buckaroo - My talk from PyData Boston is now up on youtube. www.youtube.com/watch?v=Htah...
youtube.com
Paddy Mullen - Looking at the Data with Buckaroo (PyData Boston 2025)
YouTube video by PyData
110
Paddy Mullen @paddymullen.bsky.social · 11/08/2025
Getting the script plumbed into consult-mode was a bear. customizing consult-mode requires returning a builder function, that returns another function, that is called by consult. None of the args are documented. #emacs
010
Paddy Mullen @paddymullen.bsky.social · 11/08/2025
This is really useful because I want to find whatever string i'm looking for, but I'd rather find it in my project's source first. github.com/minad/consul...
github.com
How I plumbed in a custom find-grep scripts · minad consult · Discussion #1250
I wanted to write a custom find-grep script that did the following: returned matches from the preferred file extensions (for python -> py,js, ts, jsx, tsx) and an exhaustive exclude list then retur...
100
Paddy Mullen @paddymullen.bsky.social · 11/08/2025
After a bunch of work with prot, I got my custom find-grep script plumbed into emacs. It searches a python project for matches in preferred files first (not site-packages, node_modules, py, js, tx extensions preferred). Then it runs a much more comprehensive search of those excluded directories
100
Paddy Mullen @paddymullen.bsky.social · 11/08/2025
Life pro tip. Next time you have to assemble IKEA furniture, buy yourself a 4mm hex t-handle. and a 4mm wera (or wiha) allen key. So much easier. www.kleintools.com/catalog/t-ha...
kleintools.com
4 mm Hex Key, Journeyman™ T-Handle, 6-Inch - JTH6M4 | Klein Tools
Klein hex-keys are the tools professionals cannot afford to be without. Heat-treated and tempered for superior strength and durability. Klein hex-keys are designed for a precise fit in sockets, preven...
020
Paddy Mullen @paddymullen.bsky.social · 09/08/2025
Nope, some how my new enum_dataframe 3-5xd the memory usage in python. Moral of the story is that polars and parquet have some seriously impressive engineering behind them.
000
Paddy Mullen @paddymullen.bsky.social · 09/08/2025
After some more work, I outputted the entire dataframe, original columns + sparse where necessary. Then I looked at the size, a couple megs more than the original parquet... No big deal, my python memory usage should be less wihtout all of those strings right?
100
Paddy Mullen @paddymullen.bsky.social · 09/08/2025
sparse values. When I changed to an enum per column, the file went down to 50 MB. This was suspiciously low. I realized my new dataframe only included sparse columns, not regular columns + sparse columns.
100
Paddy Mullen @paddymullen.bsky.social · 09/08/2025
UGH, I thought I was so slick. I had a nice function that used polars to convert a csv to parquet with enum columns for sparsely populated columns. I thought it saved about 30% on a 700 meg parquet file (10G csv). Then I dug some more and found out that I was encoding all possible sparse value
100
Paddy Mullen @paddymullen.bsky.social · 01/08/2025
Some content is on Youtube and few other places (heavy equipment, dirt bikes...). But I realized generally I have been searching youtube for more topics because I trust the results more than SEO spam... And this is after I switched to ddg. Obviously there are shills on yt, but beats listicles.
000
Paddy Mullen @paddymullen.bsky.social · 05/07/2025
Link? What is that site
000
Paddy Mullen @paddymullen.bsky.social · 05/07/2025
I need to write more. I have basically no online following. Inhale gotten better distribution through medium. I do try to set my articles to not be behind a paywall
000
Paddy Mullen @paddymullen.bsky.social · 17/06/2025
There is some good stuff on facebook. Fun post from the mainframers group about an IBM Sytem 370 card for the IBM PS/2. www.facebook.com/share/p/17tx...
Some people may ask, "What is a mainframe?".  Well, the answer varies, everything from the multi-ton beasts, to the smaller versions.  How small?  May I present one of the mainframes I used to operate/administer, an IBM P/370.  That's a complete IBM S/370 processor, including a full 16M of memory, on a single MCA card, which fit into a PS/2 model 95 system, and which ran VM/SP 5.  
Once, just for grins and giggles, I IPLed a second level VM system.  Then, for more grins and giggles, I IPLed a third level VM system.  But, it drove me crazy trying to remember which prompt to use for which system.  Still, it demonstrated how thorough of an architecture implementation the card was. 
https://en.wikipedia.org/.../PC-based_IBM_mainframe...
It was great having my own personal mainframe, which I used for product development and testing.  It allowed me to do test installs of software on a real mainframe.  
Anyway, since some of y'all are posting really great pictures of machine rooms and big mainframes, I thought I'd offer a glimpse of what may be one of the smallest mainframes.  🙂
Oh, yeah, there was a follow-on product, the P/390, which was a S/390 system on a card.  I lusted after one for many years, but was never able to justify it.  And, there were a couple of proceeding products, the 7437, which was a full S/370 in a box about the size of a PS/2 model 95.  There was also the XT/370, which was a PC/XT, with a card-set in it which was a partial implementation of the S/370 instruction set (one of which I happen to personally own).I was the team leader for the 7437 project and designed most of the CPU. It used an IBM bit slice (similar to the AMD 2900), and an IBM chip called FLAINE which implemented the floating point instructions. The rest of the CPU was built using mostly 7400F and 7400LS chips. The CPU had a writeable control store of 8K words of 96 bits each. Initially, we implemented dynamic address translation in 7400 logic, but then one of our engineers, Tak Ng, designed a CMOS chip to do the translation, and that went into the product-level release. The chip replaced a whole card worth of 7400 logic. The 7437 consisted of 6 cards, each approximately 9x12". The cards were the Instruction Unit, Execution Unit, Control Store, PS2 Interface and two 8 MB memory cards. Below is a photo of the writeable control store card. It's the only piece of hardware I still have.
The XT/370 was based on a Motorola 68000 that IBM paid Motorola to modify to execute 370 instructions. The chip only could run CMS and not CP as it did not support supervisor state. Instead CP functions were provided by X86 code running on the PC.
The P/370 and the P/390 chips were also designed by Tak Ng.
The 7437, P/370 and P/390 were all done in Bill Beausoleil's IBM Fellow department by a team of 6 to 10 engineers and programmers.
Bill wanted IBM to sell the 7437 for about $5,000. But that would have seriously broken IBM's pricing model vs the cost/performance of traditional mainframes like the 4341 which sold for something like $100,000. VM wanted to charge something like $100K for a 7437 license,
Upper level executives were frightened that customers would buy a slew of 7437s instead of a 3090, costing IBM millions in profit. They swore that they would never let the 7437 go out the door if it meant losing the sale of even one 3090. In the end, they severely restricted who could buy a 7437 by bundling it with a 5080 graphics workstation and pricing it at $50,000.
000
Paddy Mullen @paddymullen.bsky.social · 13/06/2025
On the toughest most annoying problems I need to try to put in a half hour to an hour a day, and let ideas percolate. But actually do that work, it's much more effective than hours of grinding. I'm thinking about devops stuff especially
000
Paddy Mullen @paddymullen.bsky.social · 13/06/2025
You hear those stories about how garbage collection on lisp machines in the 80s didn't really work and engineers just restarted the machine once a day. Then I realized I do the same thing with firefox/chrome. Need to upgrade to a 32+GB laptop
000
Paddy Mullen @paddymullen.bsky.social · 12/06/2025
I want larger wheels because I think they’ll help with surface roughness a lot. Something about rollerblading bothers my knees like skiing, and ice skating don’t.
000
Paddy Mullen @paddymullen.bsky.social · 11/06/2025
There are constructions that are very useful that I know I would have avoided because they make the type checkers life difficult. This is a very bad habit to get into. I don’t know how to generally fix it. Getting more comfortable with typing will help.
000
Paddy Mullen @paddymullen.bsky.social · 11/06/2025
I think the typing helps, a bit, but I really worry about the code that I would write if I started writing typed python. There are things I’m doing which are complex to type. The urge and incentives are to make it typed first and useful second.
100
Paddy Mullen @paddymullen.bsky.social · 11/06/2025
It’s more a matter of seeing the syntax highlighting and wanting to fix it then wanting my code typed.
100
Paddy Mullen @paddymullen.bsky.social · 11/06/2025
I have started heavily typing my python code. I’m late to the party on this one, I know, but here are some thoughts. I was prompted by finally getting eglot mode working with basedpyright in eMacs.
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
All of this will be part of my article "so you want to serialize a dataframe to JS". Which is a subsection (probably the largest) of "so you want to write a table viewer"
000
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
This should be much more modular.
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
The final cool idea though is to leverage multi indexes and polars structs. It would work like this. Say you want to color a column based on the diff with an original column. Put them into a struct, then render, coloring based on the struct column you don't display.
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
I am currently kicking around two ideas. no, maybe 3, writing helps. First would be a conversion step in python, plumbing this in is a mess. Second is different renderers, that do the same thing in the frontend. This is acutally a bit cleaner
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
But how to deal with custom coded apps that build color_maps via column comparison. By the time this gets to the frontend, those column names won't exist in the data.
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
The other key change is converting every column name from the original to a letter based encoding sequential encoding (a-z, aa, ab,). and then having each config for columns explicitly state header_name and field.
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
Rathrer than specialcasing pandas parquet. I am doing the following. Explicit `first_col_config` for table configuration. no more including "index" in summary stats (I don't currently have a way to display it either.
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
This works until you want a column named "index", or you serialize a dataframe to parquet with multi-index columns. Parquet changes "index" to the string "('index', '')".
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
for JSON, buckaroo serializes dataframes as a list of dicts, `to_json(orient='records')` or `polars.with_row_index()`. This always adds a column of "index" to the records. Parquet behaves similarly, but column oriented.
100
Paddy Mullen @paddymullen.bsky.social · 09/06/2025
Hit a snag with the multi-index refactoring. The short of it is that dataframe serialization is really tricky and relying on assumption and default behavors will bite you, always
100
Paddy Mullen @paddymullen.bsky.social · 08/06/2025
It works to arbitrary levels of nesting too
Buckaroo rendering a dataframes with a 3 level column MultiIndex
000
Paddy Mullen @paddymullen.bsky.social · 08/06/2025
And I now have pandas MultiIndex columns rendering nested locally.
Image showing buckaroo in marimo rendering a dataframe with multi index columns.  the headers are nested
100
Paddy Mullen @paddymullen.bsky.social · 08/06/2025
Bahhhh. Traitlets doesn't like typing. I'm trying to add types to my project, and I have written most of them, but my project heavily uses traits, and they don't get along with python typing
000
Paddy Mullen @paddymullen.bsky.social · 07/06/2025
Yep, lot's of advice for microshift from my reddit thread about this.
010
Paddy Mullen @paddymullen.bsky.social · 07/06/2025
All of this reminds me to finish and publish my article "So you want to write a dataframe table". It details most of the little pitfalls that I remember in building Buckaroo. it's 2k words right now
000
Paddy Mullen @paddymullen.bsky.social · 07/06/2025
Finally I remembered while looking this up that pandas can have multi-indexes for rows too. I can do that. www.ag-grid.com/javascript-d... ... but down the line. The next major bit of assumption jank to pull out of Buckaroo after multi-indexes is pandas indexes. Lot's of implicit stuff around it.
ag-grid.com
JavaScript Grid: Row Spanning | AG Grid
A single cell can be used to represent multiple contiguous leaf rows with equal values. Download AG Grid v33.3.2 today: The best JavaScript Table & JavaScript Data Grid in the world.
100
Paddy Mullen @paddymullen.bsky.social · 07/06/2025
Also, this should work really well with polars tuple types.
100
Paddy Mullen @paddymullen.bsky.social · 07/06/2025
I needed to use multi indexes to deal with extracting data from panderas validations last week. The UX of passing a dataframe to buckaroo and seeing only a stacktrace was really bad. A little reading of AG-Grid docs and here I am.
100
Paddy Mullen @paddymullen.bsky.social · 07/06/2025
I'm kneedeep in adding proper multi index-column support to Buckaroo. MultiIndexes are kind of a corner case of pandas, super powerful but often forgotten. Many tables don't display them properly because they are tricky. They don't serialize to JSON natively. #python #pandas #datascience #pydata
Mockup of multi indexes in storybook for buckaroo.  This shows the nested column headers
110
Paddy Mullen @paddymullen.bsky.social · 05/06/2025
So, apparently SRAM NX 1x12 isn't a good fit on a 20 inch wheel. The RD hits the tire in the 3rd largest gear. #cargobike #bike
Image showing limited clearance between the derailleur and tire. Here the derailleur is in the 3rd or 4th largest gear Clearance between the derailleur and the ground when the derailleur is at it's most downward extended position.Riding the cargo bike on grass.  Note the 20 inch rear wheel and 26 inch front wheel.  A coastal scene is in the background.
211
Paddy Mullen @paddymullen.bsky.social · 01/06/2025
I have been playing with beartype. It's a library for fast runtime type verification in python. The maintainer is one of the most entertaining writers I have seen. Period. His response to my feature request/(trying to understand his library) has brought so much joy. github.com/beartype/bea...
github.com
[Feature Request] New `BeartypeInferHintConf(infer_hint_dict_kind)` configuration option for fine-grained dictionary type hint inference 🧘 · Issue #529 · beartype/beartype
Is it within the scope of beartype to do the following: def takes_a_dict(foo): temp = foo['a'] + len(foo['some_str']) b = "asdf" return dict(temp=temp, b=b) infer_hint(takes_a_dict) --- collections...
000
Paddy Mullen @paddymullen.bsky.social · 20/05/2025
Also everytime I open it on osx, it starts updating, I tab over to a different app. It finishes updating, restarts and grabs focus again. Does a longer update cycle, finally it’s ready. Bring back VBulletin
000