Sign in

taupirho.bsky.social

@taupirho.bsky.social
31 followers 21 following 188 posts
PostsRepliesMedia
taupirho.bsky.social @taupirho.bsky.social · 02/10/2026
I Compacted 1,000 Apache Iceberg Files Into 6. Here’s What Happened to Query Performance. Read my latest @towardsdatascience.com article for free. track.pstmrk.it/3s/towardsda...
track.pstmrk.it
I Compacted 1,000 Apache Iceberg Files Into 6. Here’s What Happened to Query Performance.
Benchmarking the impact of fewer, larger files across three SQL workloads
000
taupirho.bsky.social @taupirho.bsky.social · 02/10/2026
Use a browser-local alternative when you only need to prepare and download a protected PDF. Preparation and pre-delivery checks: securepdf.io/resources/pa... #PDFSecurity
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 25/09/2026
One of the most in-demand data engineering skills right now is a tool called dbt. dbt lets you transform data in your warehouse using SQL, while managing dependencies and testing the results. Check out my article on @towardsdatascience.com using the link below, track.pstmrk.it/3s/towardsda...
track.pstmrk.it
Getting started with dbt
A practical guide to building, testing, and documenting SQL transformations
000
taupirho.bsky.social @taupirho.bsky.social · 25/09/2026
Choose an appropriate delivery method and avoid sending the document and password together. Preparation and pre-delivery checks: securepdf.io/resources/se... #PDFSecurity
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 22/09/2026
My latest article on the new JEV LLM - currently taking the AI world by storm - is now on the @towardsdatascience.com platform. Read it for free here: track.pstmrk.it/3s/towardsda...
track.pstmrk.it
An Introduction to Jev
The AI that makes decisions instead of generating text
000
taupirho.bsky.social @taupirho.bsky.social · 22/09/2026
Jev makes AI feel more like a programming primitive: ask for a choice, score or yes/no judgment and get structured results with probabilities. Your code controls the workflow. AI supplies the judgment. Read (for free) why this approach interests me: track.pstmrk.it/3s/towardsda...
track.pstmrk.it
An Introduction to Jev
The AI that makes decisions instead of generating text
000
taupirho.bsky.social @taupirho.bsky.social · 19/09/2026
Combine a visible recipient or confidentiality mark with PDF password protection. Preparation and pre-delivery checks: securepdf.io/resources/ad... #PDFSecurity
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 17/09/2026
From a Parquet file to a working lakehouse with DuckDB + DuckLake. My latest hands-on guide on @towardsdatascience.com covers SQL updates, schema evolution, time travel and joining local data with Amazon S3. Build it step by step: track.pstmrk.it/3s/towardsda...
track.pstmrk.it
Building a Data Lakehouse with DuckDB and DuckLake | Towards Data Science
Starting with a local Parquet file, then joining it to data stored in the cloud
020
taupirho.bsky.social @taupirho.bsky.social · 11/09/2026
A repeatable check for protecting and delivering confidential client-report PDFs. Preparation and pre-delivery checks: securepdf.io/resources/pa... #PDFSecurity
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 10/09/2026
My latest post just dropped on @towardsdatascience.com . It's an introduction to the dbt tool - how to get it and how to use it with @duckdb as the target DB. Read for free using the link below. track.pstmrk.it/3s/towardsda...
track.pstmrk.it
Getting started with dbt | Towards Data Science
A practical guide to building, testing, and documenting SQL transformations
000
taupirho.bsky.social @taupirho.bsky.social · 04/09/2026
pip install pyspark can succeed while PySpark still fails. I wrote a tested setup guide for Windows, WSL and Ubuntu. It covers VS Code, Java gateway errors and a real Parquet write test. profitabledonkeys.com/guides/pyspa...
profitabledonkeys.com
Install PySpark locally on Windows, VS Code or Ubuntu
Install PySpark locally on Windows, WSL or Ubuntu, select the right Python environment in VS Code, and test a real Spark session and Parquet write.
000
taupirho.bsky.social @taupirho.bsky.social · 04/09/2026
Understand the difference between selecting a local file and uploading it to a document service. Preparation and pre-delivery checks: securepdf.io/resources/lo... #PDFSecurity
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 02/09/2026
groupBy() reduces each group to one row. PySpark window functions calculate rankings, running totals, moving averages and previous values while keeping every original record. I explain how they work in my latest @towardsdatascience.com article: track.pstmrk.it/3s/towardsda...
track.pstmrk.it
A Practical Introduction to PySpark Window Functions | Towards Data Science
Why the standard groupBy function isn’t enough
000
taupirho.bsky.social @taupirho.bsky.social · 28/08/2026
Add PDF password protection and send the password through a separate channel. Preparation and pre-delivery checks: securepdf.io/resources/pa... #PDFSecurity
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 25/08/2026
SQL can solve more graph problems than most people realise. In my latest post on @towardsdatascience.com I use Recursive CTEs to handle hierarchies, route finding, cycle detection and shortest paths. Sometimes you don’t need Neo4j. You just need better SQL. towardsdatascience.com/six-degrees-...
towardsdatascience.com
Recursive CTEs: SQL’s Hidden Graph Traversal Engine | Towards Data Science
A practical guide to navigate hierarchies, find routes, detect cycles and calculate degrees of separation
000
taupirho.bsky.social @taupirho.bsky.social · 21/08/2026
I tested DuckDB’s Quack protocol across three EC2 servers, each with a 10-million-row database. Quack handled remote SQL execution. Python handled the concurrent fan-out, timing and result collection. Full setup, results and code: towardsdatascience.com/running-sql-...
towardsdatascience.com
Running SQL Concurrently Across Three Remote DuckDB Servers with Quack | Towards Data Science
A small experiment in remote SQL execution
000
taupirho.bsky.social @taupirho.bsky.social · 16/08/2026
What happens when you spread 30M rows across 3 DuckDB databases on 3 AWS servers—and query them concurrently over HTTP? I built “cluster-duck” to test DuckDB’s experimental Quack protocol. Read my latest @towardsdatascience.com article for free at: towardsdatascience.com/running-sql-...
towardsdatascience.com
Running SQL Concurrently Across Three Remote DuckDB Servers with Quack | Towards Data Science
A small experiment in remote SQL execution
000
taupirho.bsky.social @taupirho.bsky.social · 15/08/2026
Prepare sensitive accounts, tax documents and client reports for delivery. Secure and protect with SecurePDF without uploading to a third party. securepdf.io #PDFSecurity #DataProtection
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 11/08/2026
matplotib or ploty On the one hand, matplotlib is a proven stalwart of the Python charting ecosystem. On the other hand, Plotly offers superb interactivity in its charts. Hmmm, which is better? Find out for free in my @towardsdatascience.com article. towardsdatascience.com/matplotlib-v...
towardsdatascience.com
Matplotlib vs Plotly: Which Python Chart Tool Should You Choose? | Towards Data Science
From Static Plots to Interactive Data Exploration
000
taupirho.bsky.social @taupirho.bsky.social · 07/08/2026
The medallion data architecture has been a mainstay of modern data platforms for years. But what exactly is it, and how do you implement it? Find out by reading my latest @towardsdatascience.com article using the link below. towardsdatascience.com/the-medallio...
towardsdatascience.com
The Medallion Data Architecture: An Introduction | Towards Data Science
A practical guide to Bronze, Silver and Gold, with a working Python and DuckDB example
000
taupirho.bsky.social @taupirho.bsky.social · 07/08/2026
Choosing between Matplotlib and Plotly? I put both libraries through their paces, comparing Matplotlib’s fine-grained control with Plotly’s built-in interactivity to find out which is best. Read my conclusion for free in my @towardsdatascience.com article. towardsdatascience.com/matplotlib-v...
towardsdatascience.com
Matplotlib vs Plotly: Which Python Chart Tool Should You Choose? | Towards Data Science
From Static Plots to Interactive Data Exploration
000
taupirho.bsky.social @taupirho.bsky.social · 07/08/2026
A practical workflow for preparing individual employee payslips for secure email delivery. Preparation and pre-delivery checks: securepdf.io/resources/pa... #PDFSecurity
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 06/08/2026
How to password-protect a batch of PDF files securepdf.io/resources/ba...
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 04/08/2026
What belongs in the Bronze, Silver and Gold layers of a medallion data architecture? My latest article for @towardsdatascience.com explains in detail what the architecture is and how to implement it using @duckdb.org . Read it for free using the link below. towardsdatascience.com/the-medallio...
towardsdatascience.com
The Medallion Data Architecture: An Introduction | Towards Data Science
A practical guide to Bronze, Silver and Gold, with a working Python and DuckDB example
010
taupirho.bsky.social @taupirho.bsky.social · 01/08/2026
Password-protect several PDFs in one browser-local workflow—without sending the documents to SecurePDF. Preparation and pre-delivery checks: securepdf.io/resources/ba... #PDFSecurity
securepdf.io
SecurePDF – Confidential PDF preparation for repeat work
Protect batches of payslips, client reports and employee PDFs locally in your browser.
000
taupirho.bsky.social @taupirho.bsky.social · 31/07/2026
If you have a Medium account and are interested in practical Agentic creation and deployment on the cloud, head on over, as I just published a new article that outlines exactly how to do that on AWS. medium.com/@thomas_reid...
medium.com
Running AI Agents in the Cloud
Build and deploy an agent on AWS with Strands and AgentCore
000
taupirho.bsky.social @taupirho.bsky.social · 24/07/2026
I built an AWS document-processing system that proved ~90% more efficient than the manual workflow it replaced. In this hands-on article for @towardsdatascience.com , I recreate it using Gmail, S3, Textract, Bedrock, Lambda, Step Functions and EventBridge. towardsdatascience.com/build-and-ru...
towardsdatascience.com
Build and Run an Intelligent Document Processing (IDP) System in the Cloud | Towards Data Science
Automating the classification and extraction of PII from emails using AWS
000
taupirho.bsky.social @taupirho.bsky.social · 20/07/2026
If you have an account over at the Medium blogging platform, you should go check out my latest article that just dropped over there on the new JIT compiler that shipped with Python 3.14. medium.com/gitconnected...
medium.com
Python 3.14 and its New JIT Compiler
A technical overview and some benchmarks
000
taupirho.bsky.social @taupirho.bsky.social · 10/07/2026
PySpark jobs don’t become slow for no reason. The culprit is often data shuffles, joins, caching, or inefficient file layouts. My latest free article on @towardsdatascience.com explains the practical PySpark concepts that help you write workflows that scale. towardsdatascience.com/pyspark-for-...
towardsdatascience.com
PySpark for Beginners: Building Intermediate-Level Skills | Towards Data Science
A practical next step into partitions, shuffles, joins, caching, and execution plans.
000
taupirho.bsky.social @taupirho.bsky.social · 01/07/2026
Most agent examples stop at “it works locally.” My latest @towardsdatascience.com article goes a step further: building an agent with AWS Strands and running it in the cloud with AWS AgentCore. towardsdatascience.com/build-and-ru...
towardsdatascience.com
Build and Run Your Own AI Agent in the Cloud | Towards Data Science
Build and deploy an agent on AWS with Strands and AgentCore
000
taupirho.bsky.social @taupirho.bsky.social · 28/06/2026
User saves 100's of dollars on token usage ...
000
taupirho.bsky.social @taupirho.bsky.social · 20/06/2026
Python 3.14 is getting an experimental JIT compiler. My latest article on @towardsdatascience.com deep dives into what this means, and what kind of performance gains you can realistically expect using it. Read my article for free on towardsdatascience.com/python-3-14-...
towardsdatascience.com
Python 3.14 and its New JIT Compiler | Towards Data Science
A technical overview and some benchmarks
000
taupirho.bsky.social @taupirho.bsky.social · 12/06/2026
The second part of my beginner's guide to PySpark just dropped on @towardsdatascience.com . In it, we take your skills to the next level by discussing performant data handling. Read it for free using the link below, towardsdatascience.com/pyspark-for-...
towardsdatascience.com
PySpark for Beginners: Beyond the Basics | Towards Data Science
Take the next step to building real workflows with Spark on your laptop
000
taupirho.bsky.social @taupirho.bsky.social · 07/06/2026
There's a great Wikipedia page called Signs of AI writing ( en.wikipedia.org/wiki/Wikiped...), which does what it says on the tin. It's very useful, so I made a Codex skill file from it. Get it at my GitHub page at, github.com/taupirho/wri...
000
taupirho.bsky.social @taupirho.bsky.social · 05/06/2026
Need a free, throwaway database in the cloud? Then you need ghost, an AI first, Postgres compatible RDBMS you can interact with using using your favourite coding agent. Read my article over on Medium for free for all the details. levelup.gitconnected.com/ghost-a-data...
levelup.gitconnected.com
Ghost: A database for our time?
The first database built for AI Agents
000
taupirho.bsky.social @taupirho.bsky.social · 05/06/2026
Convert your static local app, demo or website to a globally accessible website quickly and for free. In my latest article for @towardsdatascience.com I show three different ways to do this. Read for free using the link below. towardsdatascience.com/from-local-a...
towardsdatascience.com
From Local App to Public Website in Minutes | Towards Data Science
Three free ways to quickly deploy a static web app that anyone can access
000
taupirho.bsky.social @taupirho.bsky.social · 02/06/2026
Built an app and want to share it with the world? In my latest @TDataScience article, I take a Codex-generated "Space Invaders"- style game and deploy it to GitHub Pages, Hugging Face Spaces & here.now All for free. Check it out below. towardsdatascience.com/from-local-a...
towardsdatascience.com
From Local App to Public Website in Minutes | Towards Data Science
Three free ways to quickly deploy a static web app that anyone can access
000
taupirho.bsky.social @taupirho.bsky.social · 29/05/2026
If you would like a free resource on using the AWS Agent Toolkit, click the link below towardsdatascience.com/introducing-...
towardsdatascience.com
Introducing the Agent Toolkit for Amazon Web Services | Towards Data Science
It’s like having your own personal expert AWS solutions architect and data engineer rolled into one.
000
taupirho.bsky.social @taupirho.bsky.social · 25/05/2026
I wrote about the new Agent Toolkit for AWS, which helps coding agents use current AWS docs and APIs to build cloud systems more reliably. Includes an end-to-end example: RDS -> Glue -> S3 Tables/Iceberg -> Athena. Read it on @towardsdatascience.com for free. towardsdatascience.com/introducing-...
towardsdatascience.com
Introducing the Agent Toolkit for Amazon Web Services | Towards Data Science
It’s like having your own personal expert AWS solutions architect and data engineer rolled into one.
131
taupirho.bsky.social @taupirho.bsky.social · 11/05/2026
Pandas/Polars users — have you been curious about Spark, but assumed it was only for “big data engineers” working at massive scale? My latest article for @towardsdatascience.com explains the core concepts without the usual complexity. Check it out below towardsdatascience.com/pyspark-for-...
towardsdatascience.com
PySpark for Beginners: Mastering the Basics | Towards Data Science
A step-by-step guide to understanding distributed data, lazy logic, and your first DataFrame.
000
taupirho.bsky.social @taupirho.bsky.social · 04/05/2026
Create, migrate, tune, and run PoCs on unlimited PostgreSQL DBs in the cloud using AI coding agents. That's what Ghost. promises with its offering. To find out more, read my latest article, where I dive deep into this useful product. towardsdatascience.com/ghost-a-data...
towardsdatascience.com
Ghost: A Database for Our Times? | Towards Data Science
The first database built for AI Agents
110
taupirho.bsky.social @taupirho.bsky.social · 01/05/2026
Ghost.build is the database for our ages. Built with agents in mind, it allows you to experiment with unlimited Postgres databases in the cloud for free using your favourite coding agent. Check out my @towardsdatascience.com article for more info. towardsdatascience.com/ghost-a-data...
towardsdatascience.com
Ghost: A Database for Our Times? | Towards Data Science
The first database built for AI Agents
000
taupirho.bsky.social @taupirho.bsky.social · 21/04/2026
Python and Rust coders can get the best of both worlds as I show how to call Rust from Python in my latest @towardsdatascience.com article. Read it for FREE using the link below. towardsdatascience.com/calling-rust...
towardsdatascience.com
How to Call Rust from Python | Towards Data Science
A guide to bridging the gap between ease of use and raw performance.
000
taupirho.bsky.social @taupirho.bsky.social · 17/04/2026
My latest article just dropped onto the Towards AI publication over on the Medium blogging platform. It's all about Google's new Embedding model, which can embed text, PDF, images, audio and video. If you have an account over there, you should go check it out. medium.com/towards-arti...
medium.com
Introducing Gemini Embeddings 2 Preview
One Embedding model to rule them all
000
taupirho.bsky.social @taupirho.bsky.social · 05/04/2026
Don't roll the dice when shipping Python code. Modern tools like black, ruff , pytest and others can catch bugs early in your pipeline. For a detailed guide check out my latest post on @towardsdatascience.com
towardsdatascience.com
Building a Python Workflow That Catches Bugs Before Production | Towards Data Science
Using modern tooling to identify defects earlier in the software lifecycle.
000
taupirho.bsky.social @taupirho.bsky.social · 28/03/2026
PythoC lets you write a typed subset of Python and compile to native executables via LLVM. Useful for performance-sensitive code and shipping binaries. Breakdown + tradeoffs: towardsdatascience.com/write-c-code...
towardsdatascience.com
Write C Code With Python: The Magic of PythoC | Towards Data Science
Compile native, standalone applications using the Python syntax you already know.
000
taupirho.bsky.social @taupirho.bsky.social · 23/03/2026
PythoC is an interesting library. It allows you to create, compile and execute C code all within your familiar Python eco-system. If you want to learn more, check out my @towardsdatascience.com article where I explain more and show some coding examples. towardsdatascience.com/write-c-code...
towardsdatascience.com
Write C Code With Python: The Magic of PythoC | Towards Data Science
Compile native, standalone applications using the Python syntax you already know.
000
taupirho.bsky.social @taupirho.bsky.social · 17/03/2026
Google's new embedding 2 preview model is a cracker. It can embed text, PDF, images, audio, and video. Click my @towardsdatascience.com article below for a free introductory guide and sample code to process images and audio. towardsdatascience.com/introducing-...
towardsdatascience.com
Introducing Gemini Embeddings 2 Preview | Towards Data Science
One embedding model to rule them all
001
taupirho.bsky.social @taupirho.bsky.social · 16/03/2026
The tmux utility is a fantastic multi-window terminal tool that's easy to set up and use, but very powerful and a great time-saver. If you have a Medium account, you should definitely check it out using the link below. medium.com/gitconnected...
medium.com
A beginner's guide to Tmux: a multitasking superpower for your terminal
How to install it and create sessions, windows and panes
000
taupirho.bsky.social @taupirho.bsky.social · 13/03/2026
𝗪𝗿𝗶𝘁𝗲 𝗮𝗻𝗱 𝗰𝗼𝗺𝗽𝗶𝗹𝗲 𝗖 𝗰𝗼𝗱𝗲 𝘂𝘀𝗶𝗻𝗴 𝗣𝘆𝘁𝗵𝗼𝗻 using the PythoC library. Just pip install, and you're good to go. Read my in-depth article on the @towardsdatascience.com platform for free below, towardsdatascience.com/write-c-code...
towardsdatascience.com
Write C Code With Python: The Magic of PythoC | Towards Data Science
Compile native, standalone applications using the Python syntax you already know.
011