Sign in

Kit Menke

@kitmenke.com
372 followers 395 following 59 posts

Data Engineering leader in Saint Louis, STL Big Data I.D.E.A. meetup organizer, lifelong learner and teacher. He / him #dataBS

PostsRepliesMedia
Kit Menke @kitmenke.com · 30/09/2026
If you are searching for your next book to read, check out unread.world - after rating a few books you've read the app can recommend what to read next! Made with ❤️ by some good friends.
unread.world
Books worth reading — Unread
Discover books matched to your taste with Unread. Search titles, add ratings, and discover books that are worth reading.
000
Kit Menke @kitmenke.com · 26/05/2026
Nice I saw that on LI, I will share it with the group
050
Kit Menke @kitmenke.com · 26/05/2026
Thank you! Building the data community is hard work, as you know ❤️
140
Kit Menke @kitmenke.com · 17/05/2026
Join us on the EV side!
110
Kit Menke @kitmenke.com · 24/04/2026
Void Kitties are the best!
140
Kit Menke @kitmenke.com · 18/04/2026
Unsubscribe!!
000
Kit Menke @kitmenke.com · 18/04/2026
My bluesky feed is showing me tons of posts about people's dog's dying?? How do I make it stop???
100
Kit Menke @kitmenke.com · 11/03/2026
Kinda like repomix?
100
Kit Menke @kitmenke.com · 24/02/2026
Woah what?!
100
Kit Menke @kitmenke.com · 23/02/2026
I just started Google AI Pro with Antigravity for $20 a month and it includes access to Gemini and Claude models. I've hit the limits on Claude a few times but not bad.
020
Kit Menke @kitmenke.com · 09/02/2026
Markdown for sure. You can read it as plain text or render it, it is widely accepted across the development community, and AI speaks it very well.
010
Kit Menke @kitmenke.com · 15/01/2026
Agreed, hybrid is the worst.
100
Kit Menke @kitmenke.com · 30/12/2025
Timely link because I'm currently struggling with this. The article is interesting and not at all what I expected. I've always heard Silver is supposed to be the integration layer using 3NF, not star schema. So many more questions! www.databricks.com/glossary/med...
databricks.com
What is a Medallion Architecture?
A medallion architecture is a data design pattern used to logically organize data in a lakehouse, with the goal of improving the structure and quality of data.
120
Reposted by Kit Menke
Crystal Lewis @cghlewis.bsky.social · 07/12/2025
I just want everyone to know that if we are connected on here, I’m rooting for you. For your career, your creative endeavor, your happiness, whatever you are chasing. I hope it happens for you!
4777
Reposted by Kit Menke
Catalin Cimpanu @campuscodi.risky.biz · 10/11/2025
While AI companies are allowed to slurp everything they want, Quad9 warns that legal fees are drowning DNS resolvers, which are now being targeted by copyright owners to enforce blocks on piracy sites quad9.net/news/blog/wh...
quad9.net
Quad9 | A public and free DNS service for a better security and privacy
A public and free DNS service for a better security and privacy
17245
Kit Menke @kitmenke.com · 06/11/2025
Perhaps using unnest? select unnest(value, recursive := true) from read_json('~/Data/example.json')
100
Kit Menke @kitmenke.com · 26/09/2025
Any meetup.com organizers that have successfully moved their community to another platform? I would be interested in hearing your experience
000
Kit Menke @kitmenke.com · 01/09/2025
My blogging motivation has declined a lot over the years... then the endless Hugo breaking changes pretty much killed it for good. Do you know how much work it would be to convert a Hugo blog over to Zola?
100
Kit Menke @kitmenke.com · 28/08/2025
Conspiracy theory: Databricks is deliberately making your clusters start up super slowly so that you want to pay more to use serverless. #dataBS
010
Kit Menke @kitmenke.com · 10/06/2025
I've been working on visualizing JOINs for some beginner SQL workshops. Here is LEFT JOIN. Thoughts? #databs youtu.be/ZSxtZAulogo?...
youtu.be
Visualizing a SQL LEFT JOIN
YouTube video by Kit Menke
000
Kit Menke @kitmenke.com · 21/05/2025
Moving from Azure Data Studio to VSCode but ugh... the VSCode SQL Server extensions are so frustrating to use.
000
Kit Menke @kitmenke.com · 25/04/2025
I met my wife on xanga! ♥️
150
Reposted by Kit Menke
rmoff 🏃‍♂️🫖🥓 @rmoff.net · 10/04/2025
Couple of big announcements from @cloudflare.social today for folk in #dataBS: * Acquisition of Arroyo, launch of Pipelines for streaming ingestion: blog.cloudflare.com/cloudflare-a... * Launch of R2 Data Catalog—a managed Apache Iceberg catalog for R2 blog.cloudflare.com/r2-data-cata...
blog.cloudflare.com
Just landed: streaming ingestion on Cloudflare with Arroyo and Pipelines
We’ve just shipped our new streaming ingestion service, Pipelines — and we’ve acquired Arroyo, enabling us to bring new SQL-based, stateful transformations to Pipelines and R2.
093
Kit Menke @kitmenke.com · 25/02/2025
Chispa has good diffs for PySpark dataframes github.com/MrPowers/chi...
github.com
GitHub - MrPowers/chispa: PySpark test helper methods with beautiful error messages
PySpark test helper methods with beautiful error messages - MrPowers/chispa
000
Kit Menke @kitmenke.com · 19/02/2025
Databricks recently changed the default notebook format from "source" (.py, .sql, .scala) to IPYNB which seems to indicate they will be getting rid of the source format. IMO, the ipynb format brings a few issues like difficult diffs and the potential to leak data learn.microsoft.com/en-us/azure/...
learn.microsoft.com
December 2024 - Azure Databricks
December 2024 release notes for new Azure Databricks features and improvements.
020
Kit Menke @kitmenke.com · 05/02/2025
I had an HDMI KVM but it was still annoying to switch back and forth. Plus I wanted to use the full resolution at 144Hz on my gaming PC. Now I just have a big desk with separate keyboards/mice/monitors.
010
Kit Menke @kitmenke.com · 05/02/2025
Yes, I'm working on this right now and talking about how we can potentially "upgrade" some of the dimensions without breaking everything. 🙃
020
Reposted by Kit Menke
Matthew Mullins @mmullins.coginiti.co · 05/02/2025
I found this while looking through some scratch notes. I don't now recall what the context was, but it's an interesting thought on the evolution of the data warehouse. (Though there is an equivocation imbedded in this history) #databs
1988 Data Warehouse Architecture is introduced
1994 Data Warehouse is too difficult to build -> Data Marts
2008 Data Warehouse is too small for big data -> Hadoop
2010 Data Warehouse is too structured  -> Data Lake
2012 Data Warehouse is too difficult to scale -> Cloud Data Warehouse
2016 Data Warehouse is too difficult to manage -> Data Fabric
2019 Data Warehouse is too centralized -> Data Mesh
2020 Data Warehouse is too limiting -> Data Lakehouse
2023 Data Warehouse is too monolithic -> Composable Data Platform
043
Kit Menke @kitmenke.com · 04/02/2025
Thanks for the input and I agree... Right now I'm battling a mono-repo used by a big team with limited git knowledge and no tooling. Choosing a tool like dbt/flyway/liquibase could help force some standardization.
000
Kit Menke @kitmenke.com · 04/02/2025
Do you ever feel like it is difficult to keep them in sync with what is deployed to the database? Or with many people working in the same repo?
100
Kit Menke @kitmenke.com · 04/02/2025
Is keeping the table definition valuable for only certain databases? For example in Databricks you can easily get the definition and there aren't any indexes to store. Compared to SQL Server (or similar) where it is difficult to figure out what was deployed.
100
Kit Menke @kitmenke.com · 04/02/2025
You did this only for certain breaking changes right? For example - meaning of the data in a column changed or columns removed. How did you maintain two separate versions of the schema?
200
Kit Menke @kitmenke.com · 04/02/2025
Do you version your data assets? Or is there only the current version of a database table? What about the table definition? #dataBS
300
Kit Menke @kitmenke.com · 24/01/2025
An agile ceremony / rite of passage nobody mentions: arguing about story points and what they mean.
010
Reposted by Kit Menke
Kelsey Hightower @kelseyhightower.com · 23/01/2025
1. Impact. How much revenue does my work protect or generate? 2. Quality. Does my work meet or exceed customer expectations? 3. Efficiency. Reward making the right buy versus build decision. 4. Reusability. How do others leverage my work? 5. Supportability. How much work do I create for others?
321001119
Reposted by Kit Menke
David Asboth @davidasboth.com · 22/01/2025
I'm out walking and had some thoughts about data and fun stuff and mental health that I wanted to share. #dataBS
5121
Reposted by Kit Menke
Alex Power[s] @itsnotaboutthecell.com · 14/01/2025
Gahhh it’s time! @data-dragoness.bsky.social devUp call for speakers! Let’s take over with the #PowerPlatform and #MicrosoftFabric topics! For anyone who’s in the middle west, let’s do this! sessionize.com/dev-up-2025
sessionize.com
dev up 2025: Call for Speakers
The 2025 dev up conference is being held in St. Louis, Missouri from August 6-8, 2025. We are excited to be back and we are putting out the call to ...
462
Kit Menke @kitmenke.com · 14/01/2025
In my experience I'm seeing companies using Spark switch from Scala to Python for two reasons: Python has an easier learning curve and Scala devs are much harder to find.
220
Kit Menke @kitmenke.com · 06/12/2024
Great breakdown of the new S3 Tables feature that leverages Apache Iceberg. Including an explanation of the costs... which are complicated. #dataBS bigdata.2minutestreaming.com/p/meet-your-...
bigdata.2minutestreaming.com
meet your new data lakehouse: S3 Iceberg Tables
S3 Tables and S3 Metadata are two brutal new features that compete with common Apache Iceberg Lakehouse architectures
061
Kit Menke @kitmenke.com · 06/12/2024
Great talk, RIP Strangeloop
010
Reposted by Kit Menke
David Jayatillake @jayatillake.bsky.social · 04/12/2024
The more I think about yesterday's announcement about Amazon S3 Tables, the more I think that it changes things a great deal. The gravity of data has shifted from the warehouse to cloud storage... but is there really a difference any more? 🧵1/n www.businesswire.com/news/home/20....
businesswire.com
Amazon S3 Expands Capabilities with Managed Apache Iceberg Tables for Faster Data Lake Analytics and Automatic Metadata Generation to Simplify Data Discovery and Understanding
At AWS re:Invent, Amazon Web Services, Inc. (AWS), an Amazon.com, Inc. company (NASDAQ: AMZN), today announced new Amazon Simple Storage Service (Amaz
1193
Kit Menke @kitmenke.com · 03/12/2024
Arch Data Network is hosting an event around Apache Airflow and how it is being used. This Thursday, December 5th from 5:30 - 7:30 pm in Creve Coeur www.linkedin.com/events/archd...
linkedin.com
Arch Data Network December 5th Event | LinkedIn
Apache Airflow is popular because it provides a powerful, scalable, and flexible solution for orchestrating complex data workflows and automating processes across diverse environments. Join us as we d...
030
Kit Menke @kitmenke.com · 28/11/2024
Done!
010
Kit Menke @kitmenke.com · 22/11/2024
Having just spent more hours fixing things after yet another breaking Hugo change, this is something I need to seriously consider.
020
Kit Menke @kitmenke.com · 21/11/2024
Done!
010
Kit Menke @kitmenke.com · 20/11/2024
I made a Starter Pack for people in the Saint Louis, Missouri area who are doing cool stuff in Data Engineering, Data Analytics, or Data Science. If you're doing data in STL let me know! #datasky #dataBS go.bsky.app/SZUtRw3
230
Kit Menke @kitmenke.com · 20/11/2024
In two weeks, the St. Louis Big Data I.D.E.A. meetup is hosting @chad-isenberg.bsky.social to give an overview on dbt, alternatives, and the future of "the last mile" in data management. Beginners welcome! 🗓️ When: December 4, 2024 @ 5:30 PM 📍Where: Virtually on Zoom RSVP below! #dataBS #datasky
meetup.com
[VIRTUAL] A Brief Introduction to dbt, Wed, Dec 4, 2024, 5:30 PM | Meetup
In this talk, we'll cover what dbt is, why it's useful, alternatives, and the future of "the last mile" in data management. I'll assume folks have no knowledge of dbt and m
030
Reposted by Kit Menke
Simon Willison @simonwillison.net · 20/11/2024
Foursquare just open sourced their 100 million place point of interest dataset! Some notes on poking around with it using DuckDB (it's Parquet files on S3) simonwillison.net/2024/Nov/20/...
simonwillison.net
Foursquare Open Source Places: A new foundational dataset for the geospatial community
I did not expect this! > [...] we are announcing today the general availability of a foundational open data set, Foursquare Open Source Places ("FSQ OS Places"). This base layer …
23458113
Kit Menke @kitmenke.com · 19/11/2024
Full Stack Data Engineer
010
Kit Menke @kitmenke.com · 19/11/2024
Meme from the movie Zoolander of Mugatu (Will Ferrell) saying "DuckDB so hot right now"
070