Sign in

Kit Menke

@kitmenke.com
372 followers 395 following 59 posts

Data Engineering leader in Saint Louis, STL Big Data I.D.E.A. meetup organizer, lifelong learner and teacher. He / him #dataBS

PostsRepliesMedia
Kit Menke @kitmenke.com · 30/09/2026
If you are searching for your next book to read, check out unread.world - after rating a few books you've read the app can recommend what to read next! Made with ❤️ by some good friends.
unread.world
Books worth reading — Unread
Discover books matched to your taste with Unread. Search titles, add ratings, and discover books that are worth reading.
000
Kit Menke @kitmenke.com · 18/04/2026
My bluesky feed is showing me tons of posts about people's dog's dying?? How do I make it stop???
100
Reposted by Kit Menke
Crystal Lewis @cghlewis.bsky.social · 07/12/2025
I just want everyone to know that if we are connected on here, I’m rooting for you. For your career, your creative endeavor, your happiness, whatever you are chasing. I hope it happens for you!
4777
Reposted by Kit Menke
Catalin Cimpanu @campuscodi.risky.biz · 10/11/2025
While AI companies are allowed to slurp everything they want, Quad9 warns that legal fees are drowning DNS resolvers, which are now being targeted by copyright owners to enforce blocks on piracy sites quad9.net/news/blog/wh...
quad9.net
Quad9 | A public and free DNS service for a better security and privacy
A public and free DNS service for a better security and privacy
17245
Kit Menke @kitmenke.com · 26/09/2025
Any meetup.com organizers that have successfully moved their community to another platform? I would be interested in hearing your experience
000
Kit Menke @kitmenke.com · 28/08/2025
Conspiracy theory: Databricks is deliberately making your clusters start up super slowly so that you want to pay more to use serverless. #dataBS
010
Kit Menke @kitmenke.com · 10/06/2025
I've been working on visualizing JOINs for some beginner SQL workshops. Here is LEFT JOIN. Thoughts? #databs youtu.be/ZSxtZAulogo?...
youtu.be
Visualizing a SQL LEFT JOIN
YouTube video by Kit Menke
000
Kit Menke @kitmenke.com · 21/05/2025
Moving from Azure Data Studio to VSCode but ugh... the VSCode SQL Server extensions are so frustrating to use.
000
Reposted by Kit Menke
rmoff 🏃‍♂️🫖🥓 @rmoff.net · 10/04/2025
Couple of big announcements from @cloudflare.social today for folk in #dataBS: * Acquisition of Arroyo, launch of Pipelines for streaming ingestion: blog.cloudflare.com/cloudflare-a... * Launch of R2 Data Catalog—a managed Apache Iceberg catalog for R2 blog.cloudflare.com/r2-data-cata...
blog.cloudflare.com
Just landed: streaming ingestion on Cloudflare with Arroyo and Pipelines
We’ve just shipped our new streaming ingestion service, Pipelines — and we’ve acquired Arroyo, enabling us to bring new SQL-based, stateful transformations to Pipelines and R2.
093
Kit Menke @kitmenke.com · 19/02/2025
Databricks recently changed the default notebook format from "source" (.py, .sql, .scala) to IPYNB which seems to indicate they will be getting rid of the source format. IMO, the ipynb format brings a few issues like difficult diffs and the potential to leak data learn.microsoft.com/en-us/azure/...
learn.microsoft.com
December 2024 - Azure Databricks
December 2024 release notes for new Azure Databricks features and improvements.
020
Reposted by Kit Menke
Matthew Mullins @mmullins.coginiti.co · 05/02/2025
I found this while looking through some scratch notes. I don't now recall what the context was, but it's an interesting thought on the evolution of the data warehouse. (Though there is an equivocation imbedded in this history) #databs
1988 Data Warehouse Architecture is introduced
1994 Data Warehouse is too difficult to build -> Data Marts
2008 Data Warehouse is too small for big data -> Hadoop
2010 Data Warehouse is too structured  -> Data Lake
2012 Data Warehouse is too difficult to scale -> Cloud Data Warehouse
2016 Data Warehouse is too difficult to manage -> Data Fabric
2019 Data Warehouse is too centralized -> Data Mesh
2020 Data Warehouse is too limiting -> Data Lakehouse
2023 Data Warehouse is too monolithic -> Composable Data Platform
043
Kit Menke @kitmenke.com · 04/02/2025
Do you version your data assets? Or is there only the current version of a database table? What about the table definition? #dataBS
300
Kit Menke @kitmenke.com · 24/01/2025
An agile ceremony / rite of passage nobody mentions: arguing about story points and what they mean.
010
Reposted by Kit Menke
Kelsey Hightower @kelseyhightower.com · 23/01/2025
1. Impact. How much revenue does my work protect or generate? 2. Quality. Does my work meet or exceed customer expectations? 3. Efficiency. Reward making the right buy versus build decision. 4. Reusability. How do others leverage my work? 5. Supportability. How much work do I create for others?
321001119
Reposted by Kit Menke
David Asboth @davidasboth.com · 22/01/2025
I'm out walking and had some thoughts about data and fun stuff and mental health that I wanted to share. #dataBS
5121
Reposted by Kit Menke
Alex Power[s] @itsnotaboutthecell.com · 14/01/2025
Gahhh it’s time! @data-dragoness.bsky.social devUp call for speakers! Let’s take over with the #PowerPlatform and #MicrosoftFabric topics! For anyone who’s in the middle west, let’s do this! sessionize.com/dev-up-2025
sessionize.com
dev up 2025: Call for Speakers
The 2025 dev up conference is being held in St. Louis, Missouri from August 6-8, 2025. We are excited to be back and we are putting out the call to ...
462
Kit Menke @kitmenke.com · 06/12/2024
Great breakdown of the new S3 Tables feature that leverages Apache Iceberg. Including an explanation of the costs... which are complicated. #dataBS bigdata.2minutestreaming.com/p/meet-your-...
bigdata.2minutestreaming.com
meet your new data lakehouse: S3 Iceberg Tables
S3 Tables and S3 Metadata are two brutal new features that compete with common Apache Iceberg Lakehouse architectures
061
Reposted by Kit Menke
David Jayatillake @jayatillake.bsky.social · 04/12/2024
The more I think about yesterday's announcement about Amazon S3 Tables, the more I think that it changes things a great deal. The gravity of data has shifted from the warehouse to cloud storage... but is there really a difference any more? 🧵1/n www.businesswire.com/news/home/20....
businesswire.com
Amazon S3 Expands Capabilities with Managed Apache Iceberg Tables for Faster Data Lake Analytics and Automatic Metadata Generation to Simplify Data Discovery and Understanding
At AWS re:Invent, Amazon Web Services, Inc. (AWS), an Amazon.com, Inc. company (NASDAQ: AMZN), today announced new Amazon Simple Storage Service (Amaz
1193
Kit Menke @kitmenke.com · 03/12/2024
Arch Data Network is hosting an event around Apache Airflow and how it is being used. This Thursday, December 5th from 5:30 - 7:30 pm in Creve Coeur www.linkedin.com/events/archd...
linkedin.com
Arch Data Network December 5th Event | LinkedIn
Apache Airflow is popular because it provides a powerful, scalable, and flexible solution for orchestrating complex data workflows and automating processes across diverse environments. Join us as we d...
030
Kit Menke @kitmenke.com · 20/11/2024
I made a Starter Pack for people in the Saint Louis, Missouri area who are doing cool stuff in Data Engineering, Data Analytics, or Data Science. If you're doing data in STL let me know! #datasky #dataBS go.bsky.app/SZUtRw3
230
Kit Menke @kitmenke.com · 20/11/2024
In two weeks, the St. Louis Big Data I.D.E.A. meetup is hosting @chad-isenberg.bsky.social to give an overview on dbt, alternatives, and the future of "the last mile" in data management. Beginners welcome! 🗓️ When: December 4, 2024 @ 5:30 PM 📍Where: Virtually on Zoom RSVP below! #dataBS #datasky
meetup.com
[VIRTUAL] A Brief Introduction to dbt, Wed, Dec 4, 2024, 5:30 PM | Meetup
In this talk, we'll cover what dbt is, why it's useful, alternatives, and the future of "the last mile" in data management. I'll assume folks have no knowledge of dbt and m
030
Reposted by Kit Menke
Simon Willison @simonwillison.net · 20/11/2024
Foursquare just open sourced their 100 million place point of interest dataset! Some notes on poking around with it using DuckDB (it's Parquet files on S3) simonwillison.net/2024/Nov/20/...
simonwillison.net
Foursquare Open Source Places: A new foundational dataset for the geospatial community
I did not expect this! > [...] we are announcing today the general availability of a foundational open data set, Foursquare Open Source Places ("FSQ OS Places"). This base layer …
23458113
Reposted by Kit Menke
Amanda Alvarez @gecky.me · 29/10/2024
New feed! Add #dataBS or #databsky to your post, and it'll get ingested into this custom Data BS feed, which looks back over 7 days of posts.
34520
Reposted by Kit Menke
Theodore Manassis @mamonu.bsky.social · 17/11/2024
Some work that I have been involved in the last year. I hope you like the blogpost from our lead, Soumaya as its a very interesting solution. Not all problems are nails to the hammer of Spark :) ministryofjustice.github.io/data-and-ana...
ministryofjustice.github.io
Building a transaction data lake using Amazon Athena, Apache Iceberg and dbt
How we leveraged Amazon Athena, along with the Apache Iceberg table format and the dbt SQL management framework, to build robust, scalable and maintainable ELT (extract, load, transform) pipelines.
142
Reposted by Kit Menke
Martin Kleppmann @martin.kleppmann.com · 16/11/2024
In June Elon posted this graph of the rate of likes on X. It doesn't have a unit on the y axis, but it's plausible to assume that it's events/sec. If that is true, X in June was handling about 20k likes/sec. For comparison, Bluesky is now handling about 700 likes/sec during the busy part of the day.
A graph showing the rate of likes on X over the course of a couple of hours on 12 June 2024. It fluctuates between about 17k and 20k; the y axis doesn't have a unit. A second line shows a comparison to the prior week, which is slightly lower. There are strange unexplained spikes on the hour every hour.A graph showing the rate of various types of events on Bluesky over the last day. At peak there are about 800 follow events and 700 like events per second (note the likes are scaled down by a factor of 5; the right-hand y axis is for likes).
1455176
Kit Menke @kitmenke.com · 16/11/2024
Thinking about creating a STL Data starter pack for those doing cool stuff in data engineering, data analytics, and data science...
290
Reposted by Kit Menke
David Gasquez @davidgasquez.com · 14/11/2024
New post up! ✨ Exploring AT Protocol with Python to visualize the #databs social graph! davidgasquez.com/exploring-at... Took less than 1 hour to get the data and plot it. Amazing what you can do with open APIs and great SDKs!
a gr
1210820
Kit Menke @kitmenke.com · 15/11/2024
I'm all set up with SAP PowerDesigner. Next step... World domination.
000
Reposted by Kit Menke
darth™️ @darthbluesky.bsky.social · 13/11/2024
if u are wondering about that red pin emoji here 📌 if u find a post u want to return to u can stick a pin in it by replying with the red pin emoji then use @jaz.bsky.social's red pin feed which keeps track of your red pins
239690392952
Reposted by Kit Menke
Joshua J. Friedman @joshuajfriedman.com · 13/11/2024
By popular request, the moment on May 1, 2023, when Jake Tapper went on TV and referred to a Bluesky post as a "skeet"
401174183
Kit Menke @kitmenke.com · 11/11/2024
I have to be honest, updating my website is a real bummer because there is always some breaking change that Hugo has introduced.
100
Kit Menke @kitmenke.com · 09/11/2024
Cat tax!
An orange cat curled up on a chair, sleeping
010
Reposted by Kit Menke
Star Trek Minus Context @nocontexttrek1.bsky.social · 06/11/2024
Zoomed out we see a small, lone shuttlecraft flying through a cloudy sky. Closed caption reads, "(screaming)"
6582672286
Kit Menke @kitmenke.com · 05/11/2024
Do you unit test Databricks SQL code? SQL pipelines seem impossible to test in an automated way. A local option is to use open source Spark but that seems incomplete if you're using Databricks specific syntax. #dataBS
210
Kit Menke @kitmenke.com · 03/11/2024
Where's everyone at on virtual vs in-person meetups these days?
100
Kit Menke @kitmenke.com · 30/10/2024
Hello Blue sky! I am an architect and data engineer with a focus on distributed systems and real-time streaming data pipelines. In the past few years I've worked a lot with Spark, Azure, and Databricks! Outside of work, I organize the STL Big Data I.D.E.A. meetup and do some volunteering.
150