Julien Hurault @hachej.bsky.social · 13/12/2024Indeed it s not simple unfortunately..just the a way to get started quickly atm . 060
Julien Hurault @hachej.bsky.social · 13/12/2024For iceberg catalog hard to find a simpler setup.. 110
Julien Hurault @hachej.bsky.social · 13/12/2024Nice! Can you orchestrate lambda or ecs tasks that way? 120
Julien Hurault @hachej.bsky.social · 13/12/2024Just use Pyiceberg with AWS Glue, probably the fastest way to get started. 120
Julien Hurault @hachej.bsky.social · 06/12/2024In term of volume of data exchange over the marketplace? No idea 010
Julien Hurault @hachej.bsky.social · 06/12/2024SF sales rep told me that markeplace was THE feature that helped a lot to convert 120
Julien Hurault @hachej.bsky.social · 06/12/2024For those lost in GCP terminology: I wrote a summary of Iceberg integration in GCP a couple of weeks ago: juhache.substack.com/p/gcp-and-ic...juhache.substack.comGCP & IcebergJu Data Engineering Weekly - Ep 77 000
Julien Hurault @hachej.bsky.social · 06/12/2024New blog post: Building a 0$ Data Distribution System. juhache.substack.com/p/0-data-dis...juhache.substack.com0$ Data DistributionJu Data Engineering Weekly - Ep 78 040
Julien Hurault @hachej.bsky.social · 02/12/2024check catalog.boringdata.io/dashboard/in...catalog.boringdata.ioHome 020
Julien Hurault @hachej.bsky.social · 30/11/2024" bash / make knowledge, a single instance SQL processing engine (DuckDB, CHDB or a few python scripts), a distributed file system, git and a developer workflow (CI/CD)" what s your best option to orchestate sql models in such setup? 100
Reposted by Julien HuraultJavi Santana @javisantana.bsky.social · 30/11/2024Some learnings after helping +50 companies in high performance data engineering projects javisantana.com/2024/11/30/l...javisantana.comjavisantana.com 43812
Julien Hurault @hachej.bsky.social · 30/11/2024Super good thx! "immutable workflow + atomic operation" 100%! 110
Julien Hurault @hachej.bsky.social · 30/11/2024Niiiice, your view is doing a read_parquet(*) on their bucket? Or do you copy the data? 000
Julien Hurault @hachej.bsky.social · 29/11/20241 Docker container embedding app code + SQLite DB → live chat app with 10k simultaneous users. youtu.be/0rlATWBNvMwyoutu.beDHH discusses SQLite (and Stoicism)YouTube video by Aaron Francis 010
Julien Hurault @hachej.bsky.social · 29/11/2024Where is the data stored, then? In DuckDB itself? So, if you have a 1GB dataset, does that mean you’ll share a single .duckdb file containing the entire dataset? Or either a view pointing to parquet files: CREATE VIEW... as read_parquet(*.parquet) ? 000
Julien Hurault @hachej.bsky.social · 29/11/2024Do you see DuckDB as a format? For me: • Parquet = Standard storage format • Iceberg = Standard metadata format • DuckDB = One possible distribution vector 100
Reposted by Julien HuraultJake Gold @jacob.gold · 11/11/2024Yup, there are almost fifteen million SQLite databases on Bluesky’s PDS servers. It’s wildly efficient and simple but not without trade offs of course. Makes sense for this use case in large part because each users atproto repository is self contained, with links to other repos, like a website. 2396
Julien Hurault @hachej.bsky.social · 28/11/2024just do both -> catalog.boringdata.io/dashboard/in...catalog.boringdata.ioHome 110
Reposted by Julien HuraultBen Lindsay @benjlindsay.com · 27/11/2024Here's what I put in my ~/.zshrc file to make sure my virtualenv autoactivates when I move to a directory with a .venv file. Works well for me so far. Do the rest of you do something like this? #Python #DataBS 8372
Julien Hurault @hachej.bsky.social · 28/11/2024Prediction: poeple will monetize custom feeds github.com/bluesky-soci...github.comGitHub - bluesky-social/feed-generator: ATProto Feed Generator Starter KitATProto Feed Generator Starter Kit. Contribute to bluesky-social/feed-generator development by creating an account on GitHub. 000
Reposted by Julien HuraultNikhil Benesch @benesch.bsky.social · 26/11/2024Something interesting is brewing in Iceberg-on-S3 land. 👀 lists.apache.org/thread/v7x65... cc @eatonphil.bsky.sociallists.apache.org 3295
Julien Hurault @hachej.bsky.social · 26/11/2024Many vendors do not display their price on their landing page... Probably for that reason: can t be reverted. 110
Julien Hurault @hachej.bsky.social · 25/11/2024Head for brittany not Paris :) en.m.wikipedia.org/wiki/Brittanyen.m.wikipedia.orgBrittany - Wikipedia 000
Julien Hurault @hachej.bsky.social · 25/11/2024Building a data pipeline = 50% bringing data from A to B at time t 50% making pipeline fixing easy at time t+1 100
Julien Hurault @hachej.bsky.social · 24/11/2024New blog post: GCP & Iceberg open.substack.com/pub/juhache/... 000
Julien Hurault @hachej.bsky.social · 24/11/2024Is there a starter pack for data engineering here on bluesky? 320
Julien Hurault @hachej.bsky.social · 23/11/2024merged with location.foursquare.com/products/pla...location.foursquare.comPlacesFoursquare's Places allows developers to unlock points of interest data (POI) with precision and in rich detail - all stored in your POI database. Learn more. 010
Julien Hurault @hachej.bsky.social · 23/11/20242010 — 2017: ML = pip install scikit-learn 2017 — 2023: ML = pip install torch 2023 — : ML = pip install requests 000
Julien Hurault @hachej.bsky.social · 22/11/2024FINALLY! 🎉 aws.amazon.com/blogs/comput... No more endless searching for support with intrinsic docs.aws.amazon.com/step-functio...lnkd.inLinkedInThis link will take you to a page that’s not on LinkedIn 000
Julien Hurault @hachej.bsky.social · 19/11/2024app.snowflake.com/marketplace/...app.snowflake.comSnowflake 000
Julien Hurault @hachej.bsky.social · 19/11/2024World wide address data? app.snowflake.com/marketplace/...app.snowflake.comSnowflake 100
Julien Hurault @hachej.bsky.social · 19/11/2024Some datasets, probably yes. The Snowflake Marketplace experience is cool: any dataset just one SELECT away, but it’s locked inside the Snowflake ecosystem. I love the idea of taking that concept outside of Snowflake—having any dataset available with just one ATTACH command. 110
Julien Hurault @hachej.bsky.social · 19/11/2024What are the cost for this setup? Storage + compute to bring the data to the bucket? Distribution via R2 should be free no? 010
Julien Hurault @hachej.bsky.social · 19/11/2024What about dumping Snowflake Marketplace free listings to R2 ? @jakthom.bsky.social @ssp.sh @tobilg.com 110
Julien Hurault @hachej.bsky.social · 19/11/2024The Snowflake unbundling continues: • Snowflake Storage → Iceberg • Snowflake Marketplace → Cloudflare R2 + DuckDB :) 210