Sign in

Burak

@buremba.bsky.social
92 followers 159 following 87 posts

Data Engineer - Cooking github.com/buremba/universql 🐥

PostsRepliesMedia
Burak @buremba.bsky.social · 27/05/2025
Soon to be, it looks like: youtu.be/zeonmOO9jm4?... Otherwise, there is no point of using Parquet instead of their DuckDB native format. I’m glad they didn’t ignore the “industry standards”
youtu.be
Introducing DuckLake
YouTube video by DuckDB
050
Burak @buremba.bsky.social · 27/05/2025
Is there any plan to support data compaction to data lake when data inlining is used?
140
Burak @buremba.bsky.social · 27/05/2025
I was worried about Iceberg being ignored in favor of DuckLake but looks like you fixed Iceberg’s biggest problems and still kept the compatibility. Super exciting!
130
Burak @buremba.bsky.social · 27/05/2025
Turns out the implementation wasn’t WAL but they had a new Iceberg compatible data lake extension. I like the direction they are going!
110
Burak @buremba.bsky.social · 21/05/2025
I have this one but they might have soon to be public extension to use the WAL to keep the data in sync with data lake: github.com/duckdb/duckd...
github.com
Implement WALReader by adsharma · Pull Request #17247 · duckdb/duckdb
This could be useful to external replication tools to read WAL records similar to how wal2json (Postgres) and binlog (MySQL) work. Translation to externally consumable format is not included.
110
Burak @buremba.bsky.social · 21/05/2025
Is @duckdb.org cooking native data lake integration with streaming support through WAL? That could enable DuckDB to have a multi-user mode..
130
Burak @buremba.bsky.social · 11/05/2025
After not using Facebook for years, wanted to try out Marketplace. Apparently you can send messages to people on the website but you can only see messages are sent to you on their Messenger app. I guess this is their definition of “connecting people”.
000
Burak @buremba.bsky.social · 05/05/2025
That’s a good analogy, might steal it. :) However; when the destination path is not clear (which is usually case as you need to experiment and iterate anyways) smashing can help accelerate finding the destination as you learn where not to go.
000
Burak @buremba.bsky.social · 26/04/2025
Ironically the number of stale documents in our company is increased dramatically thanks to LLM.
010
Burak @buremba.bsky.social · 17/03/2025
Oh I lost count of how much time I waste trying to infer the column names from random CSV files without a header. This is very handy!
020
Burak @buremba.bsky.social · 22/02/2025
Just found out that Databricks hired Snowflake’s Polaris (Iceberg) lead PM. It’s crazy how aggressive these guys with the competition!
050
Burak @buremba.bsky.social · 18/02/2025
Great to see Amazon implementing Iceberg REST Catalog layer for Glue! It enables read/write support on S3Tables from any Iceberg client, now everybody as a free Iceberg catalog via AWS Glue. aws.amazon.com/blogs/storag...
010
Burak @buremba.bsky.social · 15/02/2025
Exactly! I think Flight will get more popular over time as it's the most efficient implementation, but this approach can help existing RESTFul apps to adopt SQL integrations before switching over to GRPC.
100
Burak @buremba.bsky.social · 15/02/2025
The main inspirations are github.com/PostgREST/po... and @qxip.bsky.social 's DuckDB webmacro extension: duckdb.org/community_ex...
github.com
GitHub - PostgREST/postgrest: REST API for any Postgres database
REST API for any Postgres database. Contribute to PostgREST/postgrest development by creating an account on GitHub.
021
Burak @buremba.bsky.social · 15/02/2025
Released an experimental @fastapi.tiangolo.com integration with @duckdb.org today, which enables REST APIs to have bidirectional read/write support in SQL. github.com/buremba/duck...
291
Burak @buremba.bsky.social · 10/02/2025
Pretty common but if one of these languages is the “main” one, it might be more desirable to generate JSONSchema from Pydantic/TS and generate the models for other language from JSONSchema. It’s more about where you want the source of truth should be.
120
Burak @buremba.bsky.social · 06/02/2025
I had the exact same thought..
020
Burak @buremba.bsky.social · 05/02/2025
"think twice before you speak."
000
Burak @buremba.bsky.social · 29/01/2025
Thanks. I'm also a fan of your creative extensions! Quackpipe was one of the inspirations. :)
010
Burak @buremba.bsky.social · 28/01/2025
Today I had to explain my partner what @duckdb.org is because “I will fly to Amsterdam for a day to meet ducks” didn’t make any sense to her. Excited to meet with the contributors! duckdb.org/events/2025/...
duckdb.org
DuckCon #6 in Amsterdam
DuckDB is an in-process SQL database management system focused on analytical query processing. It is designed to be easy to install and easy to use. DuckDB has no external dependencies. DuckDB has bin...
1130
Burak @buremba.bsky.social · 28/01/2025
One here! 🍻
010
Burak @buremba.bsky.social · 24/01/2025
It's interesting to see many seed-stage, well-funded startups trying to "re-write X in Rust." as a business model. WarpStream, ScyllaDB, and Redpanda are successful because they're either 10x efficient or make the maintenance much easier than their alternative, not because they're written in C++
160
Burak @buremba.bsky.social · 14/01/2025
I couldn't figure out how to insert a table into an S3 Table without Spark. I tried to use the API but it requires me to create the files and update the metadata. PyIceberg can't write to S3 Tables through its S3 integration yet so I had to stick to Spark. boto3.amazonaws.com/v1/documenta...
boto3.amazonaws.com
update_table_metadata_location - Boto3 1.35.99 documentationContentsMenuExpandLight modeDark modeAuto light/dark modeClose Menu
210
Burak @buremba.bsky.social · 14/01/2025
If AWS is serious about S3 Tables, they should support Iceberg REST Catalog in it. Right now we can only create tables with Spark.
110
Burak @buremba.bsky.social · 14/01/2025
Qlik's Upsolver acquisition shows the importance of adopting new technologies as a potential acquisition target for bigger companies. It's a 10-year-old company, and they raised a ton, so I'm not sure how good the deal was for the co-founders.
000
Burak @buremba.bsky.social · 14/01/2025
dbt acquiring SDF Labs shows how important it is to have a good relationship with your competitors. SQLMesh might be more ambitious, but I'm sure it was a good exit for SDF founders in only 2 years!
010
Burak @buremba.bsky.social · 14/01/2025
It's a good day to be acquired in the data space.
210
Burak @buremba.bsky.social · 03/01/2025
For the record I checked if Motherduck notebooks ahave it but doesn’t seem to be the case, at least yet.
100
Burak @buremba.bsky.social · 03/01/2025
Look great! I would love to try out, Where is this going to be available?
100
Burak @buremba.bsky.social · 03/01/2025
People say LLM is killing low-code platforms such as Retool and Bubble, but they seem to hire more people + raise even more funding. They're better positioned to leverage LLM maybe. The AI tools like bolt.new and v0.dev work best with Next + Shacdn combination after all, so I wouldn't be surprised.
000
Burak @buremba.bsky.social · 01/01/2025
Workers AI supports models like llama-3.3-70b and it’s powered by containers according to their announcement so I hope it will TB level limits.. I also wonder how they will position container support. Wouldn’t it be better to just call it “custom workers” similar to container support in Lambda?
110
Burak @buremba.bsky.social · 31/12/2024
I also use Lambda but CF Workers is very appealing especially when the data is in R2 for me.
100
Burak @buremba.bsky.social · 30/12/2024
Can you run DuckDB on Cloudflare? I haven't tried Python worker but since it uses Pyodide I don't it's unlikely to run and AFAIK WASM version still has some more work: github.com/duckdb/duckd...
github.com
Cloudflare workers · duckdb duckdb-wasm · Discussion #430
Has anyone tried to use this within a CF worker ? I think it's a great use case, put your data in some parquet files somewhere ( web server, maybe later even integrate with CF R2 ), run your query ...
100
Burak @buremba.bsky.social · 30/12/2024
When I hand-write too much duplicated YAML, I feel like I'm too smart for the task but then after trying out these fancy config languages I feel like I'm too dumb to use them. 🫠
010
Burak @buremba.bsky.social · 30/12/2024
Yeah that's correct. I tested Starlark (github.com/bazelbuild/s...) the other day and thought I could use Python instead. I'm mostly interested in the IDE integrations (VSCode + Intellij) but TBH Python has first-class support in both IDEs and it's hard to beat.
github.com
GitHub - bazelbuild/starlark: Starlark Language
Starlark Language. Contribute to bazelbuild/starlark development by creating an account on GitHub.
010
Burak @buremba.bsky.social · 30/12/2024
I use YAML mostly at work for some internal projects and some for personal dbt-based transformations. I aim to reduce duplication by reusing the definitions & automating the YAML generation where possible. Started experimenting with PKL (github.com/apple/pkl) and CUE (cuelang.org)
github.com
GitHub - apple/pkl: A configuration as code language with rich validation and tooling.
A configuration as code language with rich validation and tooling. - apple/pkl
110
Burak @buremba.bsky.social · 30/12/2024
Is the cache local or remote? With WASM I thought people mostly rely on browser cache but I might be wrong.
110
Burak @buremba.bsky.social · 30/12/2024
Mine is writing less YAML and I am hopeful
130
Burak @buremba.bsky.social · 21/12/2024
Yeah it’s kinda black box on when the compaction kicks in and how it works. The API has relevant features but they are not surfaced anywhere in the console. boto3.amazonaws.com/v1/documenta...
boto3.amazonaws.com
get_table_maintenance_job_status - Boto3 1.35.86 documentationContentsMenuExpandLight modeDark modeAuto light/dark modeClose Menu
010
Burak @buremba.bsky.social · 21/12/2024
Is it slower than classic S3 tables for you? In my tests, it was about to be the same but I used EC2.
110
Burak @buremba.bsky.social · 19/12/2024
This is cool but how about the mission to decrease the amount of Avro files in the world? 😝
060
Burak @buremba.bsky.social · 17/12/2024
32 bytes for a number is crazy
010
Burak @buremba.bsky.social · 17/12/2024
@felixscherz.bsky.social already created the draft PR for pyiceberg here: github.com/apache/icebe... I think the right way would be S3 adopting Iceberg REST protocol natively but this would be the alternative.
github.com
Support for S3 catalog to work with S3 Tables · Issue #1404 · apache/iceberg-python
Feature Request / Improvement Amazon S3 tables have being launched, see this, and looks like that S3 tables have a managed iceberg catalog. Based on https://github.com/awslabs/s3-tables-catalog it ...
020
Burak @buremba.bsky.social · 17/12/2024
Assuming you refer to S3Tables, I believe you already know it better than me :)
000
Burak @buremba.bsky.social · 17/12/2024
If you are operating with a single table and know the path of a Iceberg metadata file, you don’t need a catalog. Here is an example: duckdb.org/docs/extensi... Catalog is for features such as time travel, atomic update/merge/insert.
duckdb.org
Iceberg Extension
The iceberg extension is a loadable extension that implements support for the Apache Iceberg format. Installing and Loading To install and load the iceberg extension, run: INSTALL iceberg; LOAD iceber...
120
Burak @buremba.bsky.social · 17/12/2024
Great article! While it focuses on first party data sharing, IMO third party data sharing is also pretty interesting with the new “Clean Room” concept. Thanks to the differential privacy and confidential computing, now you can share confidential data and make sure your collaborator doesn’t abuse it.
030
Burak @buremba.bsky.social · 17/12/2024
We are all thankful for SAP HANA in DuckDB community
350
Burak @buremba.bsky.social · 15/12/2024
It took 5 minutes to create this dashboard with @rilldata.com's AI auto-generate features. Impressive @medriscoll.com!
0133
Burak @buremba.bsky.social · 15/12/2024
id
000
Burak @buremba.bsky.social · 15/12/2024
I agree but finding the right prompt to guide AI requires different mental model compared to actually trying to fix the code for me. Maybe I will become a better “prompt engineer” over time but the constant feedback loop makes me less efficient because of context switching.
130