Sign in

Xiangpeng Hao

@xiangpeng.systems
1.1K followers 121 following 37 posts

Database/storage Flight/DataFusion/Arrow/Parquet PhD student@UW-Madison xiangpeng.systems

PostsRepliesMedia
Xiangpeng Hao @xiangpeng.systems · 08/06/2026
A system programmer’s guide to LLM inference, hope you like it! blog.xiangpeng.systems/posts/how-to...
blog.xiangpeng.systems
A system programmer’s guide to LLM inference – Xiangpeng’s blog
Let’s build a local LLM inference engine in Rust with no dependencies.
040
Reposted by Xiangpeng Hao
Andrew Lamb @andrewlamb1111.bsky.social · 23/02/2026
Simply applying basic linting rules (like don't compress pages where it doesn't help) reduces parquet files sizes by 5% and decreases decode time by 20%. @xiangpeng.systems shows how in his latest blog blog.xiangpeng.systems/posts/parque...
blog.xiangpeng.systems
parquet-linter: A better Parquet is Parquet itself – Xiangpeng’s blog
Unleash the performance potential of your Parquet files
0192
Xiangpeng Hao @xiangpeng.systems · 04/02/2026
Thank you for sharing, glad to see many people think alike!
010
Xiangpeng Hao @xiangpeng.systems · 02/02/2026
Stop building systems for agents, build systems for human. We need infrastructures to help us holding accountability of agent's code. blog.xiangpeng.systems/posts/stop-b...
blog.xiangpeng.systems
Stop building systems for agents – Xiangpeng’s blog
Build agent systems for human.
172
Xiangpeng Hao @xiangpeng.systems · 29/01/2026
After two years since publication, Bf-Tree is finally open-sourced github.com/microsoft/bf... The GitHub repo just hit the Hacker News front page 🎉🎉🎉
github.com
GitHub - microsoft/bf-tree: Bf-Tree is a modern read-write-optimized concurrent larger-than-memory range index in Rust from MS Research.
Bf-Tree is a modern read-write-optimized concurrent larger-than-memory range index in Rust from MS Research. - microsoft/bf-tree
0110
Xiangpeng Hao @xiangpeng.systems · 03/12/2025
Mitchell is the creator of the github.com/mosure/bevy_... render-pipeline plugin for Bevy, and he’ll dive into real-time radiance-field rendering, GPU data layouts, kernels, profiling, and more.
github.com
GitHub - mosure/bevy_gaussian_splatting: bevy gaussian splatting render pipeline plugin
bevy gaussian splatting render pipeline plugin. Contribute to mosure/bevy_gaussian_splatting development by creating an account on GitHub.
052
Xiangpeng Hao @xiangpeng.systems · 03/11/2025
Appreciate the kind words!
010
Xiangpeng Hao @xiangpeng.systems · 29/10/2025
Nice to see this getting shared! 🙌 Now I’m even more motivated to turn it into a full course.
050
Xiangpeng Hao @xiangpeng.systems · 24/10/2025
Just like other big cities, Madison is getting its own systems talk series. Come join us!
081
Xiangpeng Hao @xiangpeng.systems · 10/09/2025
LiquidCache a distributed pushdown cache for DataFusion, designed to cut down S3 requests for diskless databases. 💻 Code: github.com/XiangpengHao... 📄 Paper (VLDB 2026): github.com/XiangpengHao...
090
Xiangpeng Hao @xiangpeng.systems · 02/09/2025
Thanks you for sharing! slides are here 👉 what-is-liquid-cache.xiangpeng.systems
what-is-liquid-cache.xiangpeng.systems
What is LiquidCache?
020
Xiangpeng Hao @xiangpeng.systems · 01/08/2025
Hey Tyler 👋 welcome back! I'd be happy to chat, I work in the data systems space (database + storage + cloud) from the same group that also studies storage fault!
020
Xiangpeng Hao @xiangpeng.systems · 16/05/2025
Project repo: github.com/XiangpengHao...
github.com
GitHub - XiangpengHao/liquid-cache: 10x lower latency for cloud-native DataFusion
10x lower latency for cloud-native DataFusion. Contribute to XiangpengHao/liquid-cache development by creating an account on GitHub.
060
Xiangpeng Hao @xiangpeng.systems · 16/05/2025
Join my PhD prelim talk next Monday: Data-Aware Caching for Cloud Analytics 🕐 May 19, 1PM CDT 📍 CS2310 or Zoom: uwmadison.zoom.us/j/3081128886
Data-Aware Caching for Cloud Analytics
191
Reposted by Xiangpeng Hao
Andrew Lamb @andrewlamb1111.bsky.social · 04/04/2025
My manifesto on optimizing SQL and DataFrames in query engines (including an explanation of why Apache DataFusion doesn't have a complex join ordering algorithm): www.influxdata.com/blog/optimiz... www.influxdata.com/blog/optimiz...
influxdata.com
Optimizing SQL (and DataFrames) in DataFusion: Part 1
This post reviews what a Query Optimizer is, what it does, and why you need one for SQL and DataFrames. It also describes how industrial Query Optimizers are structured and standard optimization class...
061
Xiangpeng Hao @xiangpeng.systems · 24/03/2025
New blog post: "Build your own S3-Select in 400 lines of Rust" Check it out 😉: blog.xiangpeng.systems/posts/build-...
blog.xiangpeng.systems
Build your own S3-Select in 400 lines of Rust – Xiangpeng’s blog
DataFusion is ALL YOU NEED
0103
Xiangpeng Hao @xiangpeng.systems · 14/03/2025
Credit goes to github.com/excalidraw/e... for making it easy😉
github.com
GitHub - excalidraw/excalidraw: Virtual whiteboard for sketching hand-drawn like diagrams
Virtual whiteboard for sketching hand-drawn like diagrams - excalidraw/excalidraw
100
Xiangpeng Hao @xiangpeng.systems · 13/03/2025
Here's the PR: github.com/apache/arrow...
github.com
Experimental parquet decoder with first-class selection pushdown support by XiangpengHao · Pull Request #6921 · apache/arrow-rs
Which issue does this PR close? Many long lasting issues in DataFusion and Parquet. Note that this PR may or may not close these issues, but (imo) it will be the foundation to future more optimiza...
010
Xiangpeng Hao @xiangpeng.systems · 13/03/2025
I submitted a PR that cuts average ClickBench latency by 15% for DataFusion! But reviewing it wasn't straightforward due to the nature of complex performance tuning dynamics, so I made a blog post to explain why it works -- check it out: blog.xiangpeng.systems/posts/parque...
blog.xiangpeng.systems
Efficient Filter Pushdown in Parquet – Xiangpeng’s blog
How to implement efficient filter pushdown in Parquet readers and why it’s challenging in practice.
2152
Reposted by Xiangpeng Hao
Ao Li @aoli.al · 12/03/2025
We are excited to share Fray Debugger (aoli.al/blogs/deadlo...), an IntelliJ plugin that allows you to control concurrent execution deterministically! We have translated the Deadlock Empire (deadlockempire.github.io) into Java to demonstrate how to use Fray Debugger.
aoli.al
Evil Scheduler: Mastering Concurrency Through Interactive Debugging – Ao Li
TLDR Watch the video below to see how Fray debugger works! I enjoy the concept of Deadlock Empire, an interactive game that teaches the semantics of locks and other concurrency primitives. The core id...
131
Xiangpeng Hao @xiangpeng.systems · 10/03/2025
Meanwhile, as a PhD student, I still feel frustrated comparing my systems to many ideas that seem novel but lack practical impact. That said, I find “feet on the ground, head in the clouds” research very inspiring -- it’s probably what keeps me motivated to stay in academia.
010
Xiangpeng Hao @xiangpeng.systems · 10/03/2025
Thanks for the insightful points, Marc! I totally agree that academia is important in many areas. I'm planning a follow-up post discussing the kinds of research that are impactful and beneficial to people, and your examples strongly resonate with what I have in mind!
010
Xiangpeng Hao @xiangpeng.systems · 10/03/2025
Thanks for sharing your perspective! It’s always helpful to hear insights from folks who’ve spent time in industry. There’s definitely room for academia to evolve, and I’m hopeful it will :)
010
Reposted by Xiangpeng Hao
Xuanwo @xuanwo.io · 10/03/2025
@xiangpeng.systems shared a great post about system researchers. I wrote a comment on it and would like to share some thoughts here and offer complementary ideas. In short: build paper with open source. xuanwo.io/links/2025/0...
082
Xiangpeng Hao @xiangpeng.systems · 10/03/2025
Wrote a blog post reflecting my thoughts on DeepSeek, NSF funding and system research communities in general. Apologies for the bold claims -- hope they can invite some discussions. blog.xiangpeng.systems/posts/system...
blog.xiangpeng.systems
Where are we now, system researchers? – Xiangpeng’s blog
2112
Xiangpeng Hao @xiangpeng.systems · 22/02/2025
Compile to WASM is a very interesting idea! I think Fray at some point explored this a bit, not sure about the current status
000
Xiangpeng Hao @xiangpeng.systems · 22/02/2025
Current approaches need to replace std locks with framework provided locks, like the ones in shuttle: docs.rs/shuttle/late... I think binary instrumentation like the one in this paper is possible, but I'm not an expert on this. www.microsoft.com/en-us/resear...
docs.rs
shuttle - Rust
Shuttle is a library for testing concurrent Rust code, heavily inspired by Loom.
000
Xiangpeng Hao @xiangpeng.systems · 22/02/2025
I heard from Fray dev that it is getting a built-in interactive debugger, which visualizes what each threads is doing at a given moment, I can see it to be incredibly useful!
010
Xiangpeng Hao @xiangpeng.systems · 22/02/2025
Yes, Loom and shuttle: github.com/awslabs/shut... They are incredibly useful at identifying and reproducing bugs, but I find it quite hard to use them with a debugger, as lldb needs frequently jump to different stacks and I soon lost track of what's going on...
github.com
GitHub - awslabs/shuttle: Shuttle is a library for testing concurrent Rust code
Shuttle is a library for testing concurrent Rust code - awslabs/shuttle
010
Xiangpeng Hao @xiangpeng.systems · 22/02/2025
Checkout the underneath framework: github.com/cmu-pasta/fray Looking forward to a future Rust support😉
github.com
GitHub - cmu-pasta/fray: A controlled concurrency testing framework for the JVM
A controlled concurrency testing framework for the JVM - cmu-pasta/fray
390
Xiangpeng Hao @xiangpeng.systems · 24/11/2024
It uses Gemini free tier API to translate natural language to SQL: ai.google.dev/pricing#1_5f...
ai.google.dev
Gemini API pricing  |  Google AI for Developers
The Gemini API for developers offers a robust free tier and flexible pricing as you scale.
000
Xiangpeng Hao @xiangpeng.systems · 24/11/2024
Reading S3 files (through OpenDAL) is planned for next weekend :-)
020
Xiangpeng Hao @xiangpeng.systems · 24/11/2024
My weekend project now comes with AI super power! Now you can explore Parquet data with natural language! parquet-viewer.haoxp.xyz
221
Xiangpeng Hao @xiangpeng.systems · 22/11/2024
I helped on the string view part, along with many others!
120
Xiangpeng Hao @xiangpeng.systems · 21/11/2024
This is amazing -- an open source query engine build on open standard is now the fastest, and it is in Rust! datafusion.apache.org/blog/2024/11...
datafusion.apache.org
Apache DataFusion is now the fastest single node engine for querying Apache Parquet files
<!–
2324
Reposted by Xiangpeng Hao
Alex Miller @alexmillerdb.bsky.social · 20/11/2024
New blog post on the fun new hardware advancements which databases can leverage for great gains, and why the cloud means it doesn't matter that they exist. 🫠 transactional.blog/b...
35418
Reposted by Xiangpeng Hao
Alex Miller @alexmillerdb.bsky.social · 06/11/2024
My shoddy ASCII art about writing data to disk was surprisingly popular, so I finished off an old set of notes talking more comprehensively about durably writing data to disk and added a better version of the diagram. transactional.blog/h...
26311
Xiangpeng Hao @xiangpeng.systems · 04/11/2024
Checkout my weekend project on Parquet explorer: xiangpenghao.github.io/parquet-expl... It compiles Rust Parquet into WebAssembly, allowing you to explore the structure of Parquet files directly in your browser!
040
Xiangpeng Hao @xiangpeng.systems · 28/10/2024
Thanks for the heads-up! I’ll reach out to them when I get there😀
010
Xiangpeng Hao @xiangpeng.systems · 28/10/2024
Will try SlateDB!
110
Xiangpeng Hao @xiangpeng.systems · 28/10/2024
Seriously working on it!
120
Xiangpeng Hao @xiangpeng.systems · 28/10/2024
New blog post on caching in DataFusion! See how my research is advancing DataFusion’s capabilities and what’s next: blog.haoxp.xyz/posts/cachin...
1265
Xiangpeng Hao @xiangpeng.systems · 25/10/2024
DataFusion implements one of the most advanced Parquet readers🚀, checkout how: blog.haoxp.xyz/posts/parque...
How DataFusion prune Parquet files based on its metadata can query semantics
1213