Sign in

Kostas Pardalis

@cpard.bsky.social
1.1K followers 23 following 79 posts

Building typedef.ai | host @ techontherocks.show | Done some cool stuff with trinodb | ex-RudderStack | previously CEO @ Blendo

PostsRepliesMedia
Kostas Pardalis @cpard.bsky.social · 18/10/2025
That’s a different concept though, right? As you said, here you have a proxy and you pick a different query engine over the same storage. I think using the term federation in this case will confuse people. I can see how this pattern can work.
020
Kostas Pardalis @cpard.bsky.social · 18/10/2025
Full scans on different data sources that then need to be joined and a much closer to ETL workload. This will kill every federated query engine. Plus what do you do when you have different semantic between different query engines? Let’s say how you handle decimal overflows.
030
Kostas Pardalis @cpard.bsky.social · 18/10/2025
Oh no. Trino tried tried to do that. You really can’t do it. The problem with federated queries is that they work well when you can push computation down to the query engine you federate at and get out a highly reduced dataset. That’s not the case with ETL though.
230
Kostas Pardalis @cpard.bsky.social · 09/09/2025
fenic 0.4.0 brings fenic and its expressive API for working with data, to agents. With tooling becoming a catalog artifact, MCP servers and toolsets being available with just a cli command you can turn any data set you have into well curated context for your agents. check it out!
000
Reposted by Kostas Pardalis
Tech on the Rocks @totrrocks.bsky.social · 08/09/2025
New episode: chatting with bauplan founders Jacopo Tagliabue and Ciro Greco on shipping AI with real-world data constraints. Why listen 1. Data pipelines determine model effectiveness, far more than most teams admit.
143
Kostas Pardalis @cpard.bsky.social · 20/08/2025
030
Kostas Pardalis @cpard.bsky.social · 07/08/2025
7/7 Give it a try, ⭐ the repo, open issues and join the community! 👉 t.co/zDj8rBO5Ce
t.co
https://github.com/typedef-ai/fenic
010
Kostas Pardalis @cpard.bsky.social · 07/08/2025
6/7 Performance & DX Rust optimizations plus leaner default configs deliver performance gains and a frictionless setup experience. so you spend less time tuning and more time building.
100
Kostas Pardalis @cpard.bsky.social · 07/08/2025
5/7 New Functions & Models Access built-in summarization, new semantic APIs, and multiple embedding providers (e.g. Cohere, Google Gemini) out of the box. This broadens your toolkit, so you can prototype and productionize a wider range of AI workflows quickly.
100
Kostas Pardalis @cpard.bsky.social · 07/08/2025
4/7 Composable Pipelines Save intermediate DataFrames as persistent views in the fenic catalog. Reuse and chain complex transformations across jobs without rewriting or rerunning upstream logic, accelerating iteration and collaboration.
100
Kostas Pardalis @cpard.bsky.social · 07/08/2025
3/7 Typed Semantics Define your output schema once with Pydantic and get back validated, strongly typed results. This enforces consistency, surfaces errors early, and eliminates manual parsing of LLM responses.
110
Kostas Pardalis @cpard.bsky.social · 07/08/2025
2/7 Robust Fuzzy Text Matching Ground LLM outputs against your existing data: record linkage, deduplication, and typo-tolerant joins become first-class operations. This improves precision in extraction pipelines and slashes downstream error rates.
100
Kostas Pardalis @cpard.bsky.social · 07/08/2025
Here's a bit more information on each of the new 🦊 fenic 🦊 features. 1/7 🧵 Dynamic Templating Turn any column struct or array into a live prompt fragment. No more string concatenation hacks. You get per row, data driven prompts with minimal code, boosting relevance and reducing boilerplate.
100
Kostas Pardalis @cpard.bsky.social · 06/08/2025
check the repo for more information and give it a try! github.com/typedef-ai/f...
github.com
GitHub - typedef-ai/fenic: Build reliable AI and agentic applications with DataFrames
Build reliable AI and agentic applications with DataFrames - typedef-ai/fenic
011
Kostas Pardalis @cpard.bsky.social · 06/08/2025
fenic v0.3.0 is out and it's a release I'm really excited about! Here are a few of the things that this release is introducing. Jinja as a column function Robust Fuzzy Text Matching Full Pydantic support in all semantic operators Persistent views More Functions & Models Perf & DX improvements
Using Jinja templates to dynamically create prompts for semantic filtering in fenic.
141
Reposted by Kostas Pardalis
Tech on the Rocks @totrrocks.bsky.social · 28/07/2025
@steveklabnik.com Joined us on an episode where we discussed about Why: • Cargo & friendly errors > benchmarks • 6-week releases > years-long committees • How Rust united Ruby, FP & C++ devs • Next-gen picks and many more! Check the episode on your favorite platform!
041
Kostas Pardalis @cpard.bsky.social · 07/06/2025
Everyone’s heads down on AI these days, but please take a break and soak in some deep systems wisdom from Josh Howards. He’s one of the folks behind R2 at Cloudflare. After all, whatever you build in AI will sit on top of these foundations. check @totrrocks.bsky.social for the episode link.
020
Reposted by Kostas Pardalis
Tech on the Rocks @totrrocks.bsky.social · 16/05/2025
Startups and new products increasingly prioritize serverless models to reduce user friction and accelerate adoption. @philippemnoel.bsky.social from ep.12
011
Reposted by Kostas Pardalis
Tech on the Rocks @totrrocks.bsky.social · 14/05/2025
The value proposition of formal methods becomes clear when dealing with complex distributed transactions involving multiple independent services. Jayaprabhakar(JP) Kadarkarai from ep.5
011
Reposted by Kostas Pardalis
Tech on the Rocks @totrrocks.bsky.social · 14/05/2025
User experience and developer interaction with complex data abstractions remain a significant challenge beyond the technical integration. Nikhil Simha & Varant Zanoyan from ep.2
021
Reposted by Kostas Pardalis
Tech on the Rocks @totrrocks.bsky.social · 13/05/2025
Successful AI developer tools must balance synchronous co-pilot style assistance with asynchronous autonomous agent workflows. @ivanburazin.bsky.social from ep.9
021
Reposted by Kostas Pardalis
Tech on the Rocks @totrrocks.bsky.social · 12/05/2025
Managing AI access and permissions requires careful role-based controls to prevent over-privileged AI actions in enterprise environments. Well said, even before hashtag#MCP was as popular as today. @ivanburazin.bsky.social from ep.9
011
Kostas Pardalis @cpard.bsky.social · 09/05/2025
I had the rare opportunity to sit down and chat with someone who helped shape that story of Splunk, co-founder Erik Swan. There's a lot to learn from him but what inspired me the most is his energy. Even after a success like Splunk, still learning and building listen here @totrrocks.bsky.social
010
Kostas Pardalis @cpard.bsky.social · 25/04/2025
Incremental materialization has stumped the industry for decades. Epsio led by Gilad , is changing that: product-first, real-world incremental views. If real-time data infra matters to you, check out my chat with Gilad on @totrrocks.bsky.social
050
Reposted by Kostas Pardalis
Chris @chris.blue · 17/04/2025
Just came across this! "transaction isolation in the presence of IVM remains underspecified." I was literally talking about this with @frankmcsherry.bsky.social 6 hours ago.
arxiv.org
Streaming Democratized: Ease Across the Latency Spectrum with Delayed View Semantics and Snowflake Dynamic Tables
Streaming data pipelines remain challenging and expensive to build and maintain, despite significant advancements in stronger consistency, event time semantics, and SQL support over the last decade. P...
171
Kostas Pardalis @cpard.bsky.social · 24/03/2025
You should definitely check the project!
000
Kostas Pardalis @cpard.bsky.social · 24/03/2025
Lakekeeper is an open source data catalog built on the Apache Iceberg REST catalog API. If data infrastructure drives you, check out the project and catch Viktor Kessler's insights on the latest @TotrRocks episode!
110
Kostas Pardalis @cpard.bsky.social · 14/03/2025
I'm always excited to chat with @apurvamehta.com about what @responsive.dev is building. Streaming and real time as terms are being constantly reinvented as the market needs change rapidly, and Apurva is one of the best to talk about that. Check the conversation here: techontherocks.show/15
techontherocks.show
Tech on the Rocks | Reinventing Stream Processing: From LinkedIn to Responsive with Apurva Mehta
SummaryIn this episode, Apurva Mehta, co-founder and CEO of Responsive, recounts his extensive journey in stream processing—from his early work at LinkedIn and Confluent to his current venture at R...
052
Kostas Pardalis @cpard.bsky.social · 14/03/2025
We'll be hosting another event at our offices in San Mateo. We want to bring together people who are interested in data and infra, from systems engineers who build data platforms, AI engineers, VCs and everything in between. Connect and have fun while we learn from each other. lu.ma/2hc1qm1v
lu.ma
Peninsula Data Happy Hour · Luma
🔥 An Unmissable Evening of Data & Magic! 🔥 🎉 Back by popular demand, it's time for the March edition of our Peninsula Data Happy Hour! This time we've got…
010
Reposted by Kostas Pardalis
Tech on the Rocks @totrrocks.bsky.social · 20/02/2025
New episode: “Semantic Layers: The Missing Link Between AI and Data” with @jayatillake.bsky.social . We discuss how semantic layers bridge raw data and AI, achieving 100% accuracy for natural language queries, and what’s next for LLM-powered data pipelines. 🎧 techontherocks.show/14 🎧
techontherocks.show
Tech on the Rocks | Semantic Layers: The Missing Link Between AI and Data with David Jayatillake from Cube
In this episode, we chat with David Jayatillake, VP of AI at Cube, about semantic layers and their crucial role in making AI work reliably with data. We explore how semantic layers act as a bridge ...
164
Kostas Pardalis @cpard.bsky.social · 22/01/2025
We will be hosting at our offices another Peninsula Data Happy Hour. This time we will have tacos, amazing people to connect with but also talks from Tomasz Tunguz, Tobiko Data and our own Rohit. Register here 👉 lu.ma/dm24afr0
lu.ma
Peninsula Data Happy Hour · Luma
🎉 Let's kick off 2025 in style! Join us at typedef's cozy San Mateo office (just a short stroll from Caltrain) for an unforgettable evening filled with tacos,…
041
Kostas Pardalis @cpard.bsky.social · 17/01/2025
I always wondered why PG’s text indexing and search is not enough. I finally got some good answers on that and many other questions about search in databases from Phillipe, co-founder and CEO of ParadeDB. Great conversation and many insights on the future of search technologies
010
Kostas Pardalis @cpard.bsky.social · 19/12/2024
Make sure you check the conversation with @davidmytton.social on TotR. He has some amazing stories to tell about observability and why it's so hard, security in a world where AI is turning everyone into a scraper and dev tooling. Check the conversation on your favorite podcast platform!
030
Kostas Pardalis @cpard.bsky.social · 05/12/2024
Viktor and his team have done an amazing job refreshing the concept of the catalog. There are some great ideas in lakekeeper and people should pay attention to it.
040
Kostas Pardalis @cpard.bsky.social · 29/11/2024
That’s some amazing dedication!
010
Kostas Pardalis @cpard.bsky.social · 29/11/2024
That’s impressive Phil. Congrats! How long it took you to grow to that audience size?
110
Kostas Pardalis @cpard.bsky.social · 25/11/2024
oh come on Polaris and Snowflake folks :/ github.com/apache/polar...
github.com
[BUG] Polaris can't use S3 when KMS is enabled · Issue #480 · apache/polaris
Is this a possible security vulnerability? This is NOT a possible security vulnerability Describe the bug When using Polaris with S3 (without KMS), everything is working fine (I can create Iceberg ...
020
Kostas Pardalis @cpard.bsky.social · 24/11/2024
It would awesome to have some examples of how to work with ballooning on nanovms + firecracker. Any guide out there?
100
Kostas Pardalis @cpard.bsky.social · 23/11/2024
Oh nice I’ll check both of them. What made you change to IBM Plex?
100
Kostas Pardalis @cpard.bsky.social · 23/11/2024
Do you have a favorite font that you are using for your terminal?
100
Kostas Pardalis @cpard.bsky.social · 22/11/2024
I had a lot of fun chatting with Roy about his experience during the early days of AWS and being part of the Meta teams training LLMs. Having experienced all this evolution of the industry is rare. You can listen to our conversation here: techontherocks.show/8
techontherocks.show
Tech on the Rocks | Evolving Data Infrastructure for the AI Era: AWS, Meta, and Beyond with Roy Ben-Alta
In this episode, we chat with Roy Ben-Alta, co-founder of Oakminer AI and former director at Meta AI Research, about his fascinating journey through the evolution of data infrastructure and AI. We ...
000
Kostas Pardalis @cpard.bsky.social · 22/11/2024
💯 I can’t wait to see how CXL will change the way we do distributed data processing.
020
Kostas Pardalis @cpard.bsky.social · 22/11/2024
What do yo think about CXL memory?
120
Kostas Pardalis @cpard.bsky.social · 22/11/2024
Why do we hate commas?
120
Kostas Pardalis @cpard.bsky.social · 19/11/2024
The problem is that data engineering is not web development and SQL is not JavaScript. We do need many more people to enter the market but the tooling is not there yet. DBT et.al. tried but wasn't enough. That's a good thing, there's plenty of opportunities out there.
050
Kostas Pardalis @cpard.bsky.social · 19/11/2024
The "analytics engineer" is a result of the commoditization of data warehousing. Suddenly we ended up with huge demand and almost non-existent supply of engineers who could build and run that stuff. Yes, zero interests amplified the effect but it wasn't the cause. DBT seized the opportunity.
110
Kostas Pardalis @cpard.bsky.social · 19/11/2024
We can’t have photos from the 17 miles drive and not have one with the lone cypress
020
Kostas Pardalis @cpard.bsky.social · 17/11/2024
It did a pretty good job
010
Kostas Pardalis @cpard.bsky.social · 17/11/2024
Typedef the keyword was one of the first attempts in improving developer productivity and experience in a PL and at Typedef we have an obsession with that. There’s a lot that can be done in a query engine to improve that and currently missing.
030
Kostas Pardalis @cpard.bsky.social · 17/11/2024
Perplexity knows too much! Need to figure out how it learned so much about Typedef. Typedef is a data platform and a query engine super focused on specific workloads (data engineering related) on data lakes. Perplexity missed the connection between Typedef the platform and the type alias keyword
250