Sign in

Pipeline To Insights

@pipeline2insights.bsky.social
45 followers 42 following 24 posts

P2I shares knowledge, tutorials, and experiences to help others grow in the data industry. We focus on learning, collaboration, and making complex data/AI topics easy to understand. Substack: pipeline2insights.substack.com

PostsRepliesMedia
Pipeline To Insights @pipeline2insights.bsky.social · 19/09/2026
As a data engineer, you probably already: Gather requirements from people who don't speak in specs Work with systems you didn't design Build quickly when requirements are unclear Translate between business problems and technical systems Those aren't just DE skills. They're core FDE skills, too.
pipeline2insights.substack.com
From Data Engineer to Forward Deployed Engineer
What data engineers already know, what they need to learn, and how to practise it in your current job.
021
Pipeline To Insights @pipeline2insights.bsky.social · 08/03/2026
What does a good GitHub portfolio actually look like? What kind of projects show real skills? How do you go beyond tutorials? And how should you explain what you’ve built? @ivanovyordan.com Answer these questions here. #databs #datasky #dataengineering
substack.com
The Data Engineer’s GitHub Portfolio (2026 Edition)
If you want the interview, your GitHub must prove you can design systems, handle infrastructure, and make engineering trade-offs
042
Reposted by Pipeline To Insights
Pipeline To Insights @pipeline2insights.bsky.social · 27/12/2025
WAP, AWAP, TAP, or Signal Tables, which one protects your production data best? I put together a deep dive into data quality design patterns used by modern data teams, including: - How each pattern works - Real implementation examples - Cost vs. safety trade-offs #datasky #dataBS #dataengineering
open.substack.com
Data Quality Design Patterns
Overview of WAP, AWAP, TAP, and More with Implementation Examples
021
Pipeline To Insights @pipeline2insights.bsky.social · 27/12/2025
WAP, AWAP, TAP, or Signal Tables, which one protects your production data best? I put together a deep dive into data quality design patterns used by modern data teams, including: - How each pattern works - Real implementation examples - Cost vs. safety trade-offs #datasky #dataBS #dataengineering
open.substack.com
Data Quality Design Patterns
Overview of WAP, AWAP, TAP, and More with Implementation Examples
021
Pipeline To Insights @pipeline2insights.bsky.social · 28/11/2025
How to Learn Data Engineering in 2026 Thinking about learning Data Engineering or transitioning into a Data Engineering role? Curious how AI advancements have changed the field, but still unsure where to start, what to focus on, or how to get going? This guide can help #databs #dataengineering
pipeline2insights.substack.com
How to Learn Data Engineering in 2026
The essential fundamentals you need to know and how to leverage AI tools to stay ahead of the curve in Data Engineering.
021
Reposted by Pipeline To Insights
Pipeline To Insights @pipeline2insights.bsky.social · 28/03/2025
Mistakes help us grow, whether ours or others'. In data, small errors can cause big issues like broken pipelines or high costs. These lessons aren’t just for data engineers, they benefit anyone working with data. #datasky #dataBS
pipeline2insights.substack.com
Common Data Engineering mistakes and how to avoid them
From broken pipelines to unexpected cloud costs, learn from real-world mistakes and lessons to level up your data engineering skills.
031
Pipeline To Insights @pipeline2insights.bsky.social · 28/03/2025
Mistakes help us grow, whether ours or others'. In data, small errors can cause big issues like broken pipelines or high costs. These lessons aren’t just for data engineers, they benefit anyone working with data. #datasky #dataBS
pipeline2insights.substack.com
Common Data Engineering mistakes and how to avoid them
From broken pipelines to unexpected cloud costs, learn from real-world mistakes and lessons to level up your data engineering skills.
031
Pipeline To Insights @pipeline2insights.bsky.social · 15/03/2025
Data Compression in SQL In this post, we’ll explore: - What is Data Compression? - Benefits of Data Compression in SQL. - Types of Compression in SQL. - Comparison of SQL Compression Techniques. - When to Use Compression and When to Avoid. #databs #datasky
pipeline2insights.substack.com
Data Compression in SQL
How to Store More and Query Faster in SQL
132
Pipeline To Insights @pipeline2insights.bsky.social · 04/02/2025
What is Zero-ETL ? What it isn’t ? Read more here : open.substack.com/pub/pipeline... #dataBS #datasky
020
Pipeline To Insights @pipeline2insights.bsky.social · 18/01/2025
In this post, we will cover: - What is Data Vault 2.0 - Common Data Vault interview questions - Bridging the fundamental models to Data Vault - A Case Study: Converting a Dimensional Model to a Data Vault #dataBS #datasky
open.substack.com
Week 6/33: Data Modelling for Data Engineering Interviews (Part #3)
What is Data Vault 2.0 and its role in Data Engineering Interviews
020
Pipeline To Insights @pipeline2insights.bsky.social · 11/01/2025
Explore key data modelling concepts for databases and data warehouses, including: - Normalisation vs. Denormalisation - 3NF - Dimensional Modeling - Star vs. Snowflake Schema comparisons. #databs #datasky
open.substack.com
Data Modelling Fundamentals: Normalisation, 3NF and Dimensional Modelling
Normalisation, 3NF, and dimensional modelling, with insights into Star and Snowflake schemas for efficient database and warehouse design
030
Pipeline To Insights @pipeline2insights.bsky.social · 04/01/2025
11 Storage Formats for Data Engineers Efficient data starts with the right storage format. Explore 11 formats every data engineer should know to match workloads and scale seamlessly. Highlights: Row & Columnar, Key-Value, Document, Graph, Time-Series, Hybrid. #databs #datasky
pipeline2insights.substack.com
11 Storage Formats for Data Engineers
How to leverage storage formats for efficient and scalable data systems
053
Pipeline To Insights @pipeline2insights.bsky.social · 19/12/2024
Data ingestion with dlt and Dagster: An end-to-end pipeline tutorial: Curious like us to see what people are sharing with #dataBS and #datasky? Check out this post to learn how to do it using dlt!" @matthausk.bsky.social @datateam.bsky.social @hgeren.bsky.social @hopefanhe.bsky.social #dlt
open.substack.com
Data ingestion with dlt and Dagster: An end-to-end pipeline tutorial
Ingest Data from Bluesky API to AWS S3 Using dlt and deploy it on Dagster in Just 15 Minutes.
0111
Pipeline To Insights @pipeline2insights.bsky.social · 17/12/2024
Week 6 of '100 Days of SQL Optimisation': Focused on DuckDB, leveraging columnar storage, sorted data, temp tables, Parquet, and optimal data types to boost efficiency. See how in-memory execution and smart structures enhance query performance! @duckdb.org #dataBS #datasky #duckdb
open.substack.com
Week #6: 100 Days of SQL Optimisation
Exploring DuckDB and Its Capabilities
141
Pipeline To Insights @pipeline2insights.bsky.social · 11/12/2024
Protect data with encryption, access controls, and monitoring. Safeguard credentials, apply least privilege, store only essential sensitive data, and ensure cloud security with IAM and encryption. Build a culture of security beyond compliance. @joereis.bsky.social
open.substack.com
Security Fundamentals for Data Engineers
The Role of Security in the Data Engineering Lifecycle
040
Pipeline To Insights @pipeline2insights.bsky.social · 08/12/2024
We are starting a 32-week Data Engineering Interview Guide program, covering everything from fundamentals to advanced topics, with sessions every Saturday. Do you think we're missing any critical topics? We're curious about your opinions😊 #dataBS #datasky
open.substack.com
Week 0/32 - A Comprehensive Data Engineering Interview Preparation Guide
Join us every Saturday on This New Journey
063
Pipeline To Insights @pipeline2insights.bsky.social · 04/12/2024
As a Data Engineer, understanding the data storage lifecycle and data retention policies is critical for designing efficient, cost-effective, and compliant data systems. @joereis.bsky.social #dataBS #datasky substack.com/@pipeline2in...
082
Pipeline To Insights @pipeline2insights.bsky.social · 03/12/2024
In our new post, we've covered 10 of the most popular data pipeline design patterns. We’d love to hear your thoughts. For more details, please check out the full post created by (@hgeren.bsky.social and @hopefanhe.bsky.social ): open.substack.com/pub/pipeline... #dataBS #datasky
open.substack.com
10 Pipeline Design Patterns for Data Engineers
How to leverage Design Patterns for scalable and efficient data pipelines
042
Pipeline To Insights @pipeline2insights.bsky.social · 01/12/2024
Discover how dlt simplifies data ingestion. Learn its origins and real-world use cases. Follow a step-by-step guide to build your first pipeline and join the growing dlt community! @matthausk.bsky.social @datateam.bsky.social @hgeren.bsky.social @hopefanhe.bsky.social #dataBS #datasky
open.substack.com
Introduction to data load tool (dlt): A Python Library for Simple Data Ingestion
Discover the basics of dlt and its role in modern data engineering workflows
293
Pipeline To Insights @pipeline2insights.bsky.social · 28/11/2024
Hi, wishing everyone a great Thanksgiving! Recently we wrote about how SQL queries are executed behind the scenes. If you are interested, check out our post: open.substack.com/pub/pipeline... #dataBS #datasky
062
Reposted by Pipeline To Insights
Hasan Geren @hgeren.bsky.social · 06/11/2024
Just joined and heard #dataBS and #datasky are where the cool kids hang. Wanted to introduce our blog where we regularly write about Data Engineering concepts, news, and tools. pipeline2insights.substack.com
2153
Pipeline To Insights @pipeline2insights.bsky.social · 26/11/2024
Storage is at the heart of Data Engineering. In this post, we explore the hierarchy of data storage from the ground up, drawing inspiration from Fundamentals of Data Engineering by @joereis.bsky.social and Matt Housley, as well as insights from the DE Professionals on Coursera. #dataBS #datasky
open.substack.com
Storage Fundamentals For Data Engineers
Why organised and durable storage is the cornerstone of Data Engineering?
3162