data-engineer-handbook
This is a repo with links to everything you'd ever want to learn about data engineering
Implement data warehouse schemas with star and snowflake models. Learn when to use denormalized dimensions for speed or normalized structures for efficiency and data integrity.
How to Build an LLM-Powered SQL Query Engine with LangChainLearn to build an LLM-powered SQL query engine with LangChain. This guide covers prompt engineering, SQL validation, and secure database execution for natural language to SQL conversion.
How to Design Modern OLAP Architectures with ClickHouse and Apache Druid: A Complete Engineering GuideDesign modern OLAP architectures with ClickHouse and Apache Druid. Learn to combine batch and streaming analytics for petabyte-scale data processing using Kafka and schema registries.
Best Practices for Data Cleaning Using Pandas: A Data Engineer Handbook GuideMaster pandas data cleaning with a five-step workflow. Learn best practices for production-ready datasets, preventing errors and data leakage. Your data engineer handbook guide.
How to Build Data Engineering Projects on Databricks and Azure: An End-to-End Architecture GuideBuild robust data engineering projects on Databricks and Azure. This guide details an end-to-end architecture from ingestion to secure storage and transformation, optimizing your data pipelines.
How to Implement Incremental Slowly Changing Dimensions Type 2 Queries in SQLMaster incremental Slowly Changing Dimensions Type 2 SQL queries. Learn to efficiently update dimension tables using CTEs and UNNEST for accurate historical and current data.
How to Design Graph-Based Data Models for Player-Game Relationships: A SQL Implementation GuideDesign flexible graph-based data models for player-game relationships with SQL. Implement property graphs for efficient player analysis, career tracking, and co-play statistics.
How to Build Data Pipelines with Apache Airflow Orchestration: A Complete GuideLearn how to build data pipelines with Apache Airflow orchestration. This guide details Python-based DAGs, task dependencies, and monitoring for reliable data workflows.
How to Implement Data Deduplication Strategies for Microbatch ProcessingLearn how to implement data deduplication strategies for microbatch processing using stateful filtering, window functions, and idempotent writes for exactly-once semantics.
How to Set Up Docker Containers for Local Spark Development EnvironmentsSet up local Spark development environments with Docker. Clone the DataExpert-io repository and use `make up` for a complete setup including Iceberg, MinIO, and Jupyter Notebook.
How to Implement Data Visualization and Impact Analysis for StakeholdersLearn to implement data visualization and impact analysis using a three layered architecture. Ingest data, transform with dbt, and expose KPIs via REST APIs for stakeholder BI tools.
How to Design KPIs and Run A/B Testing Experiments in Data Engineering: A Complete GuideDesign effective KPIs and run A/B testing experiments in data engineering. Learn to define metrics, implement user assignment, and log conversions for statistical significance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →