Building reliable, well-modeled data pipelines — from ingestion through transformation to analytics-ready warehouses — using both batch and streaming architectures. Background in BI means I design pipelines with the end report or dashboard in mind, not just the load job.
Data Engineering & Analytics Intern National Telecommunication Institute (NTI) · [Add dates, e.g. April 2026 – July 2026]
- Designed and deployed 3 interactive Power BI dashboards tracking customer-segment KPIs and telecom performance metrics, now used weekly by operations and management teams for data-driven decision-making.
- Queried and analyzed multi-source retail and telecom datasets (100K+ records) using SQL and Python (Pandas, NumPy); identified behavioral patterns that informed campaign targeting strategy adopted for the current quarter.
- Built and executed A/B testing frameworks to measure segment lift, producing validated recommendations directly adopted as campaign baseline by the marketing team.
- Automated recurring KPI reports using Python and Excel (Power Query, PivotTables), reducing manual reporting effort by 4+ hours per cycle.
- Design and build data pipelines using Airflow, Kafka, and PySpark across both batch and real-time workloads
- Model dimensional warehouses (star schema, SCD Type 1/2) in SQL Server, Snowflake, and DuckDB
- Transform and test data with dbt, following Medallion architecture (Bronze/Silver/Gold) patterns
- Deliver the last mile — Power BI dashboards and DAX measures that business users actually rely on
- Think in terms of reliability — documented SLOs, disaster recovery procedures, and access-control design
| Category | Tools |
|---|---|
| Orchestration & Streaming | Apache Airflow · Apache Kafka · Debezium (CDC) |
| Processing & Transformation | PySpark · dbt · Python (Pandas, NumPy) |
| Warehousing & Modeling | SQL Server · Snowflake · DuckDB · Star Schema · SCD Type 1 & 2 · Medallion Architecture |
| BI & Reporting | Power BI · DAX · SSAS (OLAP Cubes) · Excel |
| Infrastructure & Monitoring | Docker · Grafana |
| Languages | Python · T-SQL · DAX |
Airflow · Kafka · dbt · PySpark · Databricks · Snowflake Hybrid data platform that runs on identical code locally (DuckDB/PySpark) or in the cloud (Databricks/ADLS/Snowflake), switched with a single environment variable. Implements Medallion architecture, real-time Kafka streaming, CDC replication, a dbt-built dimensional model, and operational documentation including SLOs, a tested disaster-recovery runbook, and a PAN-tokenization security design.
Snowflake · dbt · Airflow · Kafka End-to-end modern data platform combining CDC ingestion, streaming, and BI-ready transformation layers.
SSIS · SSAS · SQL Server · Power BI Star-schema data warehouse with SCD-managed dimensions and an OLAP cube for enterprise reporting.
SSIS · SSAS · Power BI Three-layer ETL pipeline supporting loyalty program analytics.
SSIS · SSAS · Power BI Full BI pipeline with SCD Type 2 history tracking and OLAP cube modeling.
- Google Data Analytics Professional Certificate
- IBM Data Science Professional Certificate
- Data Analyst in Power BI — DataCamp
- Data Analyst in Python — DataCamp