Data Engineer · Riyadh

Yazeed

Almuhlaki

I build pipelines that move data from raw to serving. Ingestion, modeling, and quality gates for national-scale Saudi data.

Download CV
Core Skills
SQL
Python
ETL / ELT
Data Modeling
AWS
PostgreSQL
About

Most data never reaches the people who need it.

My work spans lakehouse architectures on AWS, dimensional warehouse modeling on Redshift and Azure Synapse, and raster pipelines that produce governorate-level outputs. I focus on what happens between the source and the serving layer.

Skills
Data Engineering
SQLPythonETL / ELTData ModelingData Warehouse DesignData Lake DesignParquetPartitioningIncremental LoadsIdempotent LoadsData Quality Checks
Cloud and Storage
AWSAzurePostgreSQLDuckDBS3
Geospatial Engineering
PostGISGeoParquetSpatial JoinsCoordinate SystemsRaster Processing
Selected Work

STEDI Lakehouse on AWS

Problem
Raw semi-structured sensor and app data arrived from three separate sources with no curated layer to query.
Source
Semi-structured JSON from a customer website, a mobile app, and IoT devices.
Pipeline
Landed raw JSON in S3, cataloged and transformed it through AWS Glue into a landing, trusted, and curated zone structure, then served the curated tables to Athena for querying.
Quality
Records filtered at the trusted zone so that only consented customer data flows into curated tables.
AWSGlueS3AthenaPythonSparkETLLakehouse

Redshift Data Warehouse

Problem
Raw JSON event logs in S3 were not queryable for analytics without a modeled warehouse.
Source
Raw JSON log and song metadata files stored in S3.
Pipeline
Extracted JSON from S3 into Redshift staging tables, then transformed into a star schema with a songplays fact table and users, songs, artists, and time dimensions using SQL joins and filters.
AWSRedshiftSQLPythonStar SchemaETLData Modeling

Azure Bike Share Data Warehouse

Problem
Operational OLTP records for trips, riders, stations, and payments were not shaped for analytical queries.
Source
Bike share operational data covering trip records, rider profiles, station locations, and payment transactions.
Pipeline
Modeled the operational data into a galaxy schema with separate fact tables for trips and payments sharing conformed dimensions, targeting an Azure Synapse analytics environment.
AzureSynapseSQLData ModelingGalaxy SchemaDimensional Modeling

KSA Dust Health Risk

Problem
Aerosol and population data existed in incompatible raster grids with no governorate-level output.
Source
NASA MERRA-2 aerosol optical depth rasters and WorldPop 2020 population rasters, covering 147 Saudi governorates.
Pipeline
Ingested both raster sources, aligned them to a common grid, computed a weighted Dust Respiratory Exposure Index per cell, then aggregated to governorate boundaries as the serving layer.
PythonRaster ProcessingGeospatialETL
Certifications
Spatial Data Science Bootcamp
Tuwaiq Academy
2026
Data Analyst
Udacity
2026
Programming for Data Science with Python
Udacity
2025
Linear Algebra for Machine Learning and Data Science
DeepLearning.AI
2025
Geospatial Artificial Intelligence (GeoAI)
GeoFusion
2026
Contact

Let's connect.

Open to data engineering internships and roles in Riyadh.

GitHubLinkedInEmail