
Yazhen Li
Changsha, China
Yazhen Li
Expert Data Engineer
Category : Web development
Struggling with slow data processing or unreliable ETL pipelines?
I build and optimize scalable, high-performance data systems that turn large datasets into reliable, actionable insights.
I’m a Data Engineer with 4+ years of professional experience, previously working at Xiaomi (Fortune Global 500) and Shopee. I specialize in designing, building, and optimizing end-to-end data pipelines for analytics and machine learning use cases.
My focus is on solving real data problems: unstable pipelines, slow Spark jobs, poor data quality, and architectures that don’t scale.
⭐ What I do:
✔ End-to-End ETL / ELT Pipelines
Design and implement automated data pipelines using Airflow and Apache Spark (Scala / Python / Java), from ingestion to analytics-ready datasets.
✔ Spark Performance Tuning
Diagnose and optimize slow or failing Spark jobs through memory and CPU tuning, shuffle optimization, and execution plan analysis to reduce runtime and cost.
✔ Big Data Architecture
Design scalable batch and streaming data platforms using Kafka, Flink, Hadoop (HDFS, Hive), and Druid.
✔ Data Quality & Reliability
Implement validation, monitoring, and data quality checks to ensure data accuracy, consistency, and trust.
⭐ Core Skills:
✔ Big Data: Spark, Airflow, Kafka, Flink, Hadoop, Hive, Druid
✔ Programming: Scala, Python, Java, SQL
✔ Databases: HBase, Redis, relational databases
✔ Platforms: Data Lake, Data Warehouse, AWS, GCP
⭐How I work:
I’m outcome-driven and business-focused. I take time to understand goals, communicate clearly, and proactively identify risks and improvements. I value ownership, reliability, and long-term collaboration over short-term delivery.
If you’re looking for a data engineer who can deliver stable systems and measurable results, I’d be glad to work with you.
I build and optimize scalable, high-performance data systems that turn large datasets into reliable, actionable insights.
I’m a Data Engineer with 4+ years of professional experience, previously working at Xiaomi (Fortune Global 500) and Shopee. I specialize in designing, building, and optimizing end-to-end data pipelines for analytics and machine learning use cases.
My focus is on solving real data problems: unstable pipelines, slow Spark jobs, poor data quality, and architectures that don’t scale.
⭐ What I do:
✔ End-to-End ETL / ELT Pipelines
Design and implement automated data pipelines using Airflow and Apache Spark (Scala / Python / Java), from ingestion to analytics-ready datasets.
✔ Spark Performance Tuning
Diagnose and optimize slow or failing Spark jobs through memory and CPU tuning, shuffle optimization, and execution plan analysis to reduce runtime and cost.
✔ Big Data Architecture
Design scalable batch and streaming data platforms using Kafka, Flink, Hadoop (HDFS, Hive), and Druid.
✔ Data Quality & Reliability
Implement validation, monitoring, and data quality checks to ensure data accuracy, consistency, and trust.
⭐ Core Skills:
✔ Big Data: Spark, Airflow, Kafka, Flink, Hadoop, Hive, Druid
✔ Programming: Scala, Python, Java, SQL
✔ Databases: HBase, Redis, relational databases
✔ Platforms: Data Lake, Data Warehouse, AWS, GCP
⭐How I work:
I’m outcome-driven and business-focused. I take time to understand goals, communicate clearly, and proactively identify risks and improvements. I value ownership, reliability, and long-term collaboration over short-term delivery.
If you’re looking for a data engineer who can deliver stable systems and measurable results, I’d be glad to work with you.
Working hours
- Monday:08h00 To 18h00
- Tuesday:08h00 To 18h00
- Wednesday:08h00 To 18h00
- Thursday:08h00 To 18h00
- Friday:08h00 To 18h00
- Saturday:Not available
- Sunday:Not available
- 🇬🇧 English
Please sign in as a customer to give your feedback



