We’re looking for a Senior Data Engineer to join a highly technical Data Engineering team supporting Data Science initiatives. This role will focus on building and maintaining large-scale data pipelines, implementing ETL processes, and making sure data is accurate, reliable, and ready for analytics and machine learning. The ideal candidate is comfortable working with Python/PySpark, Spark DataFrames, SQL, and large datasets and enjoys digging into the data-not simply moving it from one place to another. You’ll work closely with Data Scientists and other engineers on projects involving network telemetry and other complex datasets.
What You’ll Do:
- Build and maintain scalable ETL pipelines using PySpark and Spark
- Develop Spark jobs from scratch and support existing production workflows
- Work primarily with batch processing, including daily and hourly jobs
- Use Spark DataFrames to transform, join, aggregate, and prepare large datasets
- Perform data validation and quality-control checks to ensure accurate downstream analytics
- Partner with Data Scientists to understand data requirements and deliver reliable datasets for modeling and analytics
- Troubleshoot and improve existing data pipelines and processing jobs
- Help ensure data is processed in the right format and according to business and analytical requirements
- Work with AWS-based data processing environments, including EMR and Athena
- Work within an orchestration environment such as Airflow or another workflow orchestration tool
- Collaborate with engineering, Data Science, and platform teams to deliver new and updated data solutions
Requirements
- Strong experience with Python for data engineering/analytics
- Hands-on experience with PySpark, specifically Spark DataFrames
- Experience building batch Spark jobs from scratch
- Experience with data validation and quality checks
- Strong SQL skills, including joins, aggregations, groupings, and window functions
- Experience with a workflow/orchestration tool such as Airflow, Step Functions, or a similar technology
- Strong understanding of ETL and data transformation
- Working knowledge of AWS and cloud-based data processing
Nice to Have:
- Experience with Scala/Spark
- Experience working directly with Data Science teams
- Experience with Databricks
- Experience working with telemetry, network, or other highly technical datasets
- Experience with AWS EMR and Athena
- Familiarity with MLflow or machine-learning model workflows
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
