We are looking for a Principal Production Engineer to join our team.
What you’ll do (Role Expectations)
- Design and implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments
- Drive an “automation-first” culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems
- Implement and maintain sophisticated observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets
- Act as a lead Incident Commander (TDO on-call), develop response playbooks, and conduct deep-dive post-incident analyses
- Partner with Engineering and partner teams to conduct operability reviews
Who You Are (Success Profile)
- You act like an owner with a bias for action and integrity.
- You are a pragmatic builder obsessed with creating, iterating, and shipping.
- You champion simplicity by distilling complex problems into clear, actionable plans.
- You are data-driven, valuing evidence over assumptions.
- You think at scale, building solutions and processes built to last a high-growth global organization.
What We’re Looking for (Minimum Qualifications)
- 10+ years of experience managing reliability, scalability, and availability for large-scale production services
- Deep expertise in programming (e.g., Python, Go, or C/C++)
- Strong background in networking protocols, Linux/RHEL systems, and distributed architecture
- Experience in high-stakes incident management and participation in a 24/7 on-call rotation
- Proficiency in leveraging ITIL frameworks and incident data to drive service maturity through systematic problem management and technical operability reviews
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#bb
Â
