We need someone with 3+ years of experience in SRE, Production Engineering, or Infrastructure roles who has built and owned automation, observability, and tooling systems end-to-end in production. You should be comfortable working across a multi-cloud environment with strong distributed systems instincts and a track record of improving platform reliability and reducing operational burden. Bonus points if you have exposure to GPU/AI-ML infrastructure or accelerated compute workloads.
What you’ll do
- Build and own the observability stack – dashboards, alerts, and distributed tracing using tools like OpenTelemetry, Prometheus, and Grafana – to provide high-granularity visibility into multi-cloud GPU orchestration platform
- Define and implement SLIs and SLOs across API layer and internal orchestration services, partnering with Product and Platform teams to ensure new features are designed for operability from the start
- Develop automation in Python (or Go) to eliminate repetitive operational tasks — from provider API reconciliation to automated health checks and capacity rebalancing
- Maintain and extend Terraform/Pulumi modules and Kubernetes configurations to manage a growing multi-cloud provider footprint
- Participate in on-call rotation, drive rigorous root cause analysis for production incidents, and implement durable fixes to prevent recurrence
- Work directly with the founding engineering team to shape how infrastructure engineering operates as the company scales — this is a greenfield opportunity to build the playbook, not inherit a rigid system
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
