Our Client is seeking a Platform Engineer .
Key Responsibilities:
Reliability Architecture & Systems Design:
- Improve the reliability, availability, performance, scalability, and operability of enterprise applications, platforms, and cloud services
- Partner with engineering teams to define practical SLOs, SLIs, error budget thinking, service health measures, and production readiness practices
- Review designs and implementations for failure tolerance, observability, recovery, capacity, and operational readiness
- Influence architecture decisions with reliability, resilience, and operability as first-class concerns
Service Ownership & Operational Excellence:
- Establish and maintain practical service health practices in partnership with application and platform teams
- Strengthen observability using Datadog, OpenTelemetry, Azure Monitor, Application Insights, or comparable tooling
- Participate in the incident lifecycle, including triage, troubleshooting, coordination, root cause analysis, corrective actions, and reliability follow-through
- Champion blameless incident learning and ensure findings are converted into durable reliability improvements
Automation & Engineering Enablement:
- Design and implement automation-first solutions that reduce toil and manual operational work
Build and evolve tooling for:
- Deployment safety and release confidence
- Observability, including metrics, logs, and traces
- Incident detection, response, and recovery
- Use approved AI-assisted engineering tools such as GitHub Copilot and OpenAI Codex responsibly in day-to-day delivery, and share practical patterns that accelerate reliability improvements, troubleshooting, documentation, and automation
Scale, Performance & Resilience:
- Identify and address reliability, scalability, performance, and operability risks across applications, platforms, and shared services
Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or equivalent practical experience
- 7+ years of experience in site reliability engineering, platform engineering, cloud infrastructure, DevOps, systems engineering, or a related technology role
- Experience improving reliability, availability, observability, performance, or operability of production systems in cloud or enterprise environments
- Strong automation skills using at least one scripting language, such as: PowerShell; Bash; Python
Cloud & Platform Technologies:
- Experience with Azure or a comparable cloud environment, including managed services, identity, networking, monitoring, and operational practices
- Experience with Kubernetes, containerized runtimes, managed application platforms, or similar enterprise hosting environments
- Experience with infrastructure-as-code, preferably Terraform, or equivalent tools used to provision and manage cloud infrastructure
- Working knowledge of platform security fundamentals, including identity, access control, secrets management, secure configuration, and policy-driven controls
Observability & Reliability Practices:
- Experience with reliability practices, including Metrics, logs, traces, alerting, dashboards, and service health monitoring; Incident response, root cause analysis, corrective action tracking, and runbooks; SLOs, SLIs, production readiness, resilience, and reliability improvement
- Experience with GitHub or comparable source control workflows that support reliable delivery, including branching strategies, code review practices, workflow automation, and release safety
- Ability to translate incident learning and operational signals into practical reliability improvements
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
