We are looking for a Site Reliability Engineer (SRE) to join a product suite within the Risk Intelligence business.
The successful candidate will have strong technical skills in cloud infrastructure, SRE principles, and operational tooling. This role will work closely with engineering, product, and platform teams to ensure services are reliable, scalable, and secure.
Key Responsibilities
Reliability Engineering & Operations
- Support reliability operations across platforms.
- Monitor and maintain SLOs, SLIs, and error budgets.
- Participate in incident response and post-incident reviews; contribute to root cause analysis and remediation.
- Develop automation for operational tasks, incident response, and compliance.
- Maintain and enhance CI/CD pipelines with integrated testing and deployment automation.
- Implement observability dashboards and alerts using Datadog, OpenTelemetry, and BigPanda.
- Contribute to infrastructure-as-code using Terraform and GitHub Actions.
- Support integration and maintenance of API Gateway and Snowflake data platform.
Service Management & Compliance
- Follow ITIL practices for incident, problem, change, and service request management.
- Use ServiceNow for ticketing, reporting, and workflow automation.
- Ensure runbook accuracy and DR readiness.
- Monitor system performance and cost efficiency.
- Support compliance and audit readiness activities.
Collaboration & Knowledge Sharing
- Work with engineering and product teams to embed reliability into delivery.
- Share technical knowledge through documentation and enablement sessions.
- Participate in global SRE initiatives and cross-regional collaboration.
Person Specification
- Bachelor’s degree in Computer Science, Engineering, or a related technical field or equivalent practical experience.
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
