We are looking for a Tech Ops / SRE Manager to own production stability, incident response, monitoring, uptime, root cause analysis, and reliability operations across our SaaS platforms. This role will ensure production issues are handled quickly, recurring failures are reduced, and reliability standards are implemented across AWS and GCP environments.
Key Responsibilities:
- Own production support and reliability operations for SaaS platforms.
- Define, monitor, and improve uptime, SLAs, SLOs, SLIs, MTTR, and incident trends.
- Manage incident response, escalation, communication, and post-incident reviews.
- Ensure proper root cause analysis and corrective action tracking.
- Build and maintain monitoring, alerting, and operational dashboards.
- Coordinate with Engineering, Cloud, InfoSec, and Product teams during incidents and releases.
- Create and maintain runbooks, escalation matrices, and production support processes.
- Identify recurring production issues and ensure permanent fixes are implemented.
- Support disaster recovery testing, backup validation, and business continuity readiness.
- Improve system reliability through automation, health checks, and preventive controls.
Required Experience:
- 8-12 years of experience in Tech Ops, Production Support, Infrastructure, SRE, or Platform Operations.
- 3 5 years in a lead or manager role.
- Mandatory experience supporting SaaS or cloud-based production systems.
- Strong experience with incident management, RCA, monitoring, alerting, and production operations.
- Experience with AWS, GCP, Kubernetes, Docker, and modern observability tools.
- Prior experience in fintech, payments, banking, compliance SaaS, or regulated environments is strongly preferred.
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
