We are seeking an experienced engineering leader to build and lead our Quality and Reliability organization. AI-native development practices have significantly increased our engineering velocity, and this role owns the other half of that equation: ensuring the software we ship is reliable, performant, and trustworthy as we scale. The ideal candidate will have built a reliability or quality engineering practice from the ground up and will bring deep quality expertise to an established QA team. This role is an exciting opportunity to stand up a new Site Reliability Engineering function, take formal ownership of incident management across the company, and lead a talented quality organization. The role reports to the VP of Application Development as a peer of our other engineering guild leads.
What You’ll Do:
- Build Canoe’s Site Reliability Engineering practice from the ground up, including SLOs and error budgets on critical services, observability and alerting standards, production readiness reviews, and on-call design.
- Own incident management across end to end, including process, tooling, escalation paths, retrospective quality, and follow-through on corrective actions.
- Lead our Quality Engineering organization and build lasting in-house quality leadership and expertise.
- Define a risk-based quality strategy, including quality gates in our CI/CD pipelines, and continue our shift from manual verification toward automation-first, AI-assisted testing and specification-driven development.
- Define and report executive-level KPIs for quality and reliability, such as escape rates, incident trends, and error-budget consumption, and use them to inform engineering priorities across teams.
- Serve as Technical Governance Owner for our testing and incident management topics, bringing a reliability, testability, and operability perspective to technical decisions in partnership with our architects.
What We’re Looking For:
- Minimum of 10 years of software engineering experience, including 4 years leading quality, reliability, or platform engineering teams.
- Experience building an SRE, reliability, or quality engineering practice from scratch, rather than inheriting a mature one.
- Experience formally owning incident management for a company or major platform, from incident response through high-quality retrospectives and follow-through.
- Deep quality engineering background, including ownership of quality strategy and experience leading distributed or offshore QA teams.
- Demonstrated expertise applying LLMs and agentic tools such as Claude Code to real engineering workflows, with a concrete point of view on how AI reshapes testing and verification.
- Fluency with quality and reliability metrics such as escape rate, flake rate, MTTD/MTTR, and SLO attainment, including experience reporting them at the executive level.
- Excellent leadership, communication, and interpersonal skills with geographically distributed teams across multiple time zones.
- Ability to work effectively with ambiguous or missing information and adapt quickly in a startup-like environment.
- Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent experience).
Preferred Experience
- Hands-on experience with our core technologies or similar stacks: Python, TypeScript, PHP/Laravel, Kafka, AWS, Terraform, PostgreSQL, Redis, ElasticSearch, Datadog.
- Fintech or other environments where data integrity and client trust are the product.
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
