We are seeking a seasoned cloud and platform engineering leader to drive the design, implementation, and evolution of enterprise-scale cloud and AI infrastructure. This role will be responsible for establishing architectural direction, defining platform standards, and enabling scalable, secure, and efficient deployment of modern AI and cloud-native solutions across AWS and hybrid environments.
Key Responsibilities:
- Define and communicate cloud infrastructure and AI platform strategies to technical teams, business leaders, and executive stakeholders
- Provide technical leadership for platform engineering initiatives, guiding teams on architecture, best practices, and operational excellence
- Partner with enterprise architecture and engineering teams to design and validate cloud-native solutions supporting both traditional and AI-driven workloads
- Establish standards for AI platform operations, including model lifecycle management, inference performance, resiliency, governance, and cost optimization
- Administer and optimize large-scale Kubernetes environments, including multi-cluster operations and workload orchestration
- Develop observability frameworks that provide insights into platform health, reliability, performance, and operational efficiency
- Drive automation initiatives that improve deployment speed, platform consistency, and infrastructure scalability
- Promote Infrastructure as Code and GitOps methodologies across engineering teams
Requirements
- 8+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or Cloud Infrastructure roles
- Extensive hands-on experience managing Kubernetes environments at enterprise scale
- Proven experience designing and maintaining CI/CD pipelines and GitOps-based deployment models
- Advanced experience implementing Infrastructure as Code using tools such as Terraform, AWS CDK, or equivalent frameworks
- Experience implementing and managing observability, monitoring, and logging solutions in large enterprise environments
- Experience building secure and reliable AI infrastructure and supporting production AI workloads
- Strong scripting and automation expertise using Python, Bash, or similar languages
- Demonstrated expertise in cloud-native architecture, scalability, resiliency, and operational best practices
- Expertise in AWS networking, including VPC design, routing, load balancing, security controls, and connectivity services
- Deep knowledge of AWS services including compute, networking, storage, identity management, databases, and container platforms
- Strong understanding of enterprise security principles, compliance frameworks, and audit requirements
- Strong Linux administration and troubleshooting skills
- Ability to mentor engineers, influence technical direction, and foster engineering excellence across teams
Preferred Qualifications:
- Experience deploying and supporting generative AI platforms, large language models, and AI-powered applications in production environments
- Experience with service discovery, platform networking, and distributed systems architecture
- Prior experience leading cloud, platform, DevOps, or AI infrastructure teams
- Familiarity with AWS AI services, including Bedrock and related AI orchestration technologies
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
