We are looking for a Datadog Observability & Monitoring Engineer to take ownership of their Observability platform and drive adoption across the development and infrastructure teams within the organization.
Responsibilities:
- The primary initiative is consolidating several monitoring platforms into Datadog as the client’s chosen enterprise observability platform
- This is not a traditional SRE position focused solely on monitoring and dashboards; The goal is to find someone who can own the platform, drive adoption, and build a center of excellence around observability
- Ideal candidate will lead and evolve our observability strategy across cloud and on-premises environments; You will serve as the primary owner and subject matter expert for their Datadog platform – building, scaling, and operating a comprehensive monitoring solution that continuously validates service health (24/7) across business-critical systems, including external websites and key infrastructure components (e.g., firewalls, OpenShift)
- Design and implement end-to-end observability solutions spanning logs, metrics, traces, Service Level Objectives (SLOs), synthetic monitoring, and Real User Monitoring (RUM) to improve reliability, accelerate incident response, and deliver clear visibility into service performance
- Partner closely with application, SRE/DevOps, infrastructure, and security teams and serve as the internal champion and evangelist for Datadog adoption, standards, and best practices; The environment includes an active migration from OpenView to Datadog, with workflows integrating into ServiceNow for incident routing and escalation
Requirements
- 5+ years in Observability, APM, SRE, or Platform Engineering – with at least 2-3 years of hands-on, production-grade Datadog experience.
- Deep expertise across Datadog’s core product suite: APM, Infrastructure Monitoring, Log Management, Synthetics, RUM, SLOs, Dashboards, Monitors, and Alerting.
- Proficiency in both Windows Server and Unix (Linux/Solaris) environments, including agent deployment, service instrumentation, and OS-level performance analysis.
- Strong scripting and automation skills (Python, PowerShell, Bash) with hands-on experience using the Datadog API/SDK and Terraform to manage observability configurations as code.
- Solid understanding of distributed tracing, metrics pipelines, logging standards, and SLO/error budget frameworks within Datadog.
- Experience integrating Datadog with cloud platforms (Azure and AWS) and centralizing cross-environment telemetry.
- Demonstrated ability to reduce alert noise and MTTR through Datadog monitor tuning, correlation, and enrichment strategies.
- BS/BA in Computer Science, Information Systems, Engineering, or equivalent experience.
Nice to Have
- Datadog certifications (e.g., Datadog Fundamentals, APM, or Log Management).
- Experience migrating from legacy monitoring platforms (e.g., OpenView, AppDynamics, Nagios) to Datadog.
- Familiarity with .NET development (C#), including Datadog instrumentation patterns for .NET applications.
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
