Back to jobs

Site Reliability Engineer

Queen Square RecruitmentYork, United KingdomPosted 2w ago
Site Reliability Engineer
Contractor
Apply on company siteCraft my tailored resume free

Free to start · No card

Job description

Site Reliability Engineer Location: York, Hybrid - 3 days per week onsite Start Day: ASAP Contract Rate: £335 per day inside IR35 Working Pattern: On-call shifts Duration: 6 months initially Role Overview Our client is seeking a Datadog SME to lead observability initiatives across hybrid cloud and microservices environments. You will be responsible for implementing and optimising Datadog solutions, establishing SLOs and SLIs, managing observability pipelines, and supporting engineering teams with monitoring and telemetry best practices. Key Responsibilities Deploy and manage Datadog APM, Infrastructure Monitoring, Logs, RUM, Synthetics, and DBM.Define observability standards, frameworks, and monitoring strategies.Establish SLOs, SLIs, and error budgets to improve service reliability and reduce MTTR.Optimise log management, data ingestion, indexing, and retention.Provision Datadog dashboards, monitors, and configurations using Terraform or OpenTofu.Mentor teams on OpenTelemetry, SRE best practices, and CI/CD integrations. Skills & Experience 10+ years' experience in SRE, DevOps, or Systems Architecture.3+ years' specialist experience with Datadog administration and configuration.Strong knowledge of Datadog Infrastructure Monitoring, APM, Logs, Metrics, Synthetics, and Security Monitoring.Experience with AWS, Azure, or GCP and production Kubernetes environments.Proficiency in Python, Go, Bash, JavaScript, or similar scripting/programming languages.Experience instrumenting applications for APM and troubleshooting complex distributed microservices.Excellent stakeholder management and communication skills.Datadog certifications and experience with Datadog Security products would be advantageous. If you have the relevant skills and experience, please do apply promptly to be considered.