Back to jobs

Remote | Site Reliability Engineer (SRE & Incident Management) — $70–$110/hour
24-MAGAnywherePosted 1w ago
Site Reliability Engineer
Contractor
Remote
Apply on company siteCraft my tailored resume free
Free to start · No card · 5 credits the moment you sign up
Job description
This a Full Remote job, the offer is available from: New York (USA)
We are sharing a specialised part-time consulting opportunity for experienced Site Reliability Engineering and incident management professionals with strong expertise in production reliability, incident response, operational resilience, root-cause analysis, and service performance.
This role focuses on reviewing professional documents, spreadsheets, and presentation materials related to SRE, reliability engineering, and incident management. Selected experts will assess outputs for technical accuracy, operational rigour, reliability best practices, presentation quality, and overall professional credibility.
Key Responsibilities
Site Reliability Engineering
Evaluate work products involving production reliability, service availability, and operational resilience
Assess whether recommendations reflect sound SRE principles and realistic production environments
Review approaches to reliability, scalability, performance, and service health
Identify technically weak assumptions, operational gaps, and impractical recommendations
Apply professional judgement grounded in real-world Site Reliability Engineering experience
Incident Management & Response
Review incident-response plans, escalation workflows, and operational procedures
Assess incident classification, prioritisation, ownership, and coordination
Evaluate whether proposed response actions are appropriate for severity and business impact
Identify gaps in communication, escalation, containment, or recovery
Review incident-management approaches for speed, clarity, and operational effectiveness
Root Cause & Post-Incident Analysis
Evaluate root-cause analyses and post-incident reviews
Assess whether conclusions are supported by technical and operational evidence
Identify shallow causal analysis, unsupported assumptions, or missed contributing factors
Review corrective and preventive actions for practicality and effectiveness
Evaluate whether lessons learned translate into meaningful reliability improvements
Reliability Metrics & Service Health
Review analyses involving availability, latency, reliability, and service-performance metrics
Assess use of SLIs, SLOs, error budgets, and related reliability measures where relevant
Evaluate whether metrics appropriately reflect service health and user impact
Identify inconsistencies between underlying data and reported conclusions
Review whether reliability targets and operational recommendations are realistic
Monitoring & Operational Readiness
Evaluate monitoring, alerting, observability, and operational-readiness approaches
Assess whether alerts are actionable and aligned with meaningful service conditions
Review escalation paths, runbooks, and response procedures
Identify gaps in detection, diagnosis, or operational preparedness
Evaluate whether proposed controls support reliable production operations
Resilience & Failure Management
Review scenarios involving outages, degraded performance, capacity constraints, and system failures
Assess mitigation, recovery, and resilience strategies
Evaluate trade-offs between reliability, performance, complexity, and operational cost
Identify single points of failure or poorly addressed dependencies
Review recommendations for reducing recurrence and improving service resilience
Documents & Presentation Review
Evaluate incident reports, reliability analyses, operational documents, spreadsheets, and slide decks for accuracy and completeness
Review stakeholder and executive presentations for clarity and decision usefulness
Identify factual, technical, analytical, aesthetic, and formatting issues
Assess whether charts, tables, and visuals accurately represent underlying operational information
Ensure conclusions and recommendations are clearly supported by evidence
Structured Evaluation & Feedback
Assess assigned outputs against domain-specific quality criteria
Identify technical, operational, analytical, and presentation weaknesses
Distinguish substantive reliability issues from minor editorial concerns
Provide clear, structured written feedback explaining identified strengths and weaknesses
Apply evaluation standards consistently across different SRE and incident-management work products
Ideal Profile
5+ years of relevant professional experience in Site Reliability Engineering, incident management, reliability engineering, DevOps, production engineering, systems engineering, or a closely related field
Strong practical understanding of SRE and production reliability
Hands-on experience with incident response, escalation, post-incident review, and root-cause analysis
Experience with monitoring, observability, service health, and operational readiness
Strong understanding of availability, performance, resilience, and service-level objectives
Ability to assess technical recommendations for operational feasibility and reliability impact
Highly proficient with Microsoft Office and Google Workspace
Advanced proficiency with PowerPoint / Google Slides
Strong spreadsheet and analytical skills
Native or professional fluency in English
Excellent written communication and ability to provide precise, structured feedback
Strong attention to technical, operational, analytical, and presentation detail
Master's degree or higher from a recognised institution is advantageous
Engagement Details
Part-time independent contractor engagement
Fully remote
Flexible scheduling based on project requirements
Compensation: $70–$110/hour
Work includes evaluation of incident-management plans, reliability analyses, operational documentation, spreadsheets, and presentation materials
Projects may be extended, shortened, or concluded based on project needs and performance
Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
H1-B and STEM OPT support is unavailable for this engagement
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.
This offer from "24-MAG" has been enriched by Jobgether.com and got a 74% flex score.