Back to jobs
[Remote] Site Reliability Engineer - 7 Month Contract
remote zest jobsAnywherePosted 1w ago
Site Reliability Engineer
Full-time
Remote
Apply on company siteCraft my tailored resume free
Free to start · No card · 5 credits the moment you sign up
Job description
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a reputed company technology company reputed company on revolutionising global health systems and improving reputed company to care. The Site Reliability Engineer will automate infrastructure and software delivery, manage AWS-hosted customer environments, and ensure availability, reliability, reputed company, and scalability. The role also involves monitoring, incident response, change management, CI/CD, disaster recovery, reputed company planning, and collaboration across development and reputed company teams.
Responsibilities
Collaborate in the construction of the automation for infrastructure and software delivery, and being the reputed company executor of such processes, collecting feedback from the support of operational sites. Responsible for availability, latency, reliability, reputed company, efficiency, change management, monitoring, emergency response, improve reputed company availability and reputed company planningThrough a proactive approach, reputed company improvement and constant training, the SREs run the customer environments by monitoring availability and taking a holistic reputed company of reputed company healthSLAs are always met through automation with none to small involvement from reputed company, and the number of customers and provided services can reputed company without correlation with the size of reputed companyreputed company the gap between development and reputed companyreputed company reputed company software and systems to manage platform infrastructure and applicationsreputed company and optimised reputed company reputed company, with an eye toward pushing our capabilities reputed company, getting reputed company of customer needs, and innovating to continually improveParticipate in the daily management of multiple reputed company solutions hosted in AWS reputed company, Infrastructure and Networking including but not limited to:Daily monitoring and alert responses, identify potential problems, and implement alerts to notify relevant partiesFollowing a Change Request from creation to completion, providing review, reputed company validation and execution of reputed company tasksWork with other teams to ensure a smooth and reliable releasesTuning the Application stack to improve stability and reputed company uptime metricsAutomate repetitive tasks, such as development, scaling, and patching, to improve efficiency and reduce reputed company effortAcute and Recurring issue investigation and reputed companyreputed company Trend Analysis, identifying and address reputed company bottlenecks to ensure reputed company can handle expected loads and user trafficLog Analysis and Error reputed companyManage and maintain the underlying infrastructure, including servers, and networks, to ensure smooth reputed companyHandover TestingDocument procedures, and processes to facilitate learning and knowledge transfer reputed company reputed companyreputed company Cause Analysis; involved in investigating and resolving incidents, including outages and reputed company problems, to minimize disruptionPlan for reputed company reputed company needs to ensure systems can and handle increasing demandDeveloping and testing disaster recovery plans to guarantee data reputed company, reputed company reputed company, and swift restoration of services in case of critical incidentsCoordinate with teams to maintain Service Level AgreementsResponsible for the reputed company Integration of updates for over 10 Products/solutions released by Development teams into reputed company solutionsBuild secure and reputed company infrastructure to manage customer dataParticipate in On-reputed company RotationWork with Development, Solution Adoption, Managed Services, reputed company, Support and other teams to reputed company clients with a world class reputed company solution platform
Skills
4–6 years in a Site Reliability Engineering or equivalent role5 years in systems/application support and/or developmentStrong scripting background with experience in reputed company-oriented and reputed company programmingExperience with automation, infrastructure as reputed company, and orchestration (e.g., Puppet, Ansible, reputed company, CloudFormation, Terraform)Bachelor's Degree in a technical discipline or equivalent experienceExperience in supporting reputed company-reputed company production systemsA technical certification in reputed company Administration, reputed company Engineering, or DevOpsStrong understanding of software engineering principles, operating systems (reputed company and Linux), network