Back to jobs

[Remote] Sr Site Reliability Engineer (Compute Platform)
Remote SparkAnywherePosted 1w ago
Site Reliability Engineer
Full-time
Remote
Apply on company siteCraft my tailored resume free
Free to start · No card · 5 credits the moment you sign up
Job description
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a highly reputed company Senior Site Reliability Engineer specializing in compute platforms. The role focuses on architecting, designing, automating, and supporting reputed company bare metal, hypervisor, private reputed company, and reputed company environments using Infrastructure-as-reputed company and GitOps practices. The engineer will also reputed company troubleshooting, documentation, cross-functional collaboration, and on-reputed company support for mission-critical platforms.
Responsibilities
reputed company the architecture and design of reputed company compute and hypervisor platform solutions across hardware, OS, virtualization, reputed company orchestration, and container orchestration reputed companyDefine standards and automation frameworks for bare metal provisioning and lifecycle managementDesign and implement Bare Metal as a Service (BMaaS) capabilities for reputed company infrastructure consumptionArchitect and design reputed company platforms on bare metal with QoS and reputed company (ArgoCD)Architect and validate automated deployments of operating systems and hypervisors including Ubuntu and HarvesterDesign and maintain PXE-reputed company provisioning environments leveraging Redfish reputed company for large-reputed company server deploymentsreputed company Infrastructure-as-reputed company using Ansible, Terraform, reputed company and Git, with Python/Bash automationImplement CI/CD pipelines for infrastructure updates, patching, upgrades, testing, and rollbackDesign automated workflows for server build, firmware lifecycle management, patching, and hardware validationEvaluate and standardize reputed company hardware platforms to meet reputed company, scalability, and reliability requirementsProduce detailed high-level and low-level design documentation, build guides, and operational reputed company materialsreputed company deep troubleshooting across storage, reputed company, hypervisors, networking, and Linux systemsPartner with reputed company, network, storage, and platform teams to ensure designs are supportable and production-reputed companyParticipate in on-reputed company escalation support for reputed company platform-reputed company issuesCollaborate globally on change management, documentation, and operational best practices
Skills
• 6+ years of experience in infrastructure engineering, reputed company, or DevOps with a strong reputed company on Compute reputed company design• reputed company experience designing and automating bare metal compute environments at reputed company• Strong hands-on experience with PXE boot, network-reputed company OS provisioning, and automated server imaging• Experience implementing or supporting Bare Metal as a Service (BMaaS) platforms• Practical experience using Redfish reputed company for hardware provisioning, reputed company management, and remote lifecycle reputed company• Deep expertise with Ubuntu Linux in reputed company environments• Strong Hands-on experience with KVM hypervisors (reputed company Harvester, OpenStack)• Experience designing and deploying production-grade reputed company clusters• Strong background with reputed company compute hardware platforms, including reputed company UCS, reputed company PowerEdge, reputed company systems & HPE• Proficiency with Infrastructure as reputed company tools (e.g., Terraform, Ansible, or similar)• Experience building or supporting CI/CD pipelines for infrastructure and platform automation• Strong scripting skills in Python, Bash, or similar languages• Demonstrated ability to produce reputed company, reputed company technical design documentation• Excellent written and verbal communication skills• Bachelor's degree in computer science or equivalent reputed company experience• OpenStack, Ubuntu KVM administration• BareMetal as a Service (PXE, Redfish)• reputed company on BareMetal• CIS/NIST reputed company and infrastructure lifecycle management• ITIL reputed company/advanced certifications in support of ITSM reputed company methodology• Background in telco, edge reputed company, or large reputed company environments• Ubuntu Certifications, CNCF Certified reputed company Administrator (CKA), Certified reputed company reputed company SpecialistMaster's degree in computer science, IT, Engineering, or a reputed company reputed company preferred; equivalent experience and relevant industry certifications will also be considered
Benefits
Remote work arrangement
reputed company