寻找你的下一个职业机会

按职位、技能和地点搜索招聘信息。准备申请前,先仔细了解职位要求。

找到一个吸引人的职位名称只是求职的起点。请将工作职责、招聘要求和工作条件与你的实际经历进行比较。本指南帮助你筛选机会、准备有针对性的申请材料,并确认每份申请应该在哪里提交。

此界面为简体中文。雇主发布的职位名称和描述保留原文,可能为英文。

清除筛选

搜索结果: 6,190

← 返回搜索结果

Senior/Staff Platform Engineer

Virtasant

地点
Anywhere
发布日期
2026年10月5日

申请前,请在雇主网站确认职位仍在招聘,并检查完整要求和条件。

职位描述

此界面为简体中文。雇主发布的职位名称和描述保留原文,可能为英文。

Senior/ Staff Platform Engineer Type: Remote Coverage: Pacific Hours (8:00 AM – 5:00 PM PST and On-call every 4-5 weeks) Job Description: We are looking for a Senior/Staff Platform Engineer to build, operate, and evolve large-scale production infrastructure. This is a hands-on platform and reliability engineering role for someone who has deep experience operating Kubernetes and cloud infrastructure, troubleshooting complex production systems, and building the automation and tooling that keeps those systems reliable. The work spans Kubernetes, Linux, cloud infrastructure, networking, observability, CI/CD, reliability, and production operations. You will write production code and automation in Go, Python, or Java, but this is not primarily a software development role. We are looking for an engineer who understands the systems underneath the applications and can independently diagnose and solve infrastructure problems across multiple layers. This is a highly autonomous role. You will work directly with technical stakeholders, own ambiguous infrastructure initiatives from design through production, and be trusted to drive technical decisions and critical issues without requiring constant direction. Key Responsibilities:Platform and Kubernetes Engineering: Design, build, operate, and improve production Kubernetes platforms. Own platform-level concerns including cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability. Troubleshoot Kubernetes beyond the application layer, including networking/CNI, scheduling, node behaviour, resource constraints, controllers, and cluster-level failures. Operate and improve large-scale, highly available infrastructure across cloud, hybrid, virtualised, and/or bare-metal environments. Diagnose complex issues spanning Kubernetes, containers, Linux, networking, and underlying infrastructure. Optimise platform infrastructure for reliability, performance, scalability, and operational efficiency. Software Development and Automation: Write, maintain, and improve production tooling and automation using Go, Python, or Java. Build software and automation that improves platform operations, reliability, deployment, troubleshooting, and developer experience. Read, debug, and contribute to existing production codebases. Develop internal services, APIs, integrations, and operational tooling where needed. Automate repetitive operational processes and reduce manual intervention across the platform. Apply sound software engineering practices, including testing, code review, maintainability, and documentation. Reliability and Production Operations: Own the reliability and operational health of critical production infrastructure. Lead or contribute significantly to incident response for complex platform and infrastructure issues. Investigate root causes and implement durable remediation rather than temporary fixes. Define and improve SLOs, SLIs, alerting, and operational processes. Troubleshoot systems using logs, metrics, traces, profiling tools, and system-level diagnostics. Drive improvements in availability, performance, capacity, resilience, and operational readiness. Contribute to disaster recovery planning, testing, and continuous improvement. Infrastructure as Code and Delivery: Build and maintain infrastructure as code using Terraform and related automation technologies. Create reusable infrastructure patterns and improve automation as the platform evolves. Build and improve CI/CD and deployment workflows supporting large-scale engineering environments. Balance delivery speed with reliability, security, scalability, and operational requirements. Work across infrastructure provisioning, configuration management, deployment automation, and production operations. Participate in planning and executing production cloud or infrastructure migrations, including dependency analysis, networking, cutover, rollback, and production validation. Observability: Build and maintain production monitoring, metrics, dashboards, alerting, logging, and distributed tracing. Improve observability so engineers can identify and diagnose failures quickly. Use production telemetry to identify reliability, capacity, and performance problems before they become major incidents. Continuously improve incident detection and reduce time to diagnosis and recovery. Collaboration and Technical leadership: Work directly with customer and internal engineering teams to understand requirements, investigate problems, and drive technical solutions. Communicate architecture, technical decisions, risks, trade-offs, and progress clearly to technical stakeholders. Own complex infrastructure initiatives from initial problem definition through design, implementation, and production operation. Contribute to architecture discussions, RFCs, design reviews, and technical direction. Mentor other engineers and help improve engineering and operational practices across the team. Operate independently in ambiguous situations and take ownership when immediate technical or management direction is unavailable. Qualifications:Education and Experience: 10+ years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related fields; 10+ years is preferred for Staff-level candidates. Significant hands-on experience operating complex production infrastructure and distributed systems. Demonstrated experience building and operating production Kubernetes platforms, not only deploying applications onto existing clusters. Production programming experience with Go, Python, or Java. Strong experience with production reliability, incident response, troubleshooting, and operational ownership. Experience independently owning complex technical initiatives from an ambiguous starting point through production. Experience working directly with technical stakeholders or customers and communicating complex technical topics effectively. Degree in Computer Science, Engineering, or a related field, or equivalent practical experience. Technical Skills: Deep understanding of production Kubernetes infrastructure, including cluster architecture, networking/CNI, NetworkPolicy, scheduling, resource management, nodes, security/RBAC, and cluster behaviour. Strong Linux fundamentals and hands-on production systems troubleshooting. Strong understanding of networking concepts including DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking. Production experience with at least one major cloud platform: AWS, GCP, or Alicloud. Infrastructure as code at scale using Terraform or equivalent tooling. Configuration management and automation experience with technologies such as Ansible, Puppet, or similar. Strong production debugging and root-cause analysis skills across infrastructure and distributed systems. Observability experience using metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent. Experience with CI/CD infrastructure and modern software delivery practices. Docker and container tooling as part of the production lifecycle. Production coding and automation experience in Go, Python, or Java. Understanding of high availability, capacity planning, disaster recovery, and production resilience. Preferred: Experience planning and executing production cloud or infrastructure migrations, including cutover and rollback strategies. Experience operating large-scale or multi-cluster Kubernetes environments. Experience with hybrid cloud, on-premises, virtualised, or bare-metal infrastructure. Experience building or modifying Kubernetes controllers or operators. Advanced Kubernetes networking, CNI, NetworkPolicy, service mesh, mTLS, or workload identity experience. Multi-cloud infrastructure experience. Experience designing and testing disaster recovery strategies. Experience with large-scale CI/CD or developer infrastructure. Experience with capacity planning and performance engineering. Experience building internal platform tooling or improving developer experience. Experience with security, infrastructure hardening, IAM, or compliance requirements. Previous technical leadership, mentoring, or Staff/Principal-level engineering responsibilities. Experience working directly with external customers or stakeholders in a consulting or service-delivery environment. Soft Skills: Exceptional written and verbal technical communication. Strong analytical, debugging, and problem-solving ability. High degree of ownership and ability to operate independently. Comfortable making technical decisions and driving work forward in ambiguous environments. Able to communicate effectively with customers, engineers, and technical leadership. Strong technical judgment and ability to explain trade-offs rather than simply implement predefined solutions. Able to lead technically and influence others without requiring formal people-management authority. Comfortable working within a distributed, highly technical team. Please Note: This is a hands-on platform and reliability engineering role. We are not looking for candidates whose experience has been limited to deploying applications onto Kubernetes, consuming managed cloud services, or provisioning infrastructure without owning its production operation. The ideal candidate has personally built, operated, troubleshot, and improved production infrastructure and can clearly explain what they owned, how the underlying systems worked, and how they approached failures and architectural trade-offs. This role requires significant autonomy and strong technical communication. Engineers should be comfortable working directly with stakeholders and driving complex technical issues without continuous oversight. Production coding experience is required, but it may be in Go, Python, or Java. Go is not a requirement. Cloud migration experience is strongly preferred but is not a knockout requirement. This role is currently open to candidates based in Brazil, Mexico, or Canada.

查找职位、比较要求,再准备申请

从你希望从事的职位或运用的技能开始搜索。调整地点和职业筛选,打开职位比较工作职责。如果没有结果,可以使用更短的关键词,或逐一移除筛选条件。

区分必备要求和优先条件,检查已列出的工作安排、薪资和地点。远程职位也可能要求特定居住国家、工作许可或工作时间重合。请向雇主确认职位信息和完整条件。

选择能够回应职位要求的真实经历,说明你的贡献,只使用能够证实的数字。遵循雇主的申请说明,并在提交前检查联系方式、文档内容和 PDF。

提交申请前的检查清单

求职常见问题

为什么有些职位使用英文?

职位名称和描述由招聘企业撰写。为避免改变招聘要求或工作条件,我们保留原文。操作界面和本指南使用简体中文。如果中文搜索没有结果,可以尝试使用职位发布语言中的名称或技能,例如“software engineer”。搜索词不会自动翻译,界面语言也不代表企业要求的申请语言。

远程职位是否允许从任何国家工作?

不一定。企业可能对居住国家、工作许可或工作时段有要求。请查看原始招聘页面中的具体条件。如果未说明,应先向企业确认,再判断能否从你所在的地区工作。“远程”标签本身并不代表没有地点限制。

搜索没有结果时应该怎么办?

尝试更通用的职位名称或单个技能,并逐一移除筛选条件。不同企业可能用不同名称描述相似工作。如果某个职位已经消失,请在企业招聘页面搜索其职位编号。扩大搜索范围不会让已经关闭的职位重新开放。

申请会通过 ResumizeAI 直接提交吗?

申请按钮会打开外部网站。请按照企业或招聘服务的说明,在该网站完成并确认提交。在 ResumizeAI 中准备简历并不等于已经申请职位。如果链接只打开企业网站,请先找到对应职位,再继续申请流程。

如何针对职位调整简历和求职信?

将招聘要求与能够解释清楚的项目、任务和成果联系起来。突出相关经历,不要添加未经实际掌握的技能或虚构成绩。在求职信中用具体例子说明申请动机,并遵守企业要求的语言和文件格式。提交前检查两份文件,确保内容准确、联系方式正确、链接可用。