次の仕事の機会を見つけましょう

職種、スキル、勤務地で求人を検索できます。応募書類を準備する前に募集要件を確認しましょう。

魅力的な職種名を見つけることは、仕事探しの出発点です。業務内容、応募条件、働き方を自分の経験と照らし合わせましょう。このガイドでは、候補の絞り込み、応募書類の準備、提出先の確認を順番に進めるためのポイントを紹介します。

操作画面は日本語です。企業が掲載した求人の職種名や説明文は原文のまま表示され、英語の場合があります。

条件をクリア

検索結果: 6,190

← 検索結果に戻る

Senior/Staff Platform Engineer

Virtasant

勤務地
Anywhere
掲載日
2026年10月5日

応募する前に、企業のウェブサイトで現在も募集中かどうかと詳しい条件を確認してください。

仕事内容

操作画面は日本語です。企業が掲載した求人の職種名や説明文は原文のまま表示され、英語の場合があります。

Senior/ Staff Platform Engineer Type: Remote Coverage: Pacific Hours (8:00 AM – 5:00 PM PST and On-call every 4-5 weeks) Job Description: We are looking for a Senior/Staff Platform Engineer to build, operate, and evolve large-scale production infrastructure. This is a hands-on platform and reliability engineering role for someone who has deep experience operating Kubernetes and cloud infrastructure, troubleshooting complex production systems, and building the automation and tooling that keeps those systems reliable. The work spans Kubernetes, Linux, cloud infrastructure, networking, observability, CI/CD, reliability, and production operations. You will write production code and automation in Go, Python, or Java, but this is not primarily a software development role. We are looking for an engineer who understands the systems underneath the applications and can independently diagnose and solve infrastructure problems across multiple layers. This is a highly autonomous role. You will work directly with technical stakeholders, own ambiguous infrastructure initiatives from design through production, and be trusted to drive technical decisions and critical issues without requiring constant direction. Key Responsibilities:Platform and Kubernetes Engineering: Design, build, operate, and improve production Kubernetes platforms. Own platform-level concerns including cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability. Troubleshoot Kubernetes beyond the application layer, including networking/CNI, scheduling, node behaviour, resource constraints, controllers, and cluster-level failures. Operate and improve large-scale, highly available infrastructure across cloud, hybrid, virtualised, and/or bare-metal environments. Diagnose complex issues spanning Kubernetes, containers, Linux, networking, and underlying infrastructure. Optimise platform infrastructure for reliability, performance, scalability, and operational efficiency. Software Development and Automation: Write, maintain, and improve production tooling and automation using Go, Python, or Java. Build software and automation that improves platform operations, reliability, deployment, troubleshooting, and developer experience. Read, debug, and contribute to existing production codebases. Develop internal services, APIs, integrations, and operational tooling where needed. Automate repetitive operational processes and reduce manual intervention across the platform. Apply sound software engineering practices, including testing, code review, maintainability, and documentation. Reliability and Production Operations: Own the reliability and operational health of critical production infrastructure. Lead or contribute significantly to incident response for complex platform and infrastructure issues. Investigate root causes and implement durable remediation rather than temporary fixes. Define and improve SLOs, SLIs, alerting, and operational processes. Troubleshoot systems using logs, metrics, traces, profiling tools, and system-level diagnostics. Drive improvements in availability, performance, capacity, resilience, and operational readiness. Contribute to disaster recovery planning, testing, and continuous improvement. Infrastructure as Code and Delivery: Build and maintain infrastructure as code using Terraform and related automation technologies. Create reusable infrastructure patterns and improve automation as the platform evolves. Build and improve CI/CD and deployment workflows supporting large-scale engineering environments. Balance delivery speed with reliability, security, scalability, and operational requirements. Work across infrastructure provisioning, configuration management, deployment automation, and production operations. Participate in planning and executing production cloud or infrastructure migrations, including dependency analysis, networking, cutover, rollback, and production validation. Observability: Build and maintain production monitoring, metrics, dashboards, alerting, logging, and distributed tracing. Improve observability so engineers can identify and diagnose failures quickly. Use production telemetry to identify reliability, capacity, and performance problems before they become major incidents. Continuously improve incident detection and reduce time to diagnosis and recovery. Collaboration and Technical leadership: Work directly with customer and internal engineering teams to understand requirements, investigate problems, and drive technical solutions. Communicate architecture, technical decisions, risks, trade-offs, and progress clearly to technical stakeholders. Own complex infrastructure initiatives from initial problem definition through design, implementation, and production operation. Contribute to architecture discussions, RFCs, design reviews, and technical direction. Mentor other engineers and help improve engineering and operational practices across the team. Operate independently in ambiguous situations and take ownership when immediate technical or management direction is unavailable. Qualifications:Education and Experience: 10+ years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related fields; 10+ years is preferred for Staff-level candidates. Significant hands-on experience operating complex production infrastructure and distributed systems. Demonstrated experience building and operating production Kubernetes platforms, not only deploying applications onto existing clusters. Production programming experience with Go, Python, or Java. Strong experience with production reliability, incident response, troubleshooting, and operational ownership. Experience independently owning complex technical initiatives from an ambiguous starting point through production. Experience working directly with technical stakeholders or customers and communicating complex technical topics effectively. Degree in Computer Science, Engineering, or a related field, or equivalent practical experience. Technical Skills: Deep understanding of production Kubernetes infrastructure, including cluster architecture, networking/CNI, NetworkPolicy, scheduling, resource management, nodes, security/RBAC, and cluster behaviour. Strong Linux fundamentals and hands-on production systems troubleshooting. Strong understanding of networking concepts including DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking. Production experience with at least one major cloud platform: AWS, GCP, or Alicloud. Infrastructure as code at scale using Terraform or equivalent tooling. Configuration management and automation experience with technologies such as Ansible, Puppet, or similar. Strong production debugging and root-cause analysis skills across infrastructure and distributed systems. Observability experience using metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent. Experience with CI/CD infrastructure and modern software delivery practices. Docker and container tooling as part of the production lifecycle. Production coding and automation experience in Go, Python, or Java. Understanding of high availability, capacity planning, disaster recovery, and production resilience. Preferred: Experience planning and executing production cloud or infrastructure migrations, including cutover and rollback strategies. Experience operating large-scale or multi-cluster Kubernetes environments. Experience with hybrid cloud, on-premises, virtualised, or bare-metal infrastructure. Experience building or modifying Kubernetes controllers or operators. Advanced Kubernetes networking, CNI, NetworkPolicy, service mesh, mTLS, or workload identity experience. Multi-cloud infrastructure experience. Experience designing and testing disaster recovery strategies. Experience with large-scale CI/CD or developer infrastructure. Experience with capacity planning and performance engineering. Experience building internal platform tooling or improving developer experience. Experience with security, infrastructure hardening, IAM, or compliance requirements. Previous technical leadership, mentoring, or Staff/Principal-level engineering responsibilities. Experience working directly with external customers or stakeholders in a consulting or service-delivery environment. Soft Skills: Exceptional written and verbal technical communication. Strong analytical, debugging, and problem-solving ability. High degree of ownership and ability to operate independently. Comfortable making technical decisions and driving work forward in ambiguous environments. Able to communicate effectively with customers, engineers, and technical leadership. Strong technical judgment and ability to explain trade-offs rather than simply implement predefined solutions. Able to lead technically and influence others without requiring formal people-management authority. Comfortable working within a distributed, highly technical team. Please Note: This is a hands-on platform and reliability engineering role. We are not looking for candidates whose experience has been limited to deploying applications onto Kubernetes, consuming managed cloud services, or provisioning infrastructure without owning its production operation. The ideal candidate has personally built, operated, troubleshot, and improved production infrastructure and can clearly explain what they owned, how the underlying systems worked, and how they approached failures and architectural trade-offs. This role requires significant autonomy and strong technical communication. Engineers should be comfortable working directly with stakeholders and driving complex technical issues without continuous oversight. Production coding experience is required, but it may be in Go, Python, or Java. Go is not a requirement. Cloud migration experience is strongly preferred but is not a knockout requirement. This role is currently open to candidates based in Brazil, Mexico, or Canada.

検索して比較し、応募を準備しましょう

次の仕事で活かしたい職種やスキルから検索を始めましょう。勤務地と職種の条件を調整し、求人を開いて担当業務を比較してください。結果がない場合は、短い検索語を試すか、条件を一つずつ外してみましょう。

必須条件と歓迎条件を分けて確認し、勤務時間、報酬、勤務地が掲載されている場合は目を通しましょう。リモート勤務でも、居住国、就労資格、勤務時間帯が指定されることがあります。現在の募集状況と詳しい条件は企業に確認してください。

募集要件に関連する実際の経験を選び、自分の役割と貢献を説明しましょう。数値は根拠がある場合にのみ使ってください。企業の応募方法に従い、送信前に連絡先、内容、PDFの表示を確認しましょう。

応募前の確認リスト

仕事探しに関するよくある質問

英語の求人が表示されるのはなぜですか?

職種名や説明文は企業が作成したものです。条件や要件を変えないように原文を掲載しています。操作画面とこのガイドは日本語です。日本語の検索で見つからない場合は、「software engineer」のように、掲載言語の職種名やスキル名でも検索してみてください。検索語が自動翻訳されるわけではありません。

リモート求人なら、どの国からでも働けますか?

必ずしもそうではありません。居住国、就労資格、勤務時間帯が指定されている場合があります。元の求人ページで条件を確認してください。記載がない場合は、自分の居住地から応募できると判断する前に企業へ確認しましょう。リモートという表示だけでは、勤務地の制限がないことは分かりません。

検索結果が出ない場合はどうすればよいですか?

より一般的な職種名や一つのスキルで検索し、条件を一つずつ外してみてください。同じような仕事でも企業によって呼び方が異なります。特定の求人が見つからなくなった場合は、企業の採用ページで求人番号を探しましょう。検索条件を広げても、終了した募集が再開するわけではありません。

ResumizeAIから応募が送信されますか?

応募ボタンは外部サイトを開きます。企業または採用サービスの案内に従い、そのサイトで送信を完了してください。ResumizeAIで応募書類を作成しただけでは応募は送信されません。リンク先が企業のトップページの場合は、該当する募集を探してから手続きを進めましょう。

応募書類とカバーレターはどう調整すればよいですか?

募集要件と、自分が説明できるプロジェクト、業務、成果を結び付けます。経験していない技術や実績を加えず、関連する経験を分かりやすく示してください。カバーレターでは応募理由を具体例とともに伝えます。指定された言語とファイル形式に従い、送信前に両方の書類を確認しましょう。