Back to jobs
Senior Site Reliability Engineer (remote within EMEA)
FYULAnywherePosted 3w ago
Site Reliability Engineer
Full-time
Remote
Apply on company siteCraft my tailored resume free
Free to start · No card · 5 credits the moment you sign up
Job description
About the position
Platform Infrastructure builds, operates, and continuously evolves FYUL's container platform and cloud foundation. We foster a DevOps culture through self-service tooling, enabling product engineering teams to ship reliable, secure, and cost-efficient services as the business scales. The team owns our AWS cloud accounts, Kubernetes platform, cloud networking, observability stack, core databases, CI/CD pipelines, and infrastructure-as-code, and acts as the go-to partner for engineering teams on cloud and DevOps topics.\n\nWe're hiring a Senior SRE II to join Platform Infrastructure as one of the team's senior individual contributors. At this level, you're the go-to person for our most complex infrastructure problems: you architect and drive large-scale automation and reliability initiatives, set standards other engineers follow, and mentor Associate and mid-level SREs. You'll split your time between hands-on platform work - Kubernetes, AWS, GCP, CI/CD, observability - and technical leadership: proposing designs, reviewing others' work, and helping the team make good build-vs-buy and cost/reliability trade-offs.
Responsibilities
Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts and environments using infrastructure as code.Design and operate Amazon EKS clusters, including networking policies, persistent storage, and scaling strategies for containerized workloads.Own and evolve core platform services: cloud networking, Kubernetes, and the databases and messaging systems engineering teams depend on.Drive large-scale automation projects and set standards for using Terraform / Terragrunt and GitOps (ArgoCD) across teams.Lead adoption of automation to reduce manual operational work and keep environments consistent and repeatable.Be the go-to person for solving complex, cross-service infrastructure problems.Drive initiatives that improve reliability and observability (Grafana, Prometheus, Loki, Tempo, Mimir) so systems scale with minimal manual intervention.Participate in on-call rotation, lead incident response for production issues, and write clear runbooks, ADRs, and postmortems.Lead security efforts within the team - IAM, encryption, secure logging - and mentor others on secure infrastructure practices.Audit infrastructure spend regularly and drive cost optimization across the platform (rightsizing, autoscaling, FinOps practices).Mentor mid-level SREs, provide detailed feedback, and support onboarding of new team members.Communicate complex technical concepts clearly to both engineers and non-technical stakeholders.Partner with product engineering squads to understand their needs and represent Platform Infrastructure in cross-team initiatives.Requirements
Solid Linux systems administration background and comfort scripting in Python.Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS, and familiarity with the Well-Architected Framework; experience in multi-account AWS environments is a strong plus.Hands-on experience operating and troubleshooting Kubernetes (EKS) at production scale, including Helm chart development, CNI networking (we run Cilium), pod networking/IPAM concepts, and container security (ECR, image scanning).Proficiency with Terraform (modules, state management) and ideally Terragrunt for multi-environment management; GitOps experience with ArgoCD.Experience with Postgres, MySQL and/or MongoDB in production scale, including Aurora.CI/CD experience with Jenkins (Jenkinsfile, shared libraries) and/or GitHub Actions, and familiarity with deployment strategies such as blue-green and canary.Experience with the Grafana observability stack (Grafana, Prometheus, Loki, Tempo, Mimir) - metrics design, dashboarding, alerting, log aggregation, and distributed tracing. Not only using but also maintaining it.Practical incident management experience: on-call rotations, structured incident response, and writing runbooks/postmortems.Working knowledge of 12-Factor App principles and cost optimization / FinOps awareness.A methodical, data-driven approach to troubleshooting rather than guessing.Strong written communication - you write runbooks, ADRs, and postmortems that others can actually follow.Comfortable driving initiatives with ambiguous ownership, and taking accountability for outcomes rather than waiting to be asked.Track record of mentoring less senior engineers and giving direct, constructive feedback.Several years of hands-on production infrastructure/SRE experience, with demonstrated ownership of initiatives at a senior individual-contributor level (leading design work, setting standards, being the escalation point for hard problems).Nice-to-haves
GCP Experience.Experience with Kafka / AWS MSK.Prior experience in regulated or compliance-sensitive environments (security best practices, access reviews).Experience contributing to a platform/DevEx roadmap that other engineering teams consume as a self-service product.Benefits
A global, inclusive team that’s as supportive as it is ambitious and serious about getting things doneAn opportunity to work remotely or in a modern and welcoming office in RigaFlexible working hours (start your day as late as 11 AM)Private health insurance2 extra paid days off to focus on your mental or physical well-being1 extra paid day off to celebrate a Birthday or any other celebration of your choiceInternal and external learning opportunitiesAccess to mentorship, internal meetups, and hackathons, both on-site and onlineFree and healthy lunch if you work from the Rīga officeDesign and order your own merch using our platforms with an employee discountExciting team-building events and parties you’ll never forget!