Kubernetes Infrastructure Reliability Engineer
roche.wd3.myworkdayjobs.com
At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come.
Join Roche, where every voice matters. The Position The Kubernetes Infrastructure Reliability Engineer is a highly skilled expert responsible for solving complex business problems using advanced cloud native technologies. The engineer will build and maintain a Kubernetes-based infrastructure, enabling the modernization of business applications and processes. This role combines software and systems engineering to optimize systems, increase efficiency, and eliminate operational work through automation.
You will be part of the global CaaS infrastructure team at a leading healthcare company, working with members across different regions. The team's mandate is to deliver, maintain, and continuously improve a highly available Kubernetes platform across hybrid cloud deployments, including on-premise data centers and public clouds like AWS. In this role, you will apply software engineering principles to operations to build and run massively distributed, fault-tolerant systems, focusing heavily on automation, security, and observability.
Job Responsibilities Service Reliability and Optimization: Focus on capacity planning and launch reviews for services before they go live. Perform blameless postmortems and proactive identification of potential outages to foster iterative improvements Accountability/Problem Solving: Resolves complex problems in a global Kubernetes-based infrastructure through in-depth evaluation of variable factors, including inter-organizational impact, balanced with effective consultative engagement of key stakeholders.
Leads end-to-end design of infrastructure solutions and maintains component standards. Evaluates promising solutions via Proof of Concept (PoCs) and feasibility studies across multiple areas, and serves as an internal escalation point for major incidents Stakeholder Management: Acts as a bridge between engineering and operations. Communicates and presents complex information and potential solutions to cross-functional teams and the business in non-technical terms.
Represents the organization as a prime contact on initiatives and interacts with senior internal and external personnel. Uses deep knowledge to influence IT infrastructure vendor product evaluations and collaborates with multiple IT partners (e.g. Enterprise Architects, Solution Owners) to integrate feedback. Mentors and shares DevOps culture, guiding developers on how to create and deploy cloud-native applications Impact/Strategy: Provides technical leadership and direction for small-to-medium sized initiatives (projects, lifecycle work, PoCs).
Ensures solutions comply with Quality/Regulatory standards and that designs adhere to the organization’s Technical Architecture Framework (TAF) policies and directions. Assists in planning technology projects, estimating engineering resources, dependencies, risks and timelines for successful delivery Business / Technical ability: Applies extensive cloud native technical expertise, acting as a recognized expert in Kubernetes and maintaining in-depth knowledge across related cloud native technologies (containers, AWS, etc.).
Demonstrates a detailed understanding of how IT infrastructure impacts respective Roche business processes and outcomes Qualifications Education & Professional Experience Without Degree: 4–7 years of relevant experience Bachelor’s Degree: 2–5 years of relevant experience Master’s Degree: 1–3 years of relevant experience At least 1 year of experience working in a multinational environment; healthcare industry experience is a plus Technical Skills Kubern