Our Mission At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes. Job Summary The Cortex team builds and operates the industry’s leading SecOps platform, including XDR, XSIAM, XSOAR, and XPANSE. As a Senior Staff DevOps Engineer within the Cortex Production Engineering organization, you will play a key technical role in ensuring the reliability, scalability, availability, and operational efficiency of our large-scale cloud platform.
You will work hands-on with complex production systems while leading critical infrastructure, reliability, observability, and automation initiatives. You will partner closely with engineering teams to solve complex production challenges, improve platform architecture and operational readiness, and build automation that reduces operational toil. This role requires strong technical depth, ownership, and the ability to lead complex initiatives across multiple engineering teams while mentoring other engineers and raising the technical bar across the organization.
Qualifications Design, build, and operate highly available and scalable cloud infrastructure supporting critical Cortex services. Lead complex DevOps, infrastructure, reliability, and automation initiatives from design through production implementation. Troubleshoot complex production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed systems. Improve platform reliability through better monitoring, alerting, observability, capacity management, and operational practices.
Build automation and self-healing capabilities that reduce manual operational work and improve incident response. Partner with application and platform engineering teams to improve production readiness, scalability, and service architecture. Participate in and provide technical leadership during high-severity production incidents, including root cause analysis and corrective actions. Design and improve Infrastructure as Code, CI/CD, and GitOps-based deployment and infrastructure management.
Identify recurring operational problems and drive engineering solutions that permanently eliminate or reduce them. Mentor engineers, provide technical guidance, conduct design reviews, and help establish engineering best practices. Influence technical decisions across teams and drive adoption of scalable and reliable engineering solutions. Evaluate emerging technologies, including AI-assisted operations and automation, that can improve engineering productivity and operational efficiency.
Required Qualifications 8+ years of experience in DevOps, Site Reliability Engineering, Production Engineering, Cloud Infrastructure, or related areas. Strong hands-on experience operating large-scale production environments. Deep knowledge of Kubernetes, containers, and cloud-native architectures. Strong experience with public cloud platforms, prefe