Job Description: This is a senior technical lead role that pairs deep platform engineering expertise with the thought leadership needed to bring AI into production operations. The person owns the build, hardening, patching, and lifecycle of the operating system and platform estate, and serves as the top escalation point for complex incidents — while simultaneously identifying where AI and automation can replace manual toil, reduce incident volume, and shorten recovery times.
Success depends as much on change leadership as on engineering: shaping standards, winning stakeholder buy-in, and guiding 1st line operations through new AI-enabled ways of working without disrupting service levels. Key Responsibilities: Platform engineering depth — packaging and hardening approved OS builds, management agent stacks, patch currency, and configuration standards across the estate AI-driven operations innovation — identifying, piloting, and scaling AI/automation use cases for incident triage, root cause analysis, runbook execution, and predictive monitoring Thought leadership and evangelism — articulating an AI-for-operations vision, authoring white papers and R&D positions, staying ahead of emerging tooling Change management and adoption — driving organizational uptake of new practices, coaching 1st line teams, managing risk and resistance while protecting SLAs Incident, problem, and change discipline — major incident leadership, RCA, and rigorous change control in a regulated production environment Service improvement measurement — instrumenting and evidencing reductions in incident volume and MTTR from process redesign and automation Solution design and implementation planning — translating requirements into physical builds, costed designs, and multi-workstream delivery plans Automation and infrastructure-as-code fluency — scripting, configuration management, and pipeline-driven platform delivery Standards, compliance, and documentation — run books, technical policies, and AI governance/guardrails to document-control standards Cross-functional influence — partnering with architecture, security, development, vendors, and the business to land change .
Required Qualifications: Provides a high level of technical and subject matter expertise in one or more technologies and serves as a point of escalation for technical issues related to specialty. Produces, delivers, and maintains appropriate documentation for systems in accordance document control standards and procedures. Provide input to records, quality systems and management reports as required Contributes to the definition and implementation of improved operability on new and current systems.
Uses innovative methods including the redesign of process and providing technical solutions to reduce the volume and mean time to recover incidents in assigned business unit. Identifies risks & issues and takes ownership to deliver appropriate resolutions. Provides technical expertise for root cause analysis and problem management. Provides detailed implementation/project plans across multiple, complex work streams according to agreed standards and ensures project processes and timelines are understood and followed.
Works and cooperates with internal and external groups when required in order to fully support environments and maintain service. Adheres to change management procedures in defining, planning and implementing change in such a way that ensures appropriate coordination with other teams, minimizes service disruption, and ensures adherence to Service Level Agreements. Improves change management processes and procedures to ensure the most efficient processing of change within appropriate service risk constraints.
Provides specialist support during complex and/or major incidents. May be asked to lead recovery efforts during major incidents within business unit. Deputize for the team manager as required. Contributes to or author technical documentation such complex changes instructions.