Level 3 Product Technical Support Incident Manager
State Street
Role Summary The Level 3 Product Technical Support Incident Manager owns the end-to-end lifecycle of major (Sev1/Sev2) incidents across the SaaS APAC estate — driving speed, structure, and clarity from detection through resolution and post-incident review. The role is the single point of coordination that prevents multiple teams from chasing different root causes simultaneously, protects the client experience through disciplined communication, and converts every significant incident into durable problem-management and preventive action.
This is a process-leadership role: the Incident Manager directs the response, while L3/L4 own the technical diagnosis and fix. Primary Objective Restore service in the shortest possible time through structured command of the incident, maintain authoritative stakeholder communication, and ensure every major incident is followed by root-cause analysis (RCA), ensuring closure and documenting preventive actions.
Key Responsibilities Incident command — Own and drive the major-incident process; coordinate Support, L3/ASE, SaaS Ops, DevOps, and Infrastructure to a single resolution path. Severity & impact assessment — Rapidly establish business, client, portfolio, market, trading, and regulatory impact to confirm severity rating and determine response cadence. Communications ownership — Be the single exec-level voice for status; issue timely, accurate updates to clients, executives, and internal stakeholders throughout the incident.
Coordination & de-confliction — Determine appropriate subject matter experts immediately; prevent duplicated or conflicting investigation streams; ensure clear ownership of each workstream and decision. Timeline & audit trail — Maintain an accurate incident timeline, decisions log, and evidence trail to support review and regulatory scrutiny. Escalation management — Decide when to escalate and engage other teams such as SaaS Ops, DevOps, Infrastructure, or Engineering and manage the escalation path with urgency and judgement.
Post-incident review — Chair blameless post-mortems, capture corrective and preventive actions, and track them to closure with clear owners. Problem management — Partner with L3 to convert recurring incidents into problem records and drive systemic remediation. Continuous improvement — Monitor incident metrics (MTTR, recurrence, SLA adherence) and drive readiness through runbooks, on-call discipline, and table-top exercises.
Domain specific skills – OMS; EMS; IBOR; Portfolio Management; Market Data; Batch processing; Trade processing; Securities lifecycle Skills & Experience 7+ years in IT service management or production operations, with demonstrable major-incident command experience. Strong grasp of ITIL incident, problem, and change-management disciplines (certification preferred). Critical thinking - Analysing logic and facts to validate technical recovery plans without needing deep, hands-on coding skills.
Proven ability to lead high-pressure Sev1 bridges and remain calm, structured, and decisive. Exceptional written and verbal communication — able to brief traders, executives, and clients with clarity and confidence. Proven ability in clear process documentation and governance thereof. Financial services / investment management / trading systems context strongly preferred; CRIMS familiarity a plus.
Hands-on with ITSM tooling (e.g. ServiceNow, Salesforce, Jira) and comfortable interpreting logs, dashboards, and monitoring signals. Technical framework knowledge and application (i.e. ITIL for structured incident lifecycles) Business-impact thinking over pure technical troubleshooting; strong stakeholder-management and influencing skills. Conduct and lead thorough post incident review (including detailed RCA analysis) Demonstrated track record of continuous improvement B.S.
in a technical or business discipline, or equivalent operational experience. Additional requirements Periodic off-hours support About State Street What we do