Mediabistro logo
job logo

Managing Director - Head of Service Reliability, Operations and Platform

TIAA · Multiple locations ·

Job type:
Full Time

Production Support Engineer

Production Support Engineers are responsible for availability of systems and business capabilities including applications, infrastructure operations and data, business process execution which may include batch scheduling and processing, production of target business outputs, transmission of data including messages and files business needs requiring IT based solutions.
Key Responsibilities and Duties
Identify gaps in processing and production of outputs and availability of systems and outputs/outcomes. Monitor and alert the applications and services.
Manage environments and technology processing response to incidents and crisis management.
Responsible for triaging, recovery and remediation of incidents and problems including root cause analysis, defect and data analysis and impact assessment (i.e., end users, business processes, data, applications, and devices.
Roles at this level should be managing People Leaders with span of control of at least 7 FTE direct reports. Typically manages one or more groups that has a moderate to high impact on business results. Implements policies and procedures for the areas they manage; may participate with other areas in establishing broader policies and processes. Large contribution to strategic planning within an area of expertise. Provides advice to internal and/or external clients on implications of business trends, issues, operating environment changes and firm or business unit strategy. Manages the performance of direct reports through regular, timely feedback as well as the formal performance review process to ensure the delivery of projects and engagement, motivation and development of the team.
Demonstrates ability to apply understanding of concepts in complex situations at mastery level, serving as a key resource in providing advice, leadership, coaching and mentoring to others in the following: Think critically, analyze complex data, and develop long-term plans that align with the organization's vision and mission. Make sound decisions based on available information and data, and to effectively communicate those decisions to stakeholders. Inspire, motivate, and guide others towards achieving common goals and objectives. Communicate effectively with all levels of the organization, including the board of directors, senior management, and front-line employees. Effectively manage budgets, financial reporting, and overall financial performance of the organization. Lead and manage organizational change initiatives, including strategic planning, restructuring, and mergers and acquisitions. Attract, retain, and develop top talent within the organization, and to create a positive and productive work environment. Drive innovation and creativity within the organization, and to identify and pursue new growth opportunities. Recognize and manage one's own emotions, as well as to understand and respond effectively to the emotions of others. Understand and appreciate diverse cultural perspectives and to create a culture of inclusion and equity within the organization.
If you are part of an Agile team, you may be asked to perform the functions of Analyst (Tech), Dev, Lead Dev, or Engineer.
Educational Requirements
University (Degree) Preferred
Work Experience
10+ Years Required
Career Level 11PL
POSITION SUMMARY:
This role establishes and operates a Service Reliability Platform Organization, integrating command center operations, IT service management (ITSM), production operations, and ServiceNow platform capabilities into a unified control tower operating model.
The Managing Director ensures operational excellence through strong governance, automation, risk management, and continuous improvement, while delivering a resilient, scalable, and client-centric production environment.
Position Scope & Reporting:
Reports to: Head of Workplace Experience & Production Support
Leadership Level: Managing Director (Executive Band)
Span of Control: Multi-functional global organization (Command Center, ITSM, Production Ops, Platform Engineering, Reliability Engineering, Reporting)
Geographic Coverage: Global (24x7 follow-the-sun model)
KEY RESPONSIBILITIES:
1. Service Reliability & Executive Accountability
Serve as the single accountable executive for enterprise service stability, client-impacting incidents, and operational resilience outcomes.
Lead a closed-loop service health model covering detection, governance, execution, and prevention.
Establish KPIs and performance discipline aligned to enterprise priorities including Availability, MTTR, incident reduction, and client experience.
2. Command Center Leadership (Enterprise Control Tower)
Lead the Operational Command Center (OCC/NOC) responsible for real-time monitoring, detection, and incident command.
Operate a control tower model delivering integrated oversight across detection, triage, and response workflows.
Ensure rapid decision-making and coordinated response for major incidents with clear ownership and escalation protocols.
3. IT Service Management (Governance & Risk Discipline)
Oversee enterprise Incident, Problem, and Change Management processes with strong adherence to governance, audit, and regulatory expectations.
Ensure root cause elimination and risk mitigation via effective problem management practices.
Maintain ITSM as the governance spine of production operations.
4. Production Operations (Execution & Stability)
Lead global production operations ensuring availability, performance, and operational integrity of the enterprise platform.
Oversee execution of operational workflows including patching, batch processing, and recovery activities.
Ensure strong alignment between change governance and operational execution.
Work in the coordination and collaboration with other Ops functions sitting in the CIO and Shared Service orgs.
5. ServiceNow Platform Strategy & Automation
Lead ServiceNow as a strategic enterprise platform, operating as a system of record, workflow engine, and intelligence layer.
Drive adoption of automation, AI, and workflow orchestration to improve scale, efficiency, and responsiveness.
Ensure strong data integrity (CMDB, service mapping) supporting operational decision-making.
6. Reliability Engineering & Continuous Improvement
Establish proactive reliability engineering capabilities focused on trend analysis, service stability, and incident prevention.
Drive continuous improvement initiatives to reduce recurring failures and optimize client-facing services.
Enable transition from reactive operations to automation-driven resilience.
7. Client Experience, Reporting & Executive Communications
Deliver transparent, data-driven reporting on SLA/SLO performance and client impact.
Provide executive-facing dashboards and insights supporting strategic decisions.
Lead communications strategy during high-severity incidents to senior stakeholders and impacted clients.
Risk, Control & Governance Responsibilities (TIAA-Specific Emphasis)
Ensure alignment with enterprise risk management, regulatory requirements, and audit standards
Maintain strong change control, incident documentation, and traceability across all workflows
Partner with Risk, Compliance, and Audit teams to ensure control effectiveness and remediation tracking
Promote a culture of accountability, transparency, and operational discipline
Leadership & Talent Expectations
Build and lead a high-performing, globally distributed organization
Drive a culture of ownership, continuous improvement, and client-first thinking
Attract and develop talent across operations, platform engineering, and reliability disciplines
Maintain appropriate span of control across functional leads (typically 59 direct leaders)
Key Performance Measures
Enterprise MTTR and incident response effectiveness
Reduction in incident volume and recurrence
Percentage of automated resolutions and workflow efficiency
Change success rate and operational risk reduction
Client-impacting incident metrics and service stability
REQUIRED QUALIFICATIONS: