Location: New York, NY
Work arrangement: Hybrid
Work type: Full-time
Compensation: $170000 – $300000/yr
This role is pivotal in bridging the gap between next-generation tooling and the operational realities of a large-scale enterprise. The ideal candidate will have a strong background in Site Reliability Engineering (SRE) and setting direction for a portfolio of applications in Production Management to impact mission-critical production systems across the enterprise. You will thrive in an environment with ambiguous requirements and fixed deadlines, using your expertise to create tactical solutions for immediate challenges while simultaneously developing a long-term strategic vision for system resilience and efficiency. The role will require interfacing and influencing managing director and C-suite level stakeholders.
Key Responsibilities
Lead complex migration programs from legacy systems to new, strategic platforms. Examples include management of application recovery plans, application green zones, and batch management workflows while simultaneously enhancing their quality and reliability.
Define and evangelize a multi-year transformation roadmap, aligning technology investments with business priorities through thorough gap analysis between old and new systems.
Develop and execute a dual-horizon strategy that includes tactical approaches to meet fixed deadlines, alongside a long-term vision for tool adoption, process optimization, and automation.
Drive the operationalization of new processes and best practices. This includes onboarding teams to new platforms and ensuring adherence to standards.
Act as the primary liaison between production teams and lead engineers. Establish and manage a continuous feedback loop to inject requirements, report defects, and influence roadmaps.
Translate vague business and operational problems into clear, actionable requirements for internal and external engineering teams.
Define, implement, and manage data models, KPIs, and scorecards to track migration progress, system quality, and adoption.
Drive accountability across cross-functional teams to meet critical deadlines and objectives.
Proactively identify and mitigate risks associated with migrating to immature platforms. Develop and implement tactical solutions to overcome tooling inefficiencies and process bottlenecks, ensuring program milestones are met.
Required Qualifications
15+ years' experience in Site Reliability Engineering (SRE), Production Management, or a similar role with a focus on operational excellence and system resilience in a leadership capacity.
Demonstrated experience managing large-scale technical programs, particularly system migrations or technology transformation initiatives.
Strong understanding of modern platforms, and ability to implement technical designs for high availability, disaster recovery, and operational resilience.
Proven ability to navigate ambiguous requirements in a regulated enterprise environment, take ownership, and create structure and direction for a program.
Ability to operate at both a strategic level (Phase 2 planning, long-term vision) and a tactical level (finding workarounds, hitting immediate deadlines).
Excellent communication skills with the ability to define clear goals, hold teams accountable, and articulate complex technical challenges to a wide range of stakeholders for strategic decision-making.
Familiarity with the Software Development Lifecycle (SDLC) and experience implementing best practices in a production environment.
Build, mentor, and lead a team of Technical Project Managers, establishing standards of excellence for program execution across the portfolio.
Preferred Qualifications
Experience defining and tracking program success through data-driven scorecards and KPI metrics to ensure teams are delivering against measurable outcomes.
Hands-on experience with scripting, automation, or data modelling.
Experience working within an Agile development framework and providing feedback to project/product/engineering teams.
Familiarity with platform engineering and DevSecOps practices e.g., CI/CD, GitOps, containerization, cloud migration patterns.
Education
Bachelor's degree in Computer Science or a related field, or equivalent experience building scalable, resilient solutions to improve service reliability and operational efficiency.
Compensation: $170000 – $300000/yr