Vice President, Production Services Application Support
In this role, you’ll make an impact in the following ways:
In this role, you will play a critical part in protecting production stability, accelerating service recovery, and driving continuous improvement across mission-critical platforms.
Own end-to-end support for mission-critical applications, ensuring high availability, stability, resiliency, and timely incident resolution.
Lead complex technical triage across application, database, middleware, messaging, batch, and infrastructure layers to restore service quickly and effectively.
Support payments applications, including transaction flow validation, issue analysis, health checks, and operational readiness.
Drive incident and problem management, root cause analysis, service restoration, and permanent fix follow-up for high-impact production issues.
Partner across application development, DBA, infrastructure, network, middleware, operations, and business teams to resolve issues, reduce risk, and improve reliability.
Monitor application health and identify risks using Splunk, Grafana/Optics, AppDynamics, Moogsoft, and other observability tools.
Support batch scheduling, recovery, and operational validations using Control-M and approved support procedures.
Execute and validate deployment, release, and change activities through GitLab / CI/CD pipelines aligned to change standards.
Champion automation using Ansible and UNIX/Linux scripting to reduce manual effort, improve consistency, and strengthen operational controls.
Support MQ and Kafka platforms to ensure stable integration and operational continuity.
Maintain support documentation, runbooks, knowledge articles, escalation guides, and recovery procedures.
Participate in production readiness reviews, resiliency testing, disaster recovery exercises, and failover validations.
Use ServiceNow for incident, problem, change, service request, and follow-up tracking.
Collaborate with global teams and deliver clear, confident updates during incidents, bridge calls, and leadership communications.
Leverage AI tools such as Microsoft Copilot to elevate documentation, analysis, reporting, knowledge management, and team productivity.
To be successful in this role, we’re seeking the following:
Bachelor’s or higher degree in computer science, engineering, or a related discipline, or equivalent work experience.
10+ years of proven experience in production support, application support, or technology operations.
Demonstrated success supporting enterprise-scale, mission-critical applications in complex production environments.
Knowledge of payments domain flows, transaction processing, and production issue analysis.
Hands-on Oracle SQL experience for querying, troubleshooting, data validation, and issue investigation.
Strong UNIX/Linux skills, including log analysis, file system checks, process monitoring, and command-line troubleshooting.
Hands-on experience with Ansible for automation and operational task execution.
Knowledge of AppEngine and application runtime support.
Experience with GitLab and CI/CD pipelines for deployment, release, and validation activities.
Experience with Control-M or similar enterprise batch scheduling tools.
Experience with Splunk, Grafana/Optics, AppDynamics, Moogsoft, and related observability platforms.
Experience using ServiceNow for ITSM, including incident, problem, change, and service request management.
Proficiency in messaging technologies such as MQ and Kafka.
Exposure to AI-enabled productivity tools such as Microsoft Copilot.
Strong understanding of incident, problem, and change management, production readiness, and operational risk controls.
Ability to analyze logs, alerts, batch failures, messaging issues, transaction breaks, infrastructure events, and application errors to identify root cause and remediation.
Excellent communication and stakeholder management skills, with the ability to translate technical issues into clear updates for technology, operations, business, and leadership stakeholders.
Ability to stay composed under pressure, take ownership during critical incidents, and drive issues to closure with urgency and accountability.