Senior Associate-Technology Operations Engineering

Bengaluru, IndiaFull-timePosted Jul 22, 2026

This role is responsible for driving the reliability, resiliency, performance, and modernization of critical American Express platforms across both Mainframe and Distributed environments. You will leverage deep technical expertise in software engineering, runtime engineering, production support, and platform operations to quickly assess and remediate complex availability, performance, and operational issues.
 

As part of our technology team, you will partner with engineering, product, infrastructure, and operations teams to design, build, automate, and support highly available enterprise platforms. You will help accelerate modernization initiatives, improve operational excellence, and deliver secure, scalable, and resilient solutions that power critical customer and business capabilities.

Software Engineering & Platform Development
  • Serve as a hands-on engineer with expertise in supporting complex enterprise applications, platforms, and operational tooling across Mainframe and Distributed environments.
  • This role requires and must have 5+ years of mainframe experience and knowledge, understanding and experience of distributed application is a plus.
  • Design, develop, prototype, code, test, and implement scalable software solutions using technologies such as Java, Python, COBOL, REXX, JCL, SQL, and related frameworks.
  • Act as a technical contributor in change management reviews, root cause analysis, and troubleshooting of complex technical issues.
  • Design and implement automation solutions, and engineering practices that improve platform resiliency, operational efficiency, and security.
  • Use best practices in incident management, problem management and change management as this role focuses on application production support.
 Runtime Engineering, Reliability & Operations
  • Drive the technical roadmap for runtime systems, ensuring platform reliability, scalability, availability, recoverability, and performance.
  • Establish, monitor, and continuously improve key performance indicators (KPIs), service level objectives (SLOs), and operational metrics (MTTR, MTBF) for platform health and resiliency.
  • Lead diagnosis and resolution of production incidents, batch failures, application outages, performance bottlenecks, and infrastructure issues across Mainframe and Distributed platforms.
  • Apply Site Reliability Engineering (SRE) principles and operational excellence practices to improve system stability and reduce operational risk.
  • Support disaster recovery, high availability, workload management, capacity planning, and business continuity initiatives.
 Mainframe & Distributed Platform Engineering
  • Must have Experience in Support and optimize enterprise platforms across IBM Mainframe technologies including z/OS, JES2/JES3, TSO/ISPF, SDSF, JCL, DB2 for z/OS, IMS, VSAM, MQ, CICS, RACF, and related ecosystem tools.
  • Desirable Support and optimize Distributed Platform technologies including cloud infrastructure, Linux/Unix, containers, APIs, Java-based services, distributed databases, and modern application platforms.
  • Implement and support hybrid Mainframe and Cloud/Distributed architectures, modernization initiatives, API enablement, and enterprise integration capabilities.
  • Knowledge of cloud platforms AWS, GCP or general Cloud fundamentals.

 

Data, Integration & Automation
  • Develop and support enterprise integration solutions utilizing APIs, MQ, Connect:Direct, event-driven architectures, and batch and real-time data integration patterns.
  • Utilize relational and NoSQL databases including DB2, PostgreSQL, Redis, Couchbase, IMS DB, and VSAM to support critical business applications.
  • Automate operational processes, deployments, monitoring, reporting, and remediation activities using REXX, Python, Bash, Ansible, Jenkins, and related technologies.
  • Empower engineering teams to adopt scalable automation and self-service capabilities for deployment, monitoring, and operational support.
 Observability & Continuous Improvement
  • Implement and utilize monitoring, observability, logging, and analytics solutions using tools such as Splunk, OMEGAMON, RMF/SMF, Sysview, MainView, Dynatrace, AppDynamics, ELK, or equivalent technologies.
  • Analyze operational trends, identify opportunities for optimization, and formulate strategic recommendations to improve platform health and engineering effectiveness.
  • Drive continuous improvement initiatives focused on reliability, performance, security, operational maturity, and customer experience.
 DevOps, Security & Governance
  • Understand CI/CD pipelines and DevOps practices using tools such as Git, Jenkins, Maven, DBB, Endevor, Changeman, ISPW, UrbanCode Deploy, or equivalent platforms.
  • Apply enterprise security controls, compliance requirements, audit standards, and access management practices, including RACF, ACF2, Top Secret, and cloud security principles.
  • Ensure solutions meet non-functional requirements (NFRs) including availability, scalability, performance, security, recoverability, and maintainability.
  • Bachelor's degree in Computer Science, Computer Engineering, and/or comparable experience
  • Work experience in software engineering, app support or infrastructure operations or runtime engineering.
  • A working understanding of cloud infrastructure, distributed systems, and containerization technologies, with experience in supporting critical business applications being a plus.
  • Familiarity with monitoring and logging tools, and incident management best practices, to ensure reliability and performance of applications in a production environment.
  • Experience in JCL, COBOL & DB2 is preferred.
  • Exposure to scheduling tool (Control-M or similar), IBM mainframe utilities, SORT utilities, File Aid, Abend Aid, SPUFI, QMF , mainframe performance product (BMC’s Apptune or similar) is preferred.
  • Solid programming and scripting skills, with hands on experience to automate operational tasks using tools such as Python, Ezytrieve, REXX 
  • Knowledge of scripting languages (e.g., PowerShell, Python) for automation tasks
  • Experience in technology operations work
  • Hands on experience with relational and NoSQL databases such as DB2, Redis, Postgres, Couchbase etc.
  • Experience in cloud platforms such as AWS, Azure, or Google Cloud, Public Cloud certification is a plus

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free