Sr Lead Software Engineer - AWS - Lead AI/ML Platform Engineer
Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products.
As a Senior Lead Software Engineer at JPMorganChase within the Firmwide AI/ML Deployment Platform team, you are an integral part of a globally distributed team — spanning Glasgow, London, New Jersey, and India — that works to architect, build, and own the infrastructure that makes model deployment work at scale. You'll operate with significant autonomy: owning technical direction, engaging directly with US-based clients, and making architectural decisions with real production consequences. We build the control plane, APIs, monitoring, and deployment infrastructure that internal teams depend on. The platform is always evolving — new regions, new failure modes, new scale requirements. If you like owning problems end-to-end, making hard tradeoffs, and shipping systems that other engineers build on top of, you'll fit in.
Job responsibilities
- Drive architectural vision for platform components: control plane integration, multi-region deployment, and disaster recovery
- Design and implement APIs for retraining, scheduling, endpoint deployment, and autoscaling
- Build infrastructure for seamless integration across control plane and client accounts
- Engage directly with US-based clients — requirements, strategic solutioning, and debugging
- Make independent architectural decisions and own technical tradeoffs with minimal oversight
- Regularly provides technical guidance and direction to support the business and its technical teams, contractors, and vendors
- Drives decisions that influence the product design, application functionality, and technical operations and processes
- Drives adoption and governance of approved AI-assisted engineering practices across teams to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test acceleration, release readiness, incident/root-cause analysis), while establishing measurable validation standards (secure coding, peer review, automated testing) and promoting reuse of proven patterns and automation within the SDLC/TLM toolchain.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including approved AI-assisted development and automation capabilities, to improve the value realized by automation at scale.
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Proficiency in architecting software solutions at scale — you've designed systems that other teams depend on
- Self-directed and autonomous: you drive to outcomes without waiting for direction
- Strong client-facing communication skills, effective across time zones in a distributed team
- Deep knowledge of AWS and cloud-based infrastructure
- Track record building resilient, production-grade platform
- Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching senior engineers/leads on compliant usage patterns and controls.
Preferred qualifications, capabilities, and skills
- Comfort with ambiguity and greenfield architecture where no existing playbook applies
- Production experience with Kubernetes / EKS at scale
- Hands-on experience with AWS Sagemaker for model training and deployment
- Strong Golang skills in the context of infrastructure or platform services
- Deep understanding of networking — VPCs, DNS, cross-account connectivity
- Practical experience with LLMs — deployment, inference, or integration
- Track record delivering Terraform across multi-account, multi-region environments