Senior Lead Software Engineer - ML Engineer for Agent Platform
Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products.
As a Senior Lead Software Engineer - ML Engineer for Agent Platform at JPMorgan Chase within the Commercial and Investment Banking – Data Analytics Payments Team, you are a senior technical leader on the team that builds and runs NEO, the firm's agent runtime platform for Payments Technology. You set the architecture for how agents execute, communicate, remember, and get evaluated in a secure, stable, and scalable way. As a core technical contributor and technical direction-setter, you are responsible for the hardest technology decisions across multiple technical areas in support of the firm's business objectives, and for raising the engineering bar across the teams that build on NEO.
Job responsibilities
- Owns end-to-end architecture of the NEO agent runtime, including secure execution and isolation (micro-VMs such as Firecracker/Kata), agent-to-agent (A2A) communication, Model Context Protocol (MCP) tooling, the memory layer, and evaluation infrastructure
- Executes creative software solutions, design, development, and technical troubleshooting with the ability to think beyond routine or conventional approaches to build solutions or break down technical problems
- Develops secure and high-quality production code, and reviews and debugs code written by others; sets code and design standards adopted across teams building on NEO
- Drives team adoption of enterprise-authorized AI-assisted engineering practices to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team
- Applies and shapes the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation
- Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of the platform and the agents running on it
- Defines the permission-aware, auditable execution model for the runtime, including fine-grained authorization (OpenFGA) and runtime policy (OPA/Rego), so agents operate safely in a regulated environment
- Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
- Leads communities of practice across Software Engineering to drive awareness and use of new and leading-edge technologies, and mentors Lead and senior engineers
- Adds to team culture of diversity, opportunity, inclusion, and respect
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Hands-on practical experience delivering system design, application development, testing, and operational stability for platforms, runtimes, or distributed systems in production
- Advanced in one or more programming language(s); strong Python plus a systems language (Go or Rust) for performance-sensitive runtime work
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
- Demonstrated experience building or operating LLM/agent systems in production, including tracing, evaluations, and guardrails
- Proficient in all aspects of the Software Development Life Cycle
- Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security
- Demonstrated proficiency in software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning)
- In-depth knowledge of the financial services industry and their IT systems
- Practical cloud native experience; production Kubernetes expected
Preferred qualifications, capabilities, and skills
- Experience architecting secure code execution and sandboxing with micro-VMs (Firecracker, Kata, gVisor) for multi-tenant isolation
- Experience designing or implementing agent protocols (A2A, MCP) and multi-agent orchestration
- Experience with agent or distributed memory systems — memory nodes, episodic/semantic memory, and graph-backed retrieval (Graph RAG)
- Exposure to LLMs, RAG architectures, vector databases, and embedding-based retrieval systems
- Fine-grained authorization (OpenFGA / Zanzibar-style) and policy engines (OPA/Rego)
- Evaluation infrastructure for agents — offline/online evals, regression suites, and LLM-as-judge quality/safety gating in CI
- Proficiency with Infrastructure as Code (Terraform) and containerized deployments (Docker, Kubernetes)