Lead Software Engineer- Python / Pyspark / Java / BigData / Data Modernization / AI
We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.
As a Lead Software Engineer- Python / Pyspark / Java / BigData / Data Modernization / AI, at JPMorganChase within the Asset and Wealth Management- Global Prime Brokerage Team, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.
We look for people who are passionate about solving business problems through innovation, analytics, and an AI‑first engineering mindset—building reusable, governed analytical data products and accelerating delivery of regulatory and CEO‑priority analytics. You will define and enforce an AI-driven data product lifecycle (semantic alignment, automated lineage and data quality, pipeline/code generation, mesh registration, and self-service consumption) and will build human-in-the-loop autonomous agents to detect schema drift, propose transformations, reconcile semantics, triage data incidents, and generate governance evidence. You’ll be required to apply your depth of knowledge and expertise to all aspects of the analytics development lifecycle, and partner continuously with stakeholders across product, platform, risk, and domain teams. You will lead an AI‑first transformation of data engineering and analytics by productizing the data product lifecycle (semantics, lineage, DQ, governance) and building autonomous agents (human‑in‑the‑loop) that reduce manual toil, improve auditability, and enable self‑service consumption on the strategic data mesh. The role also owns modernization of the strategic data mesh.
Job responsibilities
- Collaborate with business and technology teams to develop AI‑first analytics and data product solutions
- Define and enforce architecture for an AI‑driven data product lifecycle: semantic extraction/alignment, automated lineage and DQ, pipeline code generation, mesh registration, and self‑service consumption
- Build and operate autonomous agents for data engineering that detect schema drift, propose transformations, reconcile semantics, triage data incidents, and maintain governance evidence under human‑in‑the‑loop controls
- Design analytics platforms capable of running reporting and other analytics; explore innovative ideas by building real‑time and batch analytics solutions
- Establish appropriate monitoring and alerting of solution events related to performance, scalability, availability, and reliability
- Provide technical leadership, guidance, and direction to other team members; build prototypes for demonstrations for peer groups, business partners, and senior leaders
- Drive team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team
- Apply knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation
- Lead migration and modernization from legacy analytics/reporting stacks to the strategic mesh ecosystem (e.g., Databricks/Iceberg/common services), reducing fragmentation and duplicated data products
- Industrialize entity resolution and parent identification with ML/LLM solutions and standardize analytical product packaging to enable reuse and monetization
- Embed governance, lineage, and DQ by design across critical domains and regulatory reporting, improving auditability and control posture
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Proven leadership delivering AI‑first analytics and data engineering at scale, including productized data mesh patterns, semantic layers, and analytical data product lifecycle ownership
- Experience developing data ingestion and integration processes, sourcing data from multiple platforms, and applying data cleansing/transformation rules for analytics-ready datasets
- Deep hands‑on experience with big data and modern data platforms (e.g., Spark, Databricks, Snowflake, Iceberg) and building robust pipelines and data lake/lakehouse frameworks
- Strong programming capability in Python and PySpark, or Java, with strong CI/CD and containerization practices
- Applied AI expertise in ML pipelines, NLP/LLMs, and agentic frameworks to build autonomous agents for engineering tasks (schema drift detection, semantic reconciliation, incident triage, governance evidence generation) under human‑in‑the‑loop controls
- Governance proficiency across lineage, data quality, and access control with evidence generation aligned to regulatory expectations (e.g., BCBS 239‑class lineage/DQ)
- Comfortable working in an agile and collaborative environment; strong written and verbal communication skills
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
Proficient in all aspects of the Software Development Life Cycle
- Python and Java