Citi is looking for a Senior Data Engineer to design and build the next-generation data processing and analytics platform that sits at the core of our AI-driven business innovation. In this role, you will architect scalable ETL/ELT pipelines using Python and PySpark, while also integrating Generative AI and Machine Learning models to push the boundaries of what our data systems can do. Working within a high-performing engineering team, you will shape how data flows, performs, and delivers value across the organisation at scale.
## Responsibilities
* Design and build highly scalable ETL/ELT pipelines in Python and PySpark to power reliable data ingestion, transformation, and integration at scale.
* Develop production-grade code and elevate team capability through structured code reviews and hands-on peer mentorship.
* Integrate Generative AI and Machine Learning models into data workflows, using AI development tools to deliver innovative solutions and resolve complex technical challenges.
* Establish and enforce data quality standards across all pipelines, ensuring accuracy, consistency, and reliability of data throughout its lifecycle.
* Implement an AI-first engineering approach by exploring and embedding Agentic AI workflows to increase development velocity and platform intelligence.
* Automate CI/CD pipelines for builds and deployments, advancing DevOps maturity across the data platform.
* Apply performance tuning and optimisation techniques to process large volumes of structured and semi-structured data efficiently and at speed.
* Collaborate directly with engineers and product owners to align technical direction with delivery goals and accelerate outcomes.
## Required Qualifications & Skills
* Deep hands-on expertise in Python and Apache Spark, including Spark Core, Spark SQL, and DataFrames/Datasets, applied to production data systems.
* Advanced SQL capability across complex query writing, data transformation logic, and query performance optimisation.
* Demonstrated experience designing, building, and maintaining ETL/ELT pipelines for large-scale data ingestion and transformation.
* Proficiency with Python data processing libraries including Pandas, NumPy, and SciPy for data wrangling and preprocessing tasks.
* Practical experience with DevOps tools and CI/CD pipeline automation, with version control managed through GitHub.
* Solid working knowledge of Linux environments, with familiarity across Windows systems.
* Ability to collaborate across multidisciplinary teams to diagnose complex technical problems and deliver effective, lasting solutions.
## Beneficial Skills & Qualifications
* Working knowledge of Machine Learning and Generative AI, with direct exposure to AI development tooling in an engineering context.
* Familiarity with Agentic AI workflows and their application to enhancing platform capabilities or development productivity.
* Experience using Jupyter Notebook for rapid prototyping and iterative development of data solutions.
## What We Offer
At Citi, you will work on technically challenging problems that matter, contributing to a platform that shapes how a global financial institution processes and derives intelligence from data. This is a role where technical depth, curiosity, and engineering rigour are genuinely valued.
* Hybrid working model with 3 days in the office and 2 days working remotely, giving you flexibility and in-person collaboration.
* Access to complex, large-scale engineering challenges at the intersection of data and AI, keeping your technical skills at the forefront of the industry.
* Opportunities to grow into AI and platform architecture disciplines, supported by a team culture that invests in continuous technical development.
* Exposure to a globally connected engineering network, collaborating with specialists across technology, data, and product functions.
* A performance-driven environment where your contributions to platform...
Want jobs like this matched to you?
Swoopd scores fresh postings against your résumé so you only see the matches that matter.