Principal Software Engineer- Telemetry/Observability
If you are looking for a game-changing career, working for one of the world's leading financial institutions, you’ve come to the right place.
As a Principal Software Engineer at JPMorganChase within the Core Foundational Platforms, you provide expertise and engineering excellence as an integral part of an agile team to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. Leverage your advanced technical capabilities and collaborate with colleagues across the organization to drive best-in-class outcomes across various technologies to support one or more of the firm’s portfolios.
Job responsibilities
- Creates complex and scalable coding frameworks using appropriate software design frameworks, with a focus on observability/telemetry platforms for Virtual Private Cloud (VPC) environments
- Develops secure and high-quality production code, and reviews and debugs code written by others across telemetry collection, enrichment, storage, and visualization services
- Leads the design and implementation of end-to-end telemetry pipelines (metrics, logs, traces, events) including standards, schemas, and instrumentation patterns for platform and application teams
- Engineers and evolves CI/CD pipelines to support reliable delivery, automated testing, policy-as-code, and safe progressive rollouts for observability services
- Designs and implements solutions that operate effectively in software-defined networking (SDN) and cloud networking environments, including troubleshooting across layered network paths and dependencies
- Builds and maintains monitoring, alerting, and SLO/SLI practices using tools such as Splunk and/or Grafana and/or Prometheus (including dashboarding, alert tuning, and operational runbooks)
- Advises cross-functional teams on technological matters within your domain of expertise, including telemetry strategy, platform integrations, and operational readiness
- Serves as the function’s go-to subject matter expert for observability and telemetry across the VPC platform, influencing standards and best practices
- Creates durable, reusable software frameworks that are leveraged across teams and functions (SDKs, libraries, templates, reference architectures, and golden paths)
- Influences leaders and senior stakeholders across business, product, and technology teams by communicating tradeoffs, risk, and roadmap decisions for platform reliability and resiliency
- Champions the firm’s culture of diversity, opportunity, inclusion, and respect
Required qualifications, capabilities, and skills
- 10+ years of software engineering experience, including building and operating platform services in production environments
- Hands-on practical experience delivering system design, application development, testing, and operational stability, ideally for platform/infrastructure or developer productivity products
- Expert-level Python development experience, including building services, automation, tooling, and integrations for telemetry and operations workflows
- Strong Linux systems and networking expertise (TCP/IP, DNS, TLS, routing, NAT, packet capture/troubleshooting), and the ability to diagnose cross-system issues in distributed environments
- Practical experience with CI/CD (e.g., automated build/test/release pipelines) and modern delivery practices (progressive delivery, automated quality gates, artifact/version management)
- Advanced knowledge of software application development and technical processes with considerable in-depth knowledge in one or more technical disciplines (e.g., cloud platforms, platform engineering, reliability engineering, networking, observability)
- Experience with observability/telemetry tooling such as Splunk and/or Grafana and/or Prometheus (including instrumentation, queries, dashboards, alerts, and operationalization)
- Experience designing or operating telemetry data pipelines (collection agents, exporters, scraping, indexing, retention, cardinality management, performance/cost controls)
- Experience applying expertise and new methods to determine solutions for complex technology problems in one or more technical disciplines
- Working knowledge of software-defined networking concepts and environments, including integrations and operational considerations for VPC-like network platforms
Preferred qualifications, capabilities, and skills
- Experience building an observability platform or “paved road” telemetry offering used by multiple engineering teams (multi-tenant design, onboarding patterns, guardrails)
- Experience with OpenTelemetry (instrumentation standards, collectors, exporters) and distributed tracing practices
- Strong understanding of SRE principles (SLOs/SLIs, error budgets, incident response, post-incident reviews, toil reduction)
- Experience with infrastructure as code and GitOps-style workflows (e.g., Terraform-like patterns, configuration as code, policy enforcement)
- Experience supporting regulated environments with strong security and control requirements (secure-by-default design, auditability, change management)