Data Engineer III
Enigma is an analytics-focused application that transforms ServiceNow-sourced ITSM data in Snowflake into curated datasets and delivers reporting, search, and insights. We’re expanding APIs for broader firm consumption and planning a revised data architecture—while preserving the fast search experience users value today.
We’re hiring a mid-level Data Engineer to build curated analytical datasets and implement high-performance “data serving” patterns for search and APIs. This role will also contribute to evaluating architecture options (e.g., Snowflake-only vs. a dedicated serving/search layer) based on latency, cost, and operational fit.
Job Responsibilities
- Build and operate ELT pipelines into Snowflake from ServiceNow-derived sources, including incremental loads, backfills, and reprocessing.
- Develop and maintain analytics models and canonical definitions (metrics, dimensions, incident taxonomy, time-windowed aggregates).
- Design low-latency data-serving patterns for search and APIs, such as:
- denormalized/search-optimized tables and materialized aggregates
- precomputed facets/filters and pagination-friendly query patterns
- strategies to minimize expensive joins at request time
- Contribute to architecture decisions for search performance:
- define measurable non-functional requirements (e.g., p95/p99 latency targets, concurrency, freshness)
- run/assist PoCs and benchmark approaches (warehouse-only vs. dedicated serving/search layer)
- document tradeoffs across latency, cost, complexity, and operational risk
- Implement data quality and observability: freshness/completeness checks, schema drift detection, reconciliation, monitoring/alerting, and clear SLAs/SLOs.
- Partner with application developers to create stable, versioned data surfaces for APIs (contract-friendly schemas, safe schema evolution).
- Create and maintain runbooks and support processes; participate in incident triage where data quality or serving performance is involved.
- Build and maintain an MCP (Model Context Protocol) server—a standardized adapter that connects AI applications (e.g., LLMs/agents) to external tools, APIs, and data sources by translating AI intent into actionable system commands.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate data pipeline/design analysis and documentation, validating outputs and handling data according to sensitivity and security requirements.
- Applies reuse-first, AI-assisted practices to strengthen SDLC-quality routines for data pipelines (e.g., test generation and control validation), ensuring traceability/auditability and alignment to resiliency and security expectations.
Required qualifications, capabilities, and skills
- 3–6 years (or equivalent) experience delivering production data pipelines and curated analytical datasets.
- Strong SQL and analytics modeling skills (incremental patterns, snapshots, aggregates).
- Proficiency in Python (preferred) and/or Java for transformations, validations, and automation.
- Experience with Snowflake, including performance/cost awareness, query tuning, and secure access patterns.
- Demonstrated ability to design for performance (benchmarking, query plan reasoning, caching/materialization strategies).
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support data engineering workflows with strong validation habits and awareness of data sensitivity.
- Ability to review and validate AI-assisted outputs (e.g., query suggestions, test ideas, or model change summaries) before use, escalating when uncertain and following data handling requirements.
- Advanced English skills
Preferred qualifications, capabilities, and skills
- Experience building integrations/servers that connect applications to tools/APIs/data sources (experience with MCP specifically is a plus).
- Experience with low-latency serving/search technologies (e.g., Elasticsearch/OpenSearch, Postgres-based serving layers, etc.).
- Familiarity with ITSM / ServiceNow datasets.
- Experience supporting data products used by multiple downstream teams (documentation, versioning, consumer enablement/change management).