Data Engineering Lead - Data Quality Systems
jobgether
India
Posted Sep 3, 2026
- Full-time
- Remote
- Security & IT
Job description
**This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Data Engineering Lead - Data Quality Systems based in India.** Lead the engineering of data quality systems that determine whether billions of records can be trusted by customers. You will combine deep hands-on engineering with technical leadership, spending approximately 80% of your time building and 20% leading a small team. Your scope will include verification pipelines, anomaly detection, scoring frameworks, LLM evaluation, and automated release gates. You will tackle complex data-quality challenges across multiple markets and large-scale production environments. The role offers significant ownership, with greenfield opportunities to establish frameworks and engineering standards from the ground up. You will work in an AI-native environment where agentic development, evaluation, observability, and automation are core engineering practices. This is an ideal opportunity for a technically strong leader who wants direct ownership of a critical data trust layer while remaining deeply involved in the code. ### Accountabilities - Architect and build continuous data-quality systems, including verification, sampling, scoring, and reconciliation pipelines operating across billions of company and people records. - Design reusable frameworks, abstractions, and technical specifications that allow engineers to create quality checks efficiently, reliably, and consistently. - Build evaluation harnesses for LLM-powered validation and extraction, including labeled evaluation sets, precision/recall measurement, judge calibration, prompt versioning, and model-drift detection. - Establish pre- and post-production release gates that identify and prevent poor-quality data from reaching customers, supported by effective failure analysis and triage tooling. - Investigate large-scale data-quality incidents, identify root causes, implement corrective solutions, and convert recurring failures into permanent automated checks. - Lead a team of 3–5 Applied AI Engineers through technical direction, code reviews, pairing, mentoring, and development of end-to-end ownership. - Set and maintain a high technical standard while remaining approximately 80% hands-on in engineering and architecture. - Apply sound judgment when choosing between deterministic rules and LLM-based validation, using structured rules where appropriate and semantic models where they add value. - Operate LLM-based quality systems as production infrastructure, with appropriate evaluation, traceability, prompt and model versioning, cost controls, and performance monitoring. - Contribute to an AI-native engineering culture based on agentic development, automated evaluation, logged traces, AI-assisted review, and reusable workflow specifications. - Establish scalable engineering practices in a lean environment characterized by high ownership, minimal process overhead, and frequent production releases. ## Requirements - 7+ years of experience building production-grade data systems in business-critical environments, including systems that operate reliably at significant scale. - Demonstrated experience working with billions of data rows and designing quality controls that remain performant and dependable at large scale. - Proven track record of **building** data-quality systems and frameworks, such as validation engines, anomaly detection, scoring systems, sampling strategies, or reconciliation mechanisms against trusted data. - Experience designing evaluation or test harnesses that are used by other engineers and can support systematic measurement of quality. - Previous experience providing technical leadership to engineers, including code reviews, technical direction, pairing, mentoring, and hands-on delivery. - Strong Python development skills and advanced SQL expertise, with an understanding of performance optimization, concurrency, and large-scale data transformations. - Practical experience operating LLMs as production systems, including evaluation sets, versioned prompts, trace logging, cost controls, and debugging model judges against precision and recall. - Proven experience using agentic development environments such as Claude Code, Cursor, or equivalent tools to build and ship production software. - Strong technical judgment regarding when to use deterministic rules versus LLM-based semantic evaluation, with the ability to clearly justify architectural decisions. - Experience with B2B data, including firmographics, people data, entity resolution, or registry matching across multiple markets, is highly valued. - Familiarity with cloud data platforms such as Snowflake, Databricks, or Redshift, together with AWS-based pipeline deployment, is advantageous. - Production-scale experience with Airflow or an equivalent orchestration platform is a plus. - Knowledge of vector databases, embeddings, retrieval patterns, matching, or deduplication is desirable. - Startup or scaleup experience, particularly in environments where engineering standards and frameworks had to be established from the ground up, is highly valued. - Strong ownership, judgment, adaptability, and communication skills suited to a fast-moving, autonomous, and highly collaborative engineering environment. ## Benefits - Fully remote position based in India. - Competitive base salary aligned with the seniority and technical scope of the role. - Meaningful equity participation and the opportunity to share in the organization's long-term growth. - Significant technical ownership over a critical data-quality and trust layer. - Greenfield engineering opportunities to define frameworks, standards, validators, evaluation systems, and release gates. - Exposure to frontier engineering challenges involving LLM evaluation, model drift, agentic development, anomaly detection, and large-scale data quality. - Opportunity to lead a small, senior engineering team while remaining deeply hands-on technically. - Lean, high-autonomy environment with minimal management layers and strong end-to-end ownership. - AI-native engineering practices, with agentic development, evaluations, traces, and AI-powered review integrated into everyday workflows. - Fast release cycles and the opportunity to make visible contributions across multiple international markets. - Equal-opportunity environment that values diverse perspectives and inclusive collaboration.