Senior AI Data Engineer, Data Products & RAG Foundations
Agilent
Job Description Agilent helps laboratories around the world advance scientific discovery, diagnostic s, and applied market solutions through instruments, software, consumables, services, and deep domain expertise . About the role: As a Senior AI Data Engineer, Data Products & RAG Foundations , you will be a data engineering SME within a cross-functional AI pod, working alongside AI engineers, domain experts, business stakeholders, data owners, and platform teams.
Your role is to build data products, pipelines, metadata, and retrieval-ready assets that power AI-enabled business and scientific workflows across the enterprise. Pods do not wait for the enterprise data foundation to be complete; they help build it through execution . Every data product created by the pod is designed for governance, reuse , and long-term value, with the next consumer in mind from day one.
Th is role goes beyond traditional da ta engineering. You will work with structured and unstructured data, semantic definitions, quality scoring , lineage, contracts, embeddings, vector search, and retrieval foundations for AI systems. You will also leverag e AI-assisted techniques, such as metadata generation, entity resolution, and conte nt classification, to create trusted, AI -ready data products at scale .
You do not need prior experience with Agilent’s internal data architecture. We are looking for a strong data engineer who understands data quality, governance, and AI-ready data foundations and is excited to help shape the future of enterprise AI at Agilent. What you will do: Data Products & Governance Build and maintain AI-ready data products and pipelines for the pod's use case, ensuring appropriate governance , lineage, metadata, access controls, and documentation from the start.
Design data products for reuse, treating every asset as a potential enterprise capability rather than a point integration. Da ta Quality and Trust Estab lish d ata quality standards , quality scoring , and model-readiness criteria that support reliable AI behavior and business outcomes. Ensure q u a lity issues are identified and addressed before the y impact downstream AI solutions.
Domain Understanding and Partnership Partner with data owners , stewards, business stakeholders, and IT teams to establish trusted definitions , authoritative sources , and domain data models. Ensure AI solutions are ground ed in validated business meaning rather than co nvenience-based access to data . Retrieval and AI Foundations Design r etrieval foundations that support AI applications , including structured and unstructured grounding, vector search, graph -based approaches, and semantic enrichment where appropriate .
Apply AI- assisted techniques such as metadata generation, entity resolution, and content classification to improve the quality, scalability, and discoverability of data assets . Eng ineering Delivery and Reuse Design and implement scalable ingestion, integration, and storage frameworks across cloud and on-premises environments. Build reusable data assets, tools, and services that support AI engineers, data scientists, and analytics teams.
Contribute reusable data products, patterns, and documentation back to the broader enterprise ecosystem. What success looks like in the first year The pod's use case is running entirely on governed, quality-scored data products, with no undocumented or unsupported data source s . Multiple data products created by the pod have been adopted, reused, or identified for reuse across additional AI or analytic s use cases.
Data q uality signals are integrated into AI evaluation and monitoring processes , influencing