Senior Machine Learning Engineer, AISWP (Hybrid)
Cisco
The application window is expected to close on: 10/29/2026 This is a hybrid position based out of Cisco's Seattle or San Jose office. Meet the Team The Cisco AI Research team brings together AI researchers, machine learning engineers, data engineers, and networking domain experts to build the next generation of AI-powered networking. We work at the intersection of generative AI, large-scale data systems, and networking, developing Large Language Models (LLMs), agents, and domain-specific AI systems.
Our work spans research and engineering, with a strong focus on translating advances in AI into scalable systems and real-world impact. Your Impact As a Senior Machine Learning Engineer, you will build and improve the data and ML systems that power our LLMs and AI models. A major focus of this role is solving one of the most important challenges in modern AI: creating high-quality training and evaluation data at scale.
You will design and build scalable data pipelines, improve human data labeling workflows, create synthetic datasets, and develop automated approaches for continuously measuring and improving dataset quality. This is a hands-on technical role at the intersection of machine learning engineering and data engineering. You will work closely with researchers, engineers, and domain experts to determine what data our models need, how to create it efficiently, and how to measure its impact on model performance.
What You’ll Do Design, build, and own end-to-end data pipelines for ML and LLM training, post-training, evaluation, and continuous model improvement. Build and improve human-in-the-loop data labeling pipelines, including data collection, task generation, annotation workflows, quality control, and feedback integration. Develop scalable approaches for synthetic data generation, augmentation, filtering, and validation to improve dataset quality, diversity, and coverage.
Apply LLMs and other ML techniques to automate data generation, labeling, filtering, scoring, and evaluation. Develop systems and metrics to continuously measure dataset quality, including label quality, duplication, contamination, coverage, bias, and distribution shifts. Build reliable and scalable infrastructure for processing large volumes of structured and unstructured data. Design experiments to understand how dataset quality and composition affect downstream model performance, and use those insights to drive improvements.
Partner with researchers and ML engineers to develop datasets for model training, fine-tuning, preference learning, evaluation, and agent development. Translate research ideas and prototypes into robust, production-ready data and ML systems. Make architectural and technical decisions across data pipelines, ML infrastructure, storage, compute, and evaluation systems. Provide technical leadership through design reviews, engineering best practices, mentoring, and cross-functional collaboration.
Minimum Qualifications Bachelor’s degree in a STEM field with 7+ years of relevant experience, OR Master’s degree in a STEM field with 4+ years of relevant experience, OR PhD in STEM or a relevant technical field with 1+ years of industry or academic research experience. 2+ years of hands-on experience building, curating, and scaling datasets for machine learning training and evaluation.
3+ years of professional programming experience using Python, C++, or Go within a production or research environment. 3+ years of experience using machine learning frameworks such as PyTorch, TensorFlow, or equivalent technologies to develop, train, evaluate, and deploy machine learning models. Preferred Qualifications Expertise in curating, scaling, and managing datasets for the entire LLM lifecycle—including synthetic data generation, augmentation, and post-training workflows like SFT and RLHF.
Proficiency in designing human-in-the-loop labeling systems and proactively mitigating complex dataset failure modes such as label n