Research Scientist - Data
ifm-us
Sunnyvale, CA
Posted Mar 25, 2025
- Full-time
Job description
About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy. **As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development.** You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers. **The Role** As a Research Scientist in the Data team, your primary responsibility is to curate high quality data at the web-scale to fuel the development of next generation foundation models. You will work on exploring andconsolidatingdata sources and collaborate with cross-functional teams to conduct in-depth data research, contributing to MBZUAI’s mission of driving impactful AI discoveries and positioning the institution as a leader in the global AI research community. Your expertise will be key in enhancing the performance of large-scale machine learning models, while supporting the development of transformative AI tools that can influence industries worldwide. ### Key Responsibilities - Pioneer web-scale data collection and curation methodologies for LLMs and multi-modal foundation models. - Design and implement novel data synthesis pipelines for code, mathematics, and agentic reasoning datasets. - Trace the impact of data from pre-training to final model capabilities and create automated quality assessment frameworks for massive datasets - Design data recipes that maximize model capabilities across diverse domains. - Optimize data-model co-design for improved training dynamics. - Contribute to research papers and represent MBZUAI at industry conferences and events, showcasing the institution’s AI research and innovation. ### Academic Qualifications - Minimum: Master’s in Computer Science, Data Science, or a related technical field, or equivalent practical experience required. - Preferred: PhD or equivalent research experience in Machine Learning, NLP, or Data Science with a focus on LLMs and data is preferred. ### Professional Experience - Experience working with large language models, including evaluation, fine-tuning, and prompt engineering. - Strong Python development skills with a focus on research-grade code and scalable data pipelines. - Familiarity with collecting and processing large-scale datasets from open-source and web resources. - Demonstrated ability to work with ML infrastructure (e.g., model evaluation, optimization, debugging). - Proactive mindset with the ability to identify impactful research questions and execute on them with minimal supervision. - Effective communication and collaboration skills for working in cross-functional teams. Preferred - Prior research experience in areas such as web data curation and mixing, synthetizing complex datasets for training, LLM evaluation, post-training data, efficient inference, LLM-as-a-judge, tokenization. - Strong publication record in leading AI conferences (e.g., NeurIPS, ICLR, ICML, EMNLP) and/or prior contributions to open-source AI research or data tools. - Hands-on experience training language/mutli-modal models from scratch.