About this role
Job title: Senior Data Scientist I
About the Role As a Senior Data Scientist, you will lead the development and deployment of GenAI, NLP and RAG solutions, building production-ready Python code and collaborating with software engineers and domain experts. You will work across the full data science lifecycle from design to production, and help validate outputs with biology and chemistry SMEs.
What You'll Do
- Data collection, data analysis, model development, defining quality metrics, quality assessment of models and regular presentations to stakeholders.
- Creating production-ready Python packages for each component of data science pipelines (pre-processing, model inference) and their deployment with software engineering teams.
- Optimizing and customizing Retrieval Augmented Generation (RAG) pipelines for content ingestion, translation, and contextualized information retrieval.
- Ingesting, preprocessing, and transforming large-scale multilingual data to ensure high-quality inputs for downstream models.
- Building AI agentic models integrated with RAG pipelines.
- Conducting rigorous testing and evaluation of AI models to ensure high performance and reliability.
- Integrating data science components and performing end-to-end quality assessments.
- Maintaining robustness of data science pipelines against model drift and ensuring consistent output quality.
- Establishing reporting processes for pipeline performance and developing automated re-training strategies for existing pipelines.
- Collaborating with cross-functional teams to integrate AI solutions into existing products and services.
- Leading and managing projects with a team of data scientists and independently executing the entire small-scale projects.
- Mentoring junior data scientists and fostering a knowledge-sharing culture within the team.
- Staying up-to-date with the latest advancements in AI, machine learning, and NLP technologies.
What We're Looking For
- Master’s or Ph.D. in Computer Science, Data Science, Artificial Intelligence, or a related field.
- 5+ years of relevant applied experience in data science, with a focus on Generative AI, NLP, and machine learning.
- Proficiency in Python for data analysis, model development, and deployment.
- Strong experience with transformer models.
- Proficiency in Generative AI technologies, including utilizing LLMs via API access, LLM evaluation tools, and prompt engineering.
- Knowledge of various Retrieval Augmented Generation (RAG) pipelines and their practical implementation.
- Experience building Agentic RAG systems (strong requirement).
- Experience with AI agent management frameworks such as LangChain, or similar tools.
- Experience with advanced algorithms in deep learning, neural networks, reinforcement learning, and transfer learning.
- Familiarity with traditional machine learning algorithms (random forests, SVM, logistic regression, and Bayesian modelling) for model building, validation, and testing.
- Familiarity with cloud platforms (Bedrock, AWS, Azure) for deployment and creating production-ready pipelines.
- Proficiency in data visualization tools and techniques.
- Experience with version control systems (GitLab or GitHub), Jira, and working in an Agile environment.
- Proficient in using OpenSearch and Databricks.
- Excellent problem-solving and analytical skills, with strong attention to detail.
- Strong communication skills and the ability to work effectively in a team-oriented environment.
Nice to Have
Compensation & Benefits
- Base Pay Range (Amsterdam, NL): €53,800 - €89,900 per year.
- Flexible working hours; wellbeing initiatives; shared parental leave; study assistance; sabbaticals.
- Additional benefits and a focus on work-life balance as part of Elsevier's offerings.