About this role
About the Role
Work alongside an AI research team to turn experimental code into production-ready systems. You will build and improve shared post-training infrastructure that supports research data generation, processing, and evaluation.
What You'll Do
-
Turn research prototypes into tested, reusable data generation, training, and evaluation pipelines.
-
Build distributed experiment support, including data loading, checkpointing, and collection of agent interactions.
-
Profile workloads and improve GPU utilization, memory efficiency, and data throughput.
-
Develop tests, experiment tracking, and debugging tools while preserving the integrity of research results.
-
Collaborate with researchers to translate ideas into maintainable software.
What We're Looking For
-
Experience building ML training, inference, or data pipelines in Python.
-
Strong Python skills and hands-on experience with PyTorch, JAX, or comparable machine learning frameworks.
-
Understanding of machine learning experiments, including how data, numerical precision, and implementation choices affect results.
-
Sound software engineering practices, including testing, profiling, version control, and documentation.
-
Distributed training or data processing experience with tools such as PyTorch Distributed, DeepSpeed, or Ray is valuable.
-
Familiarity with Hugging Face Transformers, vLLM, or experiment tracking platforms such as Weights & Biases or MLflow is useful.
-
Experience with reinforcement learning systems or multimodal datasets is a plus.
Compensation & Benefits
Competitive compensation and equity are offered. Visa sponsorship is available.
Location
On-site in Zürich, Switzerland.