About this role
Senior AI Engineer - LLM & Multi-Modal
About the job
As a member of the AI model team, you will drive innovation in architecture development for cutting-edge models of various scales, including small, large, and multi-modal systems. Your work will enhance intelligence, improve efficiency, and introduce new capabilities to advance the field.
Responsibilities
-
Large-Scale Pre-Training: Conduct foundational pre-training for LLMs and Multi-Modal models (integrating text, vision, audio, or other modalities) on large, distributed servers equipped with multi-nodes & thousands of NVIDIA GPUs.
-
Architecture & Alignment Innovation: Design, prototype, and scale innovative architectures, tokenizers, and cross-modal alignment layers to enhance model intelligence and multi-modal understanding.
-
Data Strategy: Source, filter, and curate massive-scale textual and multi-modal datasets, establishing robust data pipelines for efficient pre-training.
-
Experimental Research: Independently and collaboratively execute experiments, analyze results, and refine training methodologies for optimal performance and token efficiency.
-
Optimization & Debugging: Investigate, debug, and eliminate bottlenecks in model efficiency, computational performance, and multi-modal alignment stability during long training runs.
-
System Scalability: Contribute to the advancement of distributed training systems to ensure seamless scalability and hardware efficiency on target platforms.
-
A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
-
Hands-on experience contributing to large-scale LLM or Multi-Modal pre-training runs on large, distributed servers equipped with thousands of NVIDIA GPUs, ensuring scalability and impactful advancements in model performance.
-
Familiarity and practical experience with large-scale, distributed training frameworks, libraries and tools.
-
Deep knowledge of state-of-the-art transformer and non-transformer modifications aimed at enhancing intelligence, efficiency and scalability.
-
Strong expertise in PyTorch and Hugging Face libraries with practical experience in model development, continual pretraining, and deployment.
-
Important information for candidates
-
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
-
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
-
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
-
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
-
Double-check email addresses. All communication from us will come from emails ending in @tether.to or @tether.io
-
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
-
When in doubt, feel free to reach out through our official website.
-
Location: ABC United Kindom GB
-
Compensation: $90k - $125k estimated