Research Engineer, Large-Scale Training
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.
As a Research Engineer on the Scaling Team within Model Shaping, you will turn cutting-edge research on efficient foundation model training into robust, high-performance systems. You will profile and optimize Together's training infrastructure, identify performance bottlenecks across the stack, and implement state-of-the-art techniques from both the research literature and our own scientists in production environments. Your work will directly shape the fine-tuning experience of Together's customers. You will rapidly bring newly released open-source models onto the Model Shaping platform, ensuring they train efficiently and reliably across diverse customer workloads. Working closely with Research Scientists, you will also build the experimental infrastructure that accelerates research and enables validated ideas to be deployed reliably at scale.
Responsibilities:
- Design, implement, and optimize core components of Together's large-scale training infrastructure.
- Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
- Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
- Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
- Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
- Rapidly enable support for newly released open-source foundation models on the Together platform.
- Build and maintain experimental infrastructure that accelerates research while ensuring production-quality reliability and scalability.
Requirements:
- Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
- Strong programming skills in Python and PyTorch, with an emphasis on writing efficient, maintainable code.
- Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
- Solid understanding of ML systems fundamentals, including GPU architecture and distributed systems.
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.