Bachelor's degree or above in Computer Science, Artificial Intelligence, NLP, Machine Learning, Distributed Systems, or a related field.
LLM Pre-training Experience
Hands-on experience with complete LLM pre-training projects.
Experience participating in the pre-training of 7B+ parameter models.
Distributed Training
Strong expertise in large-scale distributed training frameworks such as Megatron-LM, DeepSpeed, or FSDP.
Practical experience with 64+ GPU training environments.
Pre-training Data Engineering
Solid understanding of large-scale pre-training data pipelines, including data cleaning, deduplication, quality filtering, tokenization, data mixing, and data quality optimization.
Training Monitoring & Debugging
Strong ability to analyze training loss, gradients, convergence, and training stability.
Experience troubleshooting large-scale distributed training issues.
Long-context Training
Familiarity with long-context training and extension techniques, including RoPE scaling, NTK-aware interpolation, and YaRN.
Preferred Qualifications
Experience with 70B+ parameter model pre-training.
Experience with MoE (Mixture-of-Experts) model pre-training.
Publications in top-tier AI/ML conferences such as NeurIPS, ICML, ICLR, ACL, or EMNLP, particularly in LLM pre-training, model architecture, or training optimization.
Experience optimizing large-scale GPU clusters and training infrastructure.