
Search by job, company or skills
Job Responsibilities
Responsible for post-training work related to the speech understanding capabilities of Omni multimodal large models. Based on existing foundation models, design and conduct post-training experiments such as SFT, DPO, and RLHF to improve speech recognition, speech understanding, long-form speech reasoning, and speech instruction-following capabilities.
Participate in the design, cleaning, construction, and quality evaluation of speech-text multimodal training data. Use Agents/tools where applicable to generate high-quality speech multimodal samples, identify bad cases, and drive dataset iteration and optimisation.
Carry out algorithm iteration for key pain points in Omni model speech understanding, including noisy speech, accents, long-form speech, mixed speech, speech Q&A, and speech instruction understanding. Analyse model weaknesses, design solutions, and validate implementation results.
Track frontier approaches in Omni models and speech large models, reproduce methods, conduct comparative experiments, develop practical algorithm solutions, prepare experiment reports, and align outcomes with business metrics.
Work with inference, data, and engineering teams to complete model performance evaluation, build metric systems, and support online validation of optimised models.
Job Requirements
Bachelor's degree or above in Computer Science, Speech Signal Processing, Artificial Intelligence, or a related field.
1-2 years of working experience in large models, speech algorithms, or related areas, with practical industry project implementation experience.
Familiar with the full large model post-training workflow, including SFT, DPO, preference alignment, training data processing, prompt engineering, and model evaluation.
Hands-on project experience in speech-side post-training for Omni or multimodal large models. Experience in speech instruction-following, long-form speech understanding, or speech dialogue alignment is preferred.
Familiar with the basic principles of speech large models and Omni multimodal models, with knowledge of speech-text alignment and speech understanding technologies.
Proficient in PyTorch and able to independently run post-training experiments.
Strong bad case analysis skills, with the ability to identify issues from the perspectives of data, model, and training strategy, and drive model performance improvement.
Job ID: 152213053