Type
Full-time
Work mode
On-site
Level
Staff
Industry
Other
Department
VinSmart Future
Salary
Thương lượng
Location
Thành phố Hồ Chí Minh, Hồ Chí Minh
Overview
- We are seeking a Senior AI Research Engineer to lead the development of state-of-the-art Vietnamese Speech AI technologies, including Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech-to-Speech Conversational AI.
- The ideal candidate has strong expertise in foundation model adaptation, pretraining, supervised fine-tuning (SFT), reinforcement learning, and knowledge distillation. You will be responsible for building SOTA Vietnamese speech models with high accuracy, naturalness, low latency
- Speech Foundation Models
- Research, develop, and optimize state-of-the-art Vietnamese ASR and TTS models.
- Adapt and improve large speech foundation models for Vietnamese language and accents.
- Work with open-source and commercial speech models, including:
- +Qwen3-ASR
- +Qwen3-TTS
- +Whisper
- +CosyVoice
- +Orpheus
- +Sesame
- +Fish Speech
- +XTTS
- +Other emerging speech foundation models
- Model Training & Fine-Tuning
- Design and implement scalable pipelines for:
- + Self-supervised pretraining
- + Continued pretraining
- + Supervised Fine-Tuning (SFT)
- + Instruction tuning
- + Domain adaptation
- Build and curate large-scale Vietnamese speech datasets.
- Develop data cleaning, alignment, and augmentation pipelines for speech training.
- Reinforcement Learning & Alignment
- Research and implement advanced optimization techniques:
- + Reinforcement Learning from Human Feedback (RLHF)
- + Direct Preference Optimization (DPO)
- + GRPO / PPO-based optimization
- + Preference learning for speech quality improvement
- Improve: ASR accuracy, TTS naturalness, Speaker similarity, Pronunciation quality, Dialogue experience
- Knowledge Distillation & Model Compression
- Distill large speech foundation models into efficient Vietnamese ASR/TTS models.
- Develop:
- + Teacher-student training frameworks
- + Representation distillation
- + Logit distillation
- + Feature matching approaches
- Optimize models using:
- + Quantization
Requirements
- Education
- Bachelor's, Master's, or PhD in: Computer Science, Artificial Intelligence, Machine Learning, Speech Processing, Computational Linguistics, Related fields
- Technical Skills
- + Strong understanding of: Deep Learning, Speech Processing, NLP, Generative AI, Transformer architectures, Experience training and fine-tuning large speech models.
- + Experience with: Self-supervised learning, Foundation models, Multimodal learning, Sequence-to-sequence architectures, Speech AI Expertise
- Hands-on experience in at least one of: Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Voice Conversion, Speech Translation, Speech-to-Speech systems
- Strong understanding of: Acoustic modeling, Language modeling, Vocoders, Speaker embeddings, Alignment methods, Reinforcement Learning & Distillation
- Practical experience with: RLHF, DPO, PPO / GRPO
Benefits
- Shape the future of intelligent AI applications across smart cities, super apps, and comprehensive intelligent systems for a top-tier company in the field.
- ● Work with a passionate team solving meaningful real-world problems.
- ● Competitive salary and benefits with clear growth paths:
- ○ Heath, social insurances
- ● Opportunities to work on state-of-the-art AI with real-world deployment across multiple high-impact verticals.
- ● Exposure to diverse technical challenges and the chance to build expertise across various AI application domains
Training
Summary of facts from the official posting. View original ↗
Interested in this role?
You'll be taken to the employer's official application page.
Apply on official site ↗
Is this your business?
Claim this page, request edits or removal
→