Skip to content
Tuyển Dụng
← Vingroup

Chuyên gia Nghiên cứu AI

Vingroup · Hồ Chí Minh
Apply on official site ↗
Type
Full-time
Work mode
On-site
Level
Staff
Industry
Other
Department
VinSmart Future
Salary
Thương lượng
Location
Thành phố Hồ Chí Minh, Hồ Chí Minh

Overview

  • We are seeking a Senior AI Research Engineer to lead the development of state-of-the-art Vietnamese Speech AI technologies, including Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech-to-Speech Conversational AI.
  • The ideal candidate has strong expertise in foundation model adaptation, pretraining, supervised fine-tuning (SFT), reinforcement learning, and knowledge distillation. You will be responsible for building SOTA Vietnamese speech models with high accuracy, naturalness, low latency
  • Speech Foundation Models
  • Research, develop, and optimize state-of-the-art Vietnamese ASR and TTS models.
  • Adapt and improve large speech foundation models for Vietnamese language and accents.
  • Work with open-source and commercial speech models, including:
  • +Qwen3-ASR
  • +Qwen3-TTS
  • +Whisper
  • +CosyVoice
  • +Orpheus
  • +Sesame
  • +Fish Speech
  • +XTTS
  • +Other emerging speech foundation models
  • Model Training & Fine-Tuning
  • Design and implement scalable pipelines for:
  • + Self-supervised pretraining
  • + Continued pretraining
  • + Supervised Fine-Tuning (SFT)
  • + Instruction tuning
  • + Domain adaptation
  • Build and curate large-scale Vietnamese speech datasets.
  • Develop data cleaning, alignment, and augmentation pipelines for speech training.
  • Reinforcement Learning & Alignment
  • Research and implement advanced optimization techniques:
  • + Reinforcement Learning from Human Feedback (RLHF)
  • + Direct Preference Optimization (DPO)
  • + GRPO / PPO-based optimization
  • + Preference learning for speech quality improvement
  • Improve: ASR accuracy, TTS naturalness, Speaker similarity, Pronunciation quality, Dialogue experience
  • Knowledge Distillation & Model Compression
  • Distill large speech foundation models into efficient Vietnamese ASR/TTS models.
  • Develop:
  • + Teacher-student training frameworks
  • + Representation distillation
  • + Logit distillation
  • + Feature matching approaches
  • Optimize models using:
  • + Quantization

Requirements

  • Education
  • Bachelor's, Master's, or PhD in: Computer Science, Artificial Intelligence, Machine Learning, Speech Processing, Computational Linguistics, Related fields
  • Technical Skills
  • + Strong understanding of: Deep Learning, Speech Processing, NLP, Generative AI, Transformer architectures, Experience training and fine-tuning large speech models.
  • + Experience with: Self-supervised learning, Foundation models, Multimodal learning, Sequence-to-sequence architectures, Speech AI Expertise
  • Hands-on experience in at least one of: Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Voice Conversion, Speech Translation, Speech-to-Speech systems
  • Strong understanding of: Acoustic modeling, Language modeling, Vocoders, Speaker embeddings, Alignment methods, Reinforcement Learning & Distillation
  • Practical experience with: RLHF, DPO, PPO / GRPO

Benefits

  • Shape the future of intelligent AI applications across smart cities, super apps, and comprehensive intelligent systems for a top-tier company in the field.
  • ● Work with a passionate team solving meaningful real-world problems.
  • ● Competitive salary and benefits with clear growth paths:
  • ○ Heath, social insurances
  • ● Opportunities to work on state-of-the-art AI with real-world deployment across multiple high-impact verticals.
  • ● Exposure to diverse technical challenges and the chance to build expertise across various AI application domains
Training

Summary of facts from the official posting. View original ↗

Interested in this role?

You'll be taken to the employer's official application page.

Apply on official site ↗
Is this your business? Claim this page, request edits or removal