- Experience
- 2+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 3 മണിക്കൂർ മുൻപ്
- Work mode
- Work from home
- Education
- PhD or equivalent experience
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
We are seeking a skilled Audio AI Engineer to devise and implement algorithms focusing on accent and voice conversion, speech synthesis, and recognition, specifically optimized for low-latency streaming environments. This role involves prototyping comprehensive audio models that improve clarity and natural sound while preserving speaker individuality. Collaboration with cross-functional product and platform teams is key to deploying these models in live communication systems. The engineer will also assess and refine performance metrics including latency, scalability, and speech quality. A commitment to ongoing research and contribution to intellectual property and internal knowledge is expected.
Team and Culture
The audio team operates worldwide including bases in the U.S., China, and Singapore, working on real-time audio solutions driven by AI technology.
Key Responsibilities
- Investigate, design, and develop algorithms tailored for accent conversion, voice transformation, speech synthesis and automatic speech recognition under streaming low-latency constraints.
- Build and fine-tune end-to-end audio models that enhance speech intelligibility and naturalness while maintaining speaker expression and identity.
- Work closely with product and platform teams to embed AI models into real-time audio and video communication frameworks.
- Analyze and optimize model outputs focusing on quality, latency, robustness, and scalability.
- Keep abreast with the latest research in speech processing and actively contribute through patents and knowledge sharing sessions.
Candidate Profile
- PhD or equivalent experience in a related area such as Streaming, Accent or Voice Conversion, TTS, or ASR; candidates with over two years of related industry experience are especially welcome.
- Proficiency with deep learning frameworks such as PyTorch or TensorFlow.
- Strong programming skills in Python, C/C++, or comparable languages.
- Knowledge of sequence modeling architectures like Transformers, RNNs, diffusion models, or conformers.
- Experience developing and deploying low-latency, real-time speech/audio models with streaming architecture and optimized pipelines.
- Understanding of model optimization methods including quantization, pruning, and distillation.
- Hands-on experience with real-time audio systems in networked communication contexts.
- Published research in leading conferences such as Icassp, Interspeech, NeurIPS, or ICLR.
Working Style
This role supports a hybrid work arrangement centered on office presence and remote flexibility.
Benefits
We offer a comprehensive benefits package designed to promote physical, emotional, mental, and financial well-being, alongside support for work-life balance and community involvement.
About the Company
We provide platforms that facilitate seamless communication and collaboration, including video conferencing and various communication tools. We foster a fast-paced, problem-solving environment with opportunities for professional growth and skill advancement.
Equal Opportunity Commitment
Our hiring practices promote fairness and support accommodations for candidates with disabilities throughout the recruitment process. Privacy and support tools are integrated to ensure a consistent and respectful interview experience.
Minimum education
Doctorate