Data Scientist - Aeronautical and Machine Learning Specialist
Singapore · Full Time
Be the first to apply
- Experience
- 3–5 yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 hours ago
- Work mode
- In office
- Education
- Bachelors in Computer Science or Information Technology
- Resume
- Required to apply
Where you'll work
Job description
Overview
Thales is a global technology leader dedicated to advancing solutions in aerospace, cybersecurity, and digital identity, trusted by governments and enterprises to address complex challenges. Operating in Singapore since 1973 with 2,000 employees across three sites, Thales provides advanced technologies in air traffic management, defence, and security, helping customers make critical decisions that enhance community safety and foster progress.
Key Responsibilities
- Perform exploratory data analysis to uncover innovative insights and identify opportunities for optimizing air traffic operations.
- Devise novel AI strategies tailored to the aeronautical domain where data volumes are limited, going beyond conventional machine learning approaches.
- Develop, train, and implement machine learning models for tasks including real-time classification, regression, and sequence forecasting using frameworks like PyTorch, TensorFlow, and Scikit-learn, utilizing transformer-based models for spatial-temporal challenges.
- Create reinforcement learning agents using algorithms such as DQN, PPO, and Actor-Critic, applying them within simulated and real-world environments through OpenAI Gym or custom platforms.
- Construct and enhance Retrieval-Augmented Generation (RAG) pipelines based on specialized documentation to facilitate AI-driven reasoning and decision-making.
- Assess outputs from large language models for hallucinations and accuracy, devising domain-specific evaluation benchmarks for ATM-related reasoning.
- Automate machine learning workflows with orchestration tools such as Kubeflow and Airflow, ensuring reproducibility and operational efficiency.
- Integrate ML models into scalable APIs and deploy them in cloud-native environments employing Docker and Kubernetes technologies.
- Monitor and maintain ML model performance over time, addressing issues such as data drift through retraining and iterative improvements.
- Manage experiment tracking, version control, and reproducibility using platforms like MLflow or Weights & Biases.
- Work collaboratively with DevOps and backend engineers to seamlessly integrate ML components into broader systems.
Candidate Requirements
- Educational background with a Bachelor's degree in Computer Science or Information Technology; Master’s degree in Computer Science or Data Science is advantageous.
- Experience or familiarity with the aeronautical domain is highly valued, or alternatively, experience in sectors with limited historical data availability rather than common domains like vision or chatbots.
- Strong understanding of data analytics including statistics, feature engineering, and domain-mapping.
- Solid mathematical foundation especially in algorithms related to prediction and optimization.
- Minimum 3 to 5 years of end-to-end machine learning project delivery experience encompassing data preparation, model training, and deployment.
- Proficiency with Python and machine learning libraries such as PyTorch or TensorFlow.
- Knowledge of advanced deep learning architectures including transformers, attention mechanisms, and encoder-decoder models.
- Experience developing and fine-tuning RAG pipelines, prompt engineering, and agent frameworks like LangChain or LlamaIndex.
- Skills in evaluating large language model (LLM) outputs for hallucination and groundedness in source data, plus designing domain-specific benchmarks for LLM reasoning evaluation.
- Practical insights into reinforcement learning fundamentals and algorithms such as PPO and DQN, including experience with OpenAI Gym or Gymnasium platforms.
- Expertise in MLOps practices involving Kubeflow Pipelines, MLflow, Docker, Kubernetes, and cloud-based environments.
- Capability to build and troubleshoot data pipelines supporting both training and inference, with knowledge of ETL/ELT tools like Apache Spark and data lakes such as S3.
- Understanding of CI/CD processes for machine learning lifecycle and model deployment strategies.
Preferred Skills and Traits
- Familiarity with additional programming languages such as Scala (2 or 3), Go, TypeScript, C, C++17, or Java17.
- Experience with AI/MLOps pipelines implemented on public cloud platforms including Azure, AWS, or GCP.
- Attributes such as adaptability, eagerness to learn, flexibility, and proactive problem-solving.
- Comfort in agile team settings with active user engagement.
Company Culture
At Thales, an inclusive and respectful work environment is paramount, fostering collaboration and passion. Employees are encouraged to bring their genuine selves, grow professionally, and contribute to technological innovations that promote safer, greener, and more inclusive futures.