Machine Learning Engineer - Training Optimization
Ireland, England, United Kingdom · На постоянной основе
Подайте заявку первыми!
- Опыт
- Любой
- Зарплата
- —
- Открытия
- 1
- Опубликовано
- 3 часа назад
- Режим работы
- В офисе
- Резюме
- Необходимо подать заявку.
Где вы будете работать
Описание работы
Position Overview
Our partner company based in Ireland is seeking a Machine Learning Engineer specialized in Training Optimization. This role offers a significant opportunity to enhance the core methodologies behind large-scale AI model training and deployment.
Key Responsibilities
- Enhance large-scale model training pipelines to boost throughput, stability, convergence rates, and computational efficiency.
- Advance distributed training strategies including data, model, and pipeline parallelism.
- Fine-tune essential training elements such as optimizers, learning rate schedulers, batch size configurations, and numerical precision techniques (bf16, fp16, fp8).
- Detect and alleviate performance bottlenecks through profiling, system diagnostics, and infrastructure enhancements.
- Work collaboratively with research teams to create architecture-aware training methods to further model performance.
- Develop and support robust training infrastructure featuring checkpointing, fault tolerance, and reproducible workflows.
- Assess and integrate advanced training approaches like gradient checkpointing, ZeRO, FSDP, and bespoke optimization methods.
- Define and monitor training performance metrics to promote ongoing improvements in efficiency.
- Transform research insights into scalable, production-ready systems.
Candidate Requirements
- Demonstrated experience in training large-scale neural networks, including extensive language models or similarly complex systems.
- Practical expertise in machine learning training optimization beyond basic model usage.
- In-depth understanding of backpropagation, optimization algorithms, training behaviors, and convergence processes.
- Experience with distributed training frameworks and large-scale computational environments.
- Proficiency with PyTorch and contemporary machine learning development pipelines.
- Ability to work near hardware boundaries including GPU performance, memory constraints, and network considerations.
- Strong programming skills to implement research concepts into reliable, production-quality code.
- Preferred experience with multi-node and multi-GPU systems.
- Familiarity with platforms such as DeepSpeed, FSDP, Megatron, or equivalent custom stacks is advantageous.
- Experience optimizing workloads on NVIDIA or AMD GPU architectures beneficial.
- Contributions to open-source machine learning projects or infrastructure recognized positively.
- Knowledge of neural network architectures beyond Transformer models is a plus.
Benefits
- Competitive remuneration with equity participation opportunities.
- Engagement with cutting-edge AI models and large-scale training ecosystems.
- High degree of ownership influencing technical trajectory and company expansion.
- Collaboration within a compact, highly skilled engineering and research team focused on innovation and quality.
- Rapid feedback cycles promoting experimentation and impactful contributions.
- Challenging environment addressing complex machine learning infrastructure at scale.
- Commitment to technical excellence, continued learning, and advanced AI development.
- Chance to contribute foundational systems shaping the future landscape of AI.
Additional Information
This listing is managed by a trusted partner company who handles all recruitment processes including application review and interviews.
Privacy and Hiring Process: Candidate applications are processed under applicable data protection regulations. AI tools help facilitate unbiased application matching but final hiring decisions rest with human evaluators.