This page was automatically translated and may contain errors. View in English.
M

AWS Data Engineer

Merck Sharp & Dohme (MSD)

Hyderabad, Telangana, India · Tempo total

Seja o primeiro a se candidatar

Experiência
8+ anos
Salário
Vagas
1
Publicado
há 3 horas
Modo de trabalho
No escritório
Educação
Qualquer graduado
Elegibilidade
Qualquer graduado
Retomar
Obrigatório candidatar-se

Onde você trabalhará

Descrição da vaga

About Merck Sharp & Dohme (MSD)

Merck Sharp & Dohme, known as Merck & Co., Inc. in the US and Canada, is a global biopharmaceutical company with over a century of history in delivering medicines and vaccines to combat critical diseases worldwide. MSD is committed to innovation through research to develop cutting-edge healthcare solutions for humans and animals alike. The company values inventive, collaborative individuals passionate about impactful healthcare advancements and career growth.

The Opportunity

  • Work from Hyderabad, India as part of a respected global healthcare firm with a 130-year legacy of ethical integrity and innovation.
  • Engage in digital transformations that support a diverse range of prescription medicines, vaccines, and animal healthcare products.
  • Contribute to a dynamic team leveraging data, analytics, and custom software to address some of the world's major health challenges.

Role Overview

We are seeking a Senior AWS Data Engineer responsible for creating scalable, reliable, and cost-effective data platforms on AWS and Databricks. This role involves developing batch and real-time data pipelines (ETL/ELT), establishing strong data quality controls, and building dimensional data models to empower analytics and reporting. You will also mentor teammates, set engineering standards, and collaborate closely with stakeholders to translate business needs into well-managed data assets. We continuously adopt modern data ecosystem patterns such as lakehouse, data mesh, and data fabric.

Key Responsibilities

  • Build and manage end-to-end ETL/ELT pipelines to ingest data into AWS data lakehouse and data warehouse infrastructure, processing both batch and streaming data.
  • Develop curated datasets applying solid data modeling techniques like star/snowflake schemas, slowly changing dimensions (SCD), and conformed dimensions to enhance BI and self-service analytics.
  • Collaborate with product managers, analysts, and data scientists to gather requirements, define mappings, and deliver accurate, accessible, and reusable data assets.
  • Define and enforce data validation rules, anomaly detection, data contracts, and SLAs; work with governance to maintain metadata catalogs and data lineage.
  • Implement orchestration, logging, monitoring, and alerting mechanisms to ensure pipeline stability and swift incident response.
  • Adopt engineering best practices including automated testing, code reviews, and CI/CD pipelines to enable safe deployments.
  • Write transformation code on Databricks using Python, PySpark, and Spark SQL, optimizing for performance and cost efficiency.
  • Create and optimize complex SQL queries for data transformations and warehouse consumption.
  • Utilize AWS services such as S3, IAM, Glue, Lambda, Step Functions, EMR/ECS/Fargate, and CloudWatch following security best practices.
  • Manage infrastructure using Terraform for provisioning and automated deployment with reusable modules.
  • Use Docker containers where appropriate to ensure consistent environments for development and deployment.
  • Employ GitHub for version control, follow branching strategies like trunk-based or GitFlow, and maintain quality pull requests.
  • Process large datasets via PySpark and formats like Delta and Parquet, ensuring reliability and scalability.
  • Create prototypes using notebooks and transition them into production-quality solutions with thorough testing and documentation.
  • Work within Agile methodologies (Scrum/Kanban), actively participating in planning, demos, and retrospectives.
  • Document technical details including data flows, runbooks, and dictionaries; contribute to team standards and operational playbooks.
  • Mentor junior data engineers through onboarding, collaborative coding, reviews, and knowledge sharing to elevate team capabilities.
  • Provide leadership by defining engineering standards, influencing technology roadmaps, and clearly communicating decisions and trade-offs.

Required Qualifications

  • Extensive experience (8+ years) in data engineering with a focus on building production-grade data platforms and pipelines.
  • Strong expertise in AWS data services including S3, Glue, Lambda, Step Functions, EMR/ECS/Fargate, and CloudWatch monitoring.
  • Proficiency in Python, PySpark, and Spark SQL with solid data engineering fundamentals.
  • Deep knowledge of data warehousing and lakehouse architectures, including dimensional modeling techniques and SCD patterns.
  • Hands-on experience with Databricks Lakehouse on AWS (Delta Lake) and familiarity with cloud data warehouses such as Redshift.
  • Demonstrated collaboration in Agile teams with clear communication of requirements and trade-offs.
  • Strong focus on data quality and operational excellence with experience in testing, monitoring, alerting, and root cause analysis.
  • Advanced SQL skills including complex joins, window functions, and query optimization.
  • Familiarity with GitHub workflows, CI/CD pipelines, Docker, and Infrastructure as Code tools like Terraform.
  • Experience mentoring engineers, providing guidance, and delivering high-quality outputs while remaining hands-on.
  • Bachelor’s degree or equivalent experience in Computer Science, Engineering, or related fields.

Preferred Qualifications

  • AWS certifications such as Developer, Solutions Architect, or Data Analytics, and/or Databricks certification.
  • Experience with data transformation frameworks like dbt and orchestration tools such as Airflow.
  • Knowledge of data product development including domain-aligned datasets, SLAs, documentation, and usage metrics.
  • Familiarity with data governance and access control mechanisms.

Additional Information

Note: This role requires exclusive experience with AWS and does not consider equivalent knowledge of Azure or GCP, even if fundamental concepts overlap.

Deixe este campo se desejar uma resposta — não o utilizaremos para mais nada.

Clique para navegar, arrastar e soltar, ou colar uma captura de tela

PNG, JPG, GIF, MP4, WebM, MOV · Máximo de 20 MB cada · Até 5 arquivos

🤖
Online · ajuda instantânea de IA