T

AWS Big Data Engineer (PySpark & EMR)

Tata Consultancy Services

Indore, Madhya Pradesh, India · Full Time

Be the first to apply

Experience
4–7 yrs
Salary
Openings
1
Posted
vor 3 Stunden
Work mode
In office
Education
Any graduate
Eligibility
Any graduate can apply for this role.
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Position Overview

We are seeking an experienced AWS Big Data Engineer proficient in PySpark and AWS EMR to join our team in Indore. The role involves designing, developing, and maintaining robust, enterprise-scale data processing pipelines leveraging AWS cloud technologies and big data frameworks.

Key Responsibilities

  • Design and build scalable data processing workflows using PySpark and various AWS services including EMR, S3, IAM, Lambda, SNS, and SQS.
  • Develop and optimize Spark-based ETL solutions for enterprise data platforms.
  • Implement complex data transformation logic efficiently with Python and PySpark.
  • Analyze and enhance Spark job performance using UI tools, debugging techniques, and tuning strategies.
  • Compose and streamline SQL queries incorporating joins, subqueries, CTEs, and other advanced features.
  • Support integration, ingestion, and reporting of data across diverse systems.
  • Contribute to solution architecture incorporating Data Vault, Data Mesh, and Data Fabric concepts.
  • Collaborate with business users, architects, and technical teams to comprehend requirements and deliver scalable solutions.
  • Engage actively in Agile ceremonies, aiding in project planning, estimations, and execution.
  • Troubleshoot production-level issues and implement performance improvements on data platforms.

Required Skills & Experience

  • 4 to 7 years of experience in data engineering or big data development roles.
  • Proficient with PySpark and Apache Spark frameworks.
  • Hands-on experience with AWS EMR and other AWS cloud services.
  • Strong proficiency in Python programming and scripting.
  • Familiarity with AWS offerings like S3, IAM, Lambda, SNS, and SQS.
  • Sound understanding of big data architecture and distributed computing principles.
  • Expertise in writing complex SQL queries including joins, subqueries, and CTEs.
  • Experience with multiple database technologies and data processing frameworks.
  • Knowledge of Spark optimization, debugging, and performance tuning techniques.

Preferred Skills

  • Familiarity with Data Vault architecture methodology.
  • Knowledge of emerging data architectural paradigms such as Data Mesh and Data Fabric.
  • Experience handling large-scale cloud data migration initiatives.
  • Exposure to Hadoop ecosystem and related big data technologies.
  • Prior experience working in Agile/Scrum project environments.

Candidate Profile

  • Strong analytical and problem-solving capabilities.
  • Excellent communication skills and ability to manage stakeholder interactions.
  • Proven ability to work closely with both business and IT teams.
  • Competent in estimating effort, scheduling work, and delivering projects successfully.
  • Self-driven, proactive, and adaptable to dynamic, fast-paced settings.

Eligibility

Any graduate can apply for this position.

Minimum education

Bachelor's Degree

Tools & software

PySpark required Apache Spark required

How they work

Communication Problem Solving Adaptability Initiative

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help