P

DevOps Support Engineer

Prodapt Solutions

Hyderabad, Telangana, India · Full Time

Be the first to apply

Experience
3–8 yrs
Salary
Openings
1
Posted
1 saat önce
Work mode
In office
Education
Any graduate
Eligibility
Applicants must hold at least a graduate degree in any discipline.
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

Prodapt Solutions Private Limited is seeking a self-motivated and technically proficient DevOps Support Engineer to handle Tier 1 & 2 application production support. Based in Hyderabad or Bangalore, this role involves providing 24/7 support for critical enterprise integration, messaging, and event-driven applications within an Azure Kubernetes ecosystem. The position requires independent incident diagnosis and resolution with minimal escalation.

Key Responsibilities

  • Deliver 24x7 incident, alert, and operational support focused on autonomous problem resolution.
  • Analyze and triage incidents using logs, monitoring tools, and platform knowledge to resolve issues without routine escalation.
  • Engage Tier 2 support exclusively for complex architectural or infrastructure problems beyond Tier 1 scope.
  • Manage client-submitted incident tickets end-to-end, including resolution and closure.
  • Respond promptly to PagerDuty and automated alerts by validating and rectifying issues before escalating.
  • Continuously monitor production and non-production environments to proactively detect anomalies and apply corrective measures.
  • Manage application support communications and operational follow-ups.
  • Provide first-responder emergency support and independently resolve incidents when possible.
  • Share client profile data and generate usage reports as requested.
  • Maintain clear communications with clients throughout incident management, including when escalation occurs.
  • Track and report operational metrics such as Mean Time To Repair (MTTR) and ticket resolution trends.
  • Create and update Tier 1 Standard Operating Procedures (SOPs) and operational runbooks based on real-world resolution experience.

Requirements

  • Minimum 3 years of experience in application production support with proven ability to independently troubleshoot and resolve incidents.
  • Comprehensive knowledge of event streaming platforms, integration middleware, AKS-hosted microservices, and observability tools.
  • Experience with incident management ticketing tools like iTrack or equivalents, including root cause analysis and documentation.
  • Proficiency in using Splunk, PagerDuty, Prometheus, and Grafana for active investigation beyond simple monitoring.
  • Hands-on operational expertise with Kubernetes, particularly Azure Kubernetes Service (AKS), including pod diagnostics and health monitoring.
  • Solid working knowledge of Confluent Kafka and Azure Event Hub, including consumer lag analysis and message flow troubleshooting.
  • Strong SQL/PostgreSQL skills for incident investigation and validation.
  • Ability to interpret Java, Spring Boot, and React logs for root cause identification.
  • Fundamental Python scripting skills to create quick operational fixes and checks.
  • Proficient Linux/Unix command line skills for real-time log analysis and system diagnostics.
  • Excellent verbal and written communication skills to provide incident updates and maintain client coordination.
  • Willingness to work in rotational 24/7 shifts on-site at Hyderabad or Bangalore.

Preferred Skills

  • Familiarity with hybrid streaming environments such as Confluent Cloud, AWS MSK, and Apache Flink.
  • Exposure to IBM Sterling Integrator for better incident context.
  • Experience supporting telecom or high-availability enterprise environments.
  • Knowledge of CI/CD deployment pipelines within a support framework.

Additional Information

The candidate must share current and expected CTC, total and relevant experience, notice period, current location, readiness to work in 24/7 shifts, and reason for job change when applying. Immediate to 30-day joining candidates are preferred. This is an on-site role requiring availability for rotational shifts.

Minimum education

Bachelor's Degree

Tools & software

PostgreSQL required Prometheus required

How they work

Communication Problem Solving Adaptability Initiative Independence
🤖
Online · instant AI help