T

Incident Manager

ThoughtFocus

Kochi, Kerala, India · Full Time

Be the first to apply

Experience
8–9 yrs
Salary
Openings
1
Posted
2 ਘੰਟੇ ਪਹਿਲਾਂ
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

We are seeking an experienced Incident Manager with strong background in ITIL and ITSM operations. The ideal candidate will possess over eight years of experience in IT service management, focusing on incident, major incident, problem, and change management processes. Expertise in ServiceNow ITSM, workflow automation, and multiple monitoring and alerting platforms will be essential.

Key Responsibilities and Skills

  • Manage ITSM operations including incident, major incident, problem, and change management.
  • Utilize ServiceNow for ITSM processes and automate workflows effectively.
  • Leverage Application Performance Management (APM) tools such as AppDynamics, SigNoz, or similar.
  • Work with monitoring tools including Grafana, Prometheus, Splunk, Application Insights, and System Center Operations Manager (SCOM).
  • Coordinate incident escalations and event management using PagerDuty and related escalation frameworks.
  • Drive monitoring tool consolidation, optimize alerting, and manage thresholds to enhance system observability.
  • Support AIOps initiatives and proactive operational strategies.
  • Collaborate across stakeholders, utilizing strong analytical and communication skills.

Industry Experience

Candidate should ideally have experience supporting banking, financial services, digital platforms, or large-scale enterprise environments.

Preferred Technical Experience

  • Monitoring infrastructure and systems with SCOM, including over 600 management packs.
  • Using Application Insights to oversee Digital Banking, NuPoint, Identity, and Tokenization platforms.
  • Employing Splunk for operational analytics in Digital Banking, ACH, FedNow, NuPoint Wire, and Imaging systems.
  • Managing PagerDuty escalations for 35+ teams, improving Mean Time To Acknowledge (MTTA) and Mean Time To Resolve (MTTR).
  • Utilizing OpManager and NetFlow to monitor network devices, traffic, and service availability.

Tools & software

Prometheus required

How they work

Communication Problem Solving

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help