A

Observability Engineer - Grafana Prometheus & VictoriaMetrics

Apptad

Navi Mumbai, Maharashtra, India · Full Time

Be the first to apply

Experience
2+ yrs
Salary
Openings
1
Posted
6 days ago
Work mode
In office
Education
Any graduate
Eligibility
Any graduate
Resume
Required to apply

Where you'll work

Job description

Job Overview

We seek an Observability Engineer to spearhead the development and upkeep of scalable monitoring frameworks leveraging Grafana, Prometheus, and VictoriaMetrics. This role is critical for improving observability, system performance, and stability across distributed systems and applications.

Grafana Responsibilities and Skills

  • Create and manage user-friendly dashboards, visual elements, and panels for live system monitoring.
  • Develop and maintain alerting rules, notification policies, and contact mechanisms for timely incident awareness.
  • Design dashboards using various visualization types including statistics, tables, pie charts, bar charts, time series, and heatmaps.
  • Apply advanced panel transformations such as filtering, joining, merging, grouping, and various calculations.
  • Incorporate dashboard variables, dynamic filtering options, and linking strategies to enhance usability.
  • Utilize JSON API and Infinity plugin with Universal Query Language to handle nested API data.
  • Integrate multiple data sources including Elasticsearch, CloudWatch, Druid, and custom JSON APIs effectively.
  • Maintain configuration as code for dashboards and alerts using version control systems like Bitbucket or GitHub.
  • Deploy monitoring changes through GoCD pipelines adhering to GitOps workflows.
  • Ensure dashboards and alerts operate correctly post-deployment through validation checks.
  • Administer role-based access control within Grafana and link security groups appropriately.
  • Implement alert silencing during scheduled maintenance and recognized downtimes.
  • Set up alert integrations with communication tools such as Email, Teams, Slack, PagerDuty, ServiceNow, and Webhooks.
  • Diagnose and resolve issues through dashboards, metrics, and log analysis using Kibana.
  • Work proficiently with PromQL for querying metrics and Lucene for log searches.

Prometheus & VictoriaMetrics Duties and Skills

  • Architect and administer monitoring solutions based on Prometheus and VictoriaMetrics platforms.
  • Handle single-node and clustered VictoriaMetrics deployments ensuring scalability and reliability.
  • Automate service and node onboarding using service discovery, Ansible, and Terraform automation tools.
  • Create and maintain Prometheus scrape jobs, relabeling, recording rules, and federation setups.
  • Configure remote_write endpoints for Prometheus and ingestion pipelines for VictoriaMetrics.
  • Manage metric data retention, scraping intervals, and optimize query performance.
  • Address challenges posed by high-cardinality metrics to optimize storage and query speed.
  • Configure core VictoriaMetrics components such as vmagent, vmalert, vmselect, and vmstorage.
  • Onboard and maintain diverse exporters including Node Exporter for Linux, Windows Exporter, SNMP Exporter for network devices, Blackbox Exporter, cAdvisor, kube-state-metrics, and JMX/WebLogic Exporters.
  • Monitor infrastructure spanning Linux, Windows, routers, switches, SD-WAN, storage arrays, JVM-based apps, Kubernetes, virtualization layers, and network attached storage platforms.
  • Resolve issues like scrape failures, exporter malfunctions, missing data points, and ingestion bottlenecks effectively.
  • Ensure secure access, authentication, and authorization mechanisms for Prometheus and VictoriaMetrics environments.

Preferred Qualifications

  • Minimum of two years of experience in observability, Site Reliability Engineering, or DevOps roles.
  • Proven expertise with Grafana, Prometheus, and VictoriaMetrics monitoring solutions.
  • Hands-on experience configuring advanced exporters such as Node, Windows, SNMP, and JMX/WebLogic.
  • Strong command of PromQL language and contemporary monitoring methodologies.
  • Experience managing configurations via Git repositories.
  • Familiarity with continuous integration/deployment tools like GoCD.
  • Knowledge of containerization and orchestration technologies including Docker and Kubernetes.
  • Competence in automation tools like Ansible or Terraform to streamline infrastructure management.
  • Collaborative skills to work effectively with development, operations, and support teams.

Eligibility

Applicants must have completed any graduate degree to be eligible for this position.

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help