X

Operations Engineer

Xsolla

Baku, Azerbaijan · Full Time

Be the first to apply

Experience
4+ yrs
Salary
Openings
1
Posted
1 час назад
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

We seek an Operations Engineer eager to join our Global Technical Operations team in Baku. This position suits someone naturally inquisitive, detail-focused, and communicative, thriving in a fast-moving, collaborative environment. The incumbent will proactively monitor and investigate production issues on a worldwide platform, enhance incident detection and response, analyze system trend data, and facilitate effective communication across teams and stakeholders during incident events.

Key Duties

  • Continuously monitor operational dashboards (Datadog) during shifts to identify anomalies by correlating application performance data, logs, metrics, synthetic tests, and real-user monitoring.
  • Assess whether alerts need incident tickets or can be resolved through quick investigation.
  • Manage production incidents by creating tickets in JIRA Service Management, conducting initial technical analysis, identifying impact scope and root cause, and routing to relevant teams.
  • Handle lower-severity incidents from detection through resolution independently using runbook procedures; escalate unresolved or complex incidents appropriately.
  • Assist during major incidents as the technical lead in war rooms by maintaining live incident documentation, surfacing real-time data, and executing mitigation steps.
  • Compose clear incident communication updates for internal teams, stakeholders, and customers on status pages.
  • Analyze post-incident data trends, identify recurring production issues, and collaborate on reports to improve product and engineering practices.
  • Prepare post-incident review documents; track and escalate outstanding action items.
  • Develop automation tools and maintain runbooks to streamline operations and incident handling.
  • Conduct thorough shift handoffs including updates on active incidents, at-risk services, deployments, and follow-ups; participate in knowledge-sharing with Site Reliability Engineers.
  • Provide backup for team leadership duties including incident escalation, communication, and classification as needed.
  • Regularly publish health reports on essential applications.

Required Qualifications

  • Over four years’ experience in roles such as SRE, DevOps, production operations, or network operations center within high-availability environments, ideally in payments, e-commerce, SaaS, or gaming sectors.
  • Advanced troubleshooting capabilities to trace issues through multiple system layers including logs, metrics, application traces, database queries, and networking paths.
  • Proficiency with observability platforms like Datadog, or alternatives such as Grafana, Splunk, New Relic, or Elastic, including building queries, reading dashboards, and configuring alerts.
  • Experience in scripting languages such as Python, Go, or Bash to automate operational tasks and interact with APIs.
  • Strong fluency in English for both written and verbal communication to produce clear, concise incident documentation and updates.
  • Familiarity with Kubernetes architecture and cloud infrastructures, preferably Google Cloud Platform, understanding deployment objects, monitoring pods, and investigating related production issues.
  • Knowledge of Service Level Objectives (SLOs), error budgets, and alert burn-rate concepts to appropriately classify incident severity.
  • Experience with incident management tools including JIRA Service Management, PagerDuty or OpsGenie, Slack, and Confluence.
  • Interest or experience with AI/ML in operations such as anomaly detection, alert correlation, predictive monitoring, or automated remediation.
  • Capability to work 24x7 shifting schedules following a global handoff model, including rotating weekend on-call duties.

Preferred Additional Skills

  • Background in gaming, payment processing, or fintech sectors with stringent uptime and transaction integrity requirements.
  • Knowledge of Datadog Service Catalog, synthetic monitoring, real user monitoring, and database technologies like MySQL, PostgreSQL, Redis, Kafka.
  • Experience supporting CI/CD pipelines and deployment tools including GitLab CI, ArgoCD, Helm.
  • Experience administering JIRA Service Management workflows and automation or an ITIL Foundation certification.

Additional Information

This role's responsibilities may evolve to align with company objectives and personal growth. Applicants consent to background investigations in compliance with applicable laws after interview stages. The company upholds strict data privacy standards, with no external sharing or sales of applicant information. For data privacy concerns, inquiries can be directed to the provided company contact.

The hiring process may incorporate AI tools to assist with application reviews and assessments, but final hiring decisions rest with human evaluators.

Tools & software

Kubernetes required

How they work

Communication Teamwork & Collaboration Problem Solving Attention to Detail Initiative

Languages

English

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help