This page was automatically translated and may contain errors. View in English.
A

IT Infrastructure Support Site Reliability Engineer II

Astreya

Ireland, England, United Kingdom • Vollzeit

Bewerben Sie sich als Erste/r!

Erfahrung
Ab 6 Jahren
Gehalt
Stellenangebote
1
Veröffentlicht
vor 18 Stunden
Arbeitsmodus
Im Büro
Wieder aufnehmen
Bewerbung erforderlich

Wo Sie arbeiten werden

Stellenbeschreibung

About the Role

We are hiring a skilled Infrastructure Automation Engineer to join our IT Infrastructure Support team focused on delivering high reliability, scalability, and performance for critical physical security infrastructure and associated systems. This role blends software engineering with operational expertise to develop and maintain automation tools, monitoring frameworks, and processes that support enterprise-grade server, network, and security device management. Collaboration with various teams to define and uphold service level objectives is essential, along with reducing manual work through intelligent automation and driving ongoing system resilience improvements. The position entails 24x5 coverage with participation in an on-call rotation to provide continuous mission-critical infrastructure support.

Key Duties

  • Collaborate with leadership to define, track, and enforce Service Level Indicators (SLIs) and Objectives (SLOs) regarding infrastructure tools such as configuration compliance, patch deployment success, and latency metrics.
  • Offer Level 3 expertise for incidents relating to tooling, focusing on automating remediation workflows to minimize Mean Time To Repair (MTTR) via advanced automation and runbook creation.
  • Automate repetitive manual infrastructure tasks aiming for significant operational overhead reduction, such as halving server build times.
  • Lead root cause analyses and blameless postmortems for major service-impacting incidents to enhance tooling reliability and overall infrastructure robustness.
  • Create and maintain automated workflows and scripts to manage and synchronize data in asset management platforms (e.g., NetBox), configuration management databases, and monitoring systems used by internal and external stakeholders.
  • Design, develop, and implement full-stack applications, custom plugins, and automation solutions that extend management and monitoring capabilities including direct device configuration.
  • Manage automated Infrastructure-as-Code for both Windows and Linux servers using tools like Ansible, Terraform, or Puppet, incorporating drift detection and automated remediation features.
  • Construct end-to-end automation pipelines for vulnerability patching, security baseline enforcement based on CIS benchmarks, and continuous audits ensuring compliance with internal and regulatory standards concerning physical security devices.
  • Develop API-driven tools for network configuration, automated firmware updates, change validation, and continuous monitoring of network health for the device fleet.
  • Deploy and standardize monitoring agents and centralized log aggregation systems alongside customizable dashboards equipped with alerting based on critical SLIs such as latency, errors, traffic, and saturation for servers and edge devices.
  • Create automation scripts integrating intelligent ticket handling, problem validation, and escalation processes within enterprise ticketing systems to consistently meet the 2-hour initial response SLA.
  • Engage in a 24x5 on-call rotation to ensure swift incident responses and uninterrupted infrastructure service.

Required Skills and Experience

  • Over six years of experience in Infrastructure Automation or Infrastructure Engineering roles.
  • Strong scripting capabilities in Python, Bash, and PowerShell; experience with Go for backend services and API development is also required.
  • Expertise with Infrastructure-as-Code solutions such as Terraform, Ansible, Chef, or Puppet, including skills in drift detection, version control, and automated remediation workflows.
  • Deep knowledge of Linux and Windows server operating systems with Tier 3 troubleshooting proficiency, system hardening, and managing enterprise-scale environments.
  • Understanding of enterprise-level networking concepts, Cisco device administration, and network automation protocols like NETCONF and RESTCONF, plus familiarity with network monitoring and flow analysis tools.
  • Hands-on experience deploying monitoring tools such as Prometheus, Grafana, and Datadog, as well as centralized logging solutions like the ELK Stack, with the ability to build tailored dashboards and alerting mechanisms.

Additional Information

The role requires availability for a 24x5 schedule with on-call duties as part of a rotation to maintain uninterrupted support for mission-critical infrastructure components, including physical security devices and network tooling.

Lassen Sie es so, wenn Sie eine Antwort wünschen – wir werden es für nichts anderes verwenden.

Zum Durchsuchen klicken, per Drag & Drop, oder Paste ein Screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Maximal 20 MB pro Datei · Bis zu 5 Dateien

🤖
Online · Sofortige KI-Hilfe