IT Infrastructure Support Site Reliability Engineer II
Ireland, England, United Kingdom • Penuh Waktu
Jadilah yang pertama mendaftar
- Pengalaman
- 6+ tahun
- Gaji
- —
- Lowongan
- 1
- Diposting
- 1 hari yang lalu
- Mode kerja
- Di kantor
- Melanjutkan
- Wajib mendaftar
Tempat Anda akan bekerja
Deskripsi pekerjaan
About the Role
We are hiring a skilled Infrastructure Automation Engineer to join our IT Infrastructure Support team focused on delivering high reliability, scalability, and performance for critical physical security infrastructure and associated systems. This role blends software engineering with operational expertise to develop and maintain automation tools, monitoring frameworks, and processes that support enterprise-grade server, network, and security device management. Collaboration with various teams to define and uphold service level objectives is essential, along with reducing manual work through intelligent automation and driving ongoing system resilience improvements. The position entails 24x5 coverage with participation in an on-call rotation to provide continuous mission-critical infrastructure support.
Key Duties
- Collaborate with leadership to define, track, and enforce Service Level Indicators (SLIs) and Objectives (SLOs) regarding infrastructure tools such as configuration compliance, patch deployment success, and latency metrics.
- Offer Level 3 expertise for incidents relating to tooling, focusing on automating remediation workflows to minimize Mean Time To Repair (MTTR) via advanced automation and runbook creation.
- Automate repetitive manual infrastructure tasks aiming for significant operational overhead reduction, such as halving server build times.
- Lead root cause analyses and blameless postmortems for major service-impacting incidents to enhance tooling reliability and overall infrastructure robustness.
- Create and maintain automated workflows and scripts to manage and synchronize data in asset management platforms (e.g., NetBox), configuration management databases, and monitoring systems used by internal and external stakeholders.
- Design, develop, and implement full-stack applications, custom plugins, and automation solutions that extend management and monitoring capabilities including direct device configuration.
- Manage automated Infrastructure-as-Code for both Windows and Linux servers using tools like Ansible, Terraform, or Puppet, incorporating drift detection and automated remediation features.
- Construct end-to-end automation pipelines for vulnerability patching, security baseline enforcement based on CIS benchmarks, and continuous audits ensuring compliance with internal and regulatory standards concerning physical security devices.
- Develop API-driven tools for network configuration, automated firmware updates, change validation, and continuous monitoring of network health for the device fleet.
- Deploy and standardize monitoring agents and centralized log aggregation systems alongside customizable dashboards equipped with alerting based on critical SLIs such as latency, errors, traffic, and saturation for servers and edge devices.
- Create automation scripts integrating intelligent ticket handling, problem validation, and escalation processes within enterprise ticketing systems to consistently meet the 2-hour initial response SLA.
- Engage in a 24x5 on-call rotation to ensure swift incident responses and uninterrupted infrastructure service.
Required Skills and Experience
- Over six years of experience in Infrastructure Automation or Infrastructure Engineering roles.
- Strong scripting capabilities in Python, Bash, and PowerShell; experience with Go for backend services and API development is also required.
- Expertise with Infrastructure-as-Code solutions such as Terraform, Ansible, Chef, or Puppet, including skills in drift detection, version control, and automated remediation workflows.
- Deep knowledge of Linux and Windows server operating systems with Tier 3 troubleshooting proficiency, system hardening, and managing enterprise-scale environments.
- Understanding of enterprise-level networking concepts, Cisco device administration, and network automation protocols like NETCONF and RESTCONF, plus familiarity with network monitoring and flow analysis tools.
- Hands-on experience deploying monitoring tools such as Prometheus, Grafana, and Datadog, as well as centralized logging solutions like the ELK Stack, with the ability to build tailored dashboards and alerting mechanisms.
Additional Information
The role requires availability for a 24x5 schedule with on-call duties as part of a rotation to maintain uninterrupted support for mission-critical infrastructure components, including physical security devices and network tooling.