Senior Observability Engineer
Galway, County Galway, Ireland (Hybrid) · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 6 ಗಂಟೆಗಳು ಹಿಂದೆ
- Work mode
- Hybrid
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Rent The Runway
Rent The Runway is revolutionizing fashion by creating the world's first shared Closet in the Cloud. Since its inception in 2009, the company has disrupted the $2.4 trillion fashion sector by promoting a more sustainable, joyful, and economical way for women to dress. Offering designer apparel and accessories through flexible subscription, one-time rentals, or purchases, Rent The Runway combines proprietary technology with a unique reverse logistics system. The company has been recognized multiple times in CNBC’s “Disruptor 50” and Fast Company’s Most Innovative Companies lists.
Rent The Runway’s European Technology Hub, located in Galway, Ireland, was launched in 2019. The Galway office expands the company's Software Engineering, Product Development, Machine Learning Engineering, and Data Science capabilities outside the US and provides various career growth opportunities.
Role Overview
The Senior Observability Engineer will lead the design, implementation, and enhancement of telemetry systems ensuring reliability, performance, and resilience of platforms. This role involves setting observability standards across engineering and infrastructure teams, improving incident response and system understanding to elevate delivery outcomes. The engineer will define key metrics, detection methods, and response strategies, owning major initiatives and fostering observability adoption throughout the company.
Main Responsibilities
- Architect, deliver, and continually upgrade observability platforms utilizing tools such as Splunk Observability Cloud and Google Cloud Observability.
- Develop scalable and automated telemetry pipelines with Terraform and Infrastructure-as-Code methodologies to ensure auditability and ease of use.
- Create and refine best practices for instrumentation of metrics, traces, logs, and events to enable precise alerting and actionable data across systems.
- Collaborate with various teams—application, platform, security, compliance—to embed observability throughout development lifecycle phases, including instrumentation, SLOs, and post-incident reviews.
- Lead the definition and use of internal standards relating to SLIs, error budgets, and overall system health monitoring.
- Promote and integrate AI-enhanced tools and workflows for faster incident detection, debugging, and root cause analysis within observability ecosystems.
- Identify and solve cross-team observability challenges by simplifying and standardizing solutions to create durable systems.
- Participate in on-call rotations with SRE/Platform teams to understand production issues and enhance alerting and response processes.
- Provide guidance and mentorship to engineers on applying telemetry principles to increase system visibility.
- Drive a reliability-first culture through reusable frameworks, shared documentation, and training initiatives that reduce operational toil while improving visibility.
Candidate Profile
- Minimum of five years in SRE, DevOps, or platform engineering with significant expertise in observability design and telemetry systems.
- Recognized authority on observability tooling and best practices, especially within distributed cloud-native systems.
- Demonstrated success in delivering impactful technical projects that amplify system visibility, decrease operational risks, and improve incident response at scale.
- Expert-level understanding of telemetry data types—metrics, logs, traces, events—and their application to both infrastructure and product services.
- Hands-on experience with telemetry pipeline management leveraging Terraform, CI/CD workflows, and Kubernetes service instrumentation.
- Strong grasp of system design elements such as service meshes, asynchronous messaging, and distributed system behaviors.
- Excellent communication skills capable of translating complex technical concepts to both technical peers and non-technical stakeholders.
- Proactive ownership mentality with a curiosity-driven approach to identifying failure scenarios, simplifying complex systems, and conducting structured root cause analysis.
- Effective collaborator able to align teams, build consensus, and lead influence across diverse groups and functions.
Benefits
- Generous paid leave programs including annual, bereavement, and family sick leave to support personal and family care.
- Universal paid parental leave for all parents along with a flexible return-to-work arrangement to accommodate family needs.
- Paid sabbatical after five consecutive years of service for rest and rejuvenation.
- Competitive pension scheme to support long-term financial security.
- Comprehensive healthcare including dental and dependent care starting from day one of employment.
- Regular company events and team outings that promote a positive and engaging workplace culture.
- Hybrid work model requiring presence in the Galway office 2-3 days per week, with the option to work remotely 2 days per week.
Additional Information
Rent The Runway is committed to equal employment opportunities and prohibits discrimination based on legally protected grounds including gender, marital status, age, disability, race, religion, sexual orientation, and other factors.
Applicants implicitly acknowledge understanding and acceptance of Rent The Runway’s Candidate Privacy Policy by submitting their application.
Level
Senior