Partly

Head of Platform

Partly

Auckland, New Zealand · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
1 hour ago
Work mode
In office
Resume
Required to apply

Where you'll work

Job description

About Partly

Partly is an innovative company creating AI infrastructure tailored to the global repair industry, beginning with the $2 trillion automotive repair market. Their unique Interpreter model is the first AI designed specifically to assess vehicle damage and determine required parts. Founded by former Rocket Lab engineers, Partly has grown rapidly, recently securing $50 million in Series B funding from prominent investors. The company is headquartered in Austin, Texas, with offices including Auckland, New Zealand.

Role Overview

The Head of Platform will lead and develop the team responsible for the internal platform that supports software development and operation at Partly. This includes overseeing infrastructure foundations, reliability operations, security protocols, and the specialised AI/ML infrastructure needed for model training and serving. The role focuses on enabling fast, efficient development through strategic platform leadership and hands-on delivery, treating the platform as an internal product to maximize developer productivity, uptime, and cost efficiency.

Key Responsibilities

  • Lead and grow the Platform function, setting strategy, managing priorities, and building the team across areas like SRE, DevEx, Infrastructure, and Security enablement.
  • Enhance developer experience by creating streamlined workflows for building, deploying, and operating services, driving adoption and productivity improvements.
  • Own reliability pillars such as observability, alerting, incident response, postmortems, SLO/error budgets, and reduce operational toil.
  • Maintain scalable, secure, and maintainable cloud and Kubernetes infrastructure using Infrastructure-as-Code and automation tools like Terraform and ArgoCD.
  • Build and oversee AI infrastructure supporting GPUs/TPUs provisioning, model training and fine-tuning orchestration, scalable serving, and cost-efficient inference.
  • Design platform capabilities that natively support AI agents and forward-deployed engineers to operate autonomously, including safe execution environments and observability.
  • Collaborate with security teams to embed secure-by-default practices such as IAM, secrets management, vulnerability handling, and policy automation.
  • Manage cost and performance trade-offs with FinOps visibility to optimize unit economics without compromising reliability or developer velocity.
  • Work cross-functionally with product and engineering leaders to drive outcomes that benefit customers and the business.
  • Contribute technically through design reviews, incident mitigation, prototyping, and setting technical standards while building a self-sufficient team.

Required Skills and Experience

  • Proven leadership in platform, infrastructure, or SRE teams with experience managing roadmaps, stakeholders, and fast-paced environments.
  • Expertise in Site Reliability Engineering practices including SLOs, incident management, observability, capacity planning, and resilience engineering.
  • In-depth knowledge of cloud platforms (preferably GCP) and Kubernetes for running production workloads at scale.
  • Hands-on experience with Infrastructure-as-Code and GitOps workflows using tools such as Terraform and ArgoCD.
  • Strong focus on developer experience, ability to create and promote ‘golden paths,’ improve workflows, and measure impact.
  • Security integration skills including IAM, secrets management, vulnerability management, and compliance (SOC2/ISO) familiarity is a plus.
  • Solid understanding of computer systems fundamentals: concurrency, networking, Linux internals, performance profiling, distributed systems, and reliability patterns.
  • Knowledge of ML/AI infrastructure components such as GPU provisioning, model training and serving orchestration, MLOps, and cost considerations of production AI workloads.
  • Experience designing platforms for AI agent users with self-service and autonomous execution capabilities, supporting rapid deployment by embedded engineers.
  • Exceptional communication skills to align senior stakeholders, explain complex trade-offs, and mentor engineers.
  • A bias for action with accountability for foundational system outcomes.

Additional Skills (Optional)

Experience scaling platform operations across multiple teams and familiarity with Partly's stack (GCP, ArgoCD, GitLab CI, Kafka, Postgres) are advantageous.

Benefits

  • Work environment with minimal bureaucracy and high trust, empowering employees through simple policies.
  • Competitive salary and equity packages rewarding full-time employees with company growth.
  • Flexible working hours and office-first culture in locations with significant teams, including Auckland.
  • Two meeting-free days weekly dedicated to uninterrupted deep work.
  • Unlimited leave policy with trust for employees to recharge as needed.
  • Well-equipped offices with amenities that foster productivity and community.
  • Opportunities for continuous learning from subject matter experts and company leaders.
  • Regular in-person team events, quarterly company-wide gatherings, and global offsites.
  • Supportive parental leave policies offering flexible return-to-work arrangements.
  • Payroll giving program promoting charitable donations.
  • Relocation assistance for those moving domestically or internationally to join the Auckland office.

Additional Information

Partly's headquarters are in Austin, TX, but onboarding occurs near your location with travel supported for quarterly team events. Relocation packages are available for successful candidates moving to Partly offices.

Work styles they’re looking for

Communication

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help