remoty.work

Archived listing

This role was posted over 30 days ago and is no longer accepting applications. We keep it for reference, but the employer may have already filled it. Browse current worldwide remote jobs or see today's verified listings.

F /100

Sr. Director, Platform & AI Infrastructure

Fortive logo
Fortive
United States 🌐 Remote

Job Description

The Sr. Director, Platform & AI Infrastructure leads the platforms that power Gordian's products and our AI transformation: cloud infrastructure, data platforms, ML/AI infrastructure, incident response, and observability.

This is a builder role. You'll stand up the AI platform that our product and engineering teams build on, modernize how we run production, and establish the incident response and observability programs that scale with us.

Key responsibilities

  • AI platform and production ML. Own the AI/ML platform: GPU capacity strategy, model serving and inference latency, training and fine-tuning infrastructure, MLOps and evaluation pipelines, vector and feature stores, and the RAG and agentic patterns our product teams build on. Partner with product engineering and architecture on build-vs-buy decisions across foundation model providers and open-source.
  • Incident and observability management. Build out the incident response program: on-call structure, severity definitions, incident command, communication standards, postmortems, and follow-through on systemic fixes. Develop the observability stack across metrics, logs, traces, and synthetics. Set SLOs and report against them.
  • Cloud infrastructure. Operate the Azure and OCI footprint. Infrastructure-as-code, capacity planning, and reliability across CPU and GPU workloads.
  • Data platforms. Operations, performance, HA/DR, and roadmap for Oracle, SQL Server, MongoDB, and similar.
  • Communication. Brief executives on reliability and risk. Lead internal communication during incidents. Speak to customers when major incidents affect them.
  • People. Lead a globally distributed team of managers and senior ICs. Maintain a strong culture and leadership bench.

Required

  • 10+ years in platform engineering, SRE, infrastructure, or AI/ML infrastructure, with 5+ leading teams.
  • Production experience running ML/AI workloads at scale, including GPU infrastructure, model serving, MLOps, or LLM/inference platforms.
  • Familiarity with the modern AI stack: vector databases, RAG, agent frameworks, evaluation, and the build-vs-buy tradeoffs across foundation model providers and open-source.
  • Built or rebuilt an incident response or observability program at scale.
  • Measurable reliability improvements (MTTR, availability, change failure rate) in a cloud environment.
  • Effective communicator with executives, the Board, customers, and the company during incidents.
  • Zero-trust and modern identity platforms.

Originally posted on Himalayas

Did you apply to this job?

What actually happened matters more than our score. Fifteen seconds, and it changes the grade the next person sees.

📱 Want jobs like this daily? Join @remotywork on Telegram — top 5 scored remote jobs every weekday, no spam.

remoTy Weekly — Every Monday

Get the best remote jobs in your inbox

Curated remote jobs scored A–F. Ghost-job alerts. Market pulse. No spam — just signal.

Prefer instant updates? Join @remotywork on Telegram — daily top jobs, no inbox clutter.