Open IT Jobs
SVP, Lead AI/MLOps Infrastructure Engineer - Full Time - Hybrid
SVP, Lead AI/MLOps Infrastructure Engineer – Full Time – Hybrid
We’re partnering with our client, a fast-growing fintech firm, on a senior-level hire to lead and own the infrastructure behind their AI and machine learning platforms.
This is a highly visible, hands-on leadership role where you’ll own the end-to-end AI/ML platform stack, from training and inference infrastructure to model serving, reliability, and cost, while setting the MLOps roadmap and standards for the team.
The priority here is MLOps and AI infrastructure first, built on deep AWS, Terraform, and production ML / GenAI experience.
What You’ll Be Doing
- Own the end-to-end AI/ML platform stack: orchestration, compute (including GPU), storage, and model serving
- Build and operate MLOps pipelines across the full model lifecycle: training, validation, versioning, and deployment
- Productionize AI/ML and GenAI (LLM) workloads in partnership with ML engineers and data scientists
- Design and manage cloud-native AWS infrastructure using Kubernetes, and own Infrastructure as Code standards (Terraform)
- Own SLAs/SLOs for model serving and inference; lead monitoring, drift detection, and incident response
- Build internal tooling and standardized environments that boost ML engineer productivity (e.g., MLflow, Kubeflow, Weights & Biases, Ray)
- Drive data governance, privacy compliance, and cost optimization across training and inference workloads
- Set the MLOps roadmap and mentor engineers on infrastructure and MLOps best practices
What They’re Looking For
- 15+ years of experience in DevOps, SRE, or platform engineering, with AWS as primary cloud
- Proven, hands-on experience building and operating MLOps pipelines in production (key priority)
- Experience with MLOps tooling, including model registries, experiment tracking, and feature stores
- Exposure to Generative AI / LLM workloads, including AWS Bedrock
- Strong Infrastructure as Code (Terraform) and scripting skills (Python or similar)
- Solid Linux, systems, and troubleshooting fundamentals
- Excellent communicator, comfortable collaborating across teams
Nice to Have
- Hands-on Kubernetes, containerized workloads, and cloud networking
- Experience in regulated or fintech environments
- Background optimizing costs for compute-intensive (GPU) workloads
Why This Role
- Own and shape the AI platform at a growing fintech
- High-impact leadership role setting MLOps strategy and standards
- Strong compensation: $200K–$230K base + bonus + equity
- Comprehensive benefits, including retirement match and unlimited PTO
- Hybrid model: 4 days onsite / 1 day remote (NYC area)
Job ID: 5540