← Back to jobs

Engineering Manager, Production Engineering

harvey

Híbrido San Francisco
Engineering Public Relations

Job Score

100 pts
Hybrid model (+80) Engineering (+10) Public Relations (+10)

Why Harvey

At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come.

This is a rare chance to help build a generational company at a true inflection point. We have strong product-market fit and world-class investor support. We’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth — personal, professional, and financial — is unmatched.

Our team moves fast, takes ownership, and is deeply committed to the mission — operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job's Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we're never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we'd love to build with you.

At Harvey, the future of professional services is being written today — and we’re just getting started.

Role Overview

Harvey is building the AI platform trusted by the world's leading law firms and enterprises. Our infrastructure is the foundation that powers every customer interaction, every model inference, and every production workload.

We're looking for a Engineering Manager to lead our Infrastructure Production Engineering organization. This team is responsible for building and operating Harvey's core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations that enable engineering teams to move quickly with confidence.

In this role, you'll own the reliability, scalability, security, and efficiency of Harvey's infrastructure platform. You'll lead a team of high-performing engineers responsible for compute fleet management, capacity planning, infrastructure automation, and production operations. You'll partner closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey's rapid growth.

You'll report to the Head of Infrastructure and play a key leadership role in shaping the future of Harvey's infrastructure platform.

At Harvey, we value Decisiveness, Simplicity, and the belief that Job's Not Finished. We move quickly, prioritize clarity, and continuously raise the bar for engineering excellence.

What You'll Do

Leadership & Strategy

  • Lead, mentor, and grow a team of high-performing infrastructure engineers responsible for Harvey's production infrastructure foundation.

  • Foster a culture of operational excellence, engineering quality, customer ownership, and continuous improvement.

  • Partner with Engineering, Security, Product, and AI Infrastructure leaders to define long-term infrastructure strategy and execution priorities.

  • Drive technical direction for compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.

  • Lead cross-functional initiatives to improve reliability, scalability, security, operational efficiency, and infrastructure cost optimization.

Infrastructure Foundation & Production Operations

  • Own and operate Harvey's global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.

  • Manage compute resources to maximize utilization, performance, and service availability while supporting rapidly growing AI workloads.

  • Lead capacity planning, demand forecasting, and fleet lifecycle management to ensure infrastructure scales efficiently with business growth.

  • Operate and continuously improve Harvey's Kubernetes platform, including cluster provisioning, upgrades, monitoring, reliability, performance, and operational automation.

  • Own Harvey's Temporal-based workflow orchestration platform, ensuring reliable, scalable, and observable execution of distributed application workflows.

  • Drive infrastructure cost optimization through capacity management, resource rightsizing, workload efficiency improvements, and utilization monitoring.

  • Build and maintain secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.

  • Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.

  • Establish comprehensive observability, monitoring, alerting, incident response, and operational readiness practices across the infrastructure platform.

What You Have

  • 7+ years of software or infrastructure engineering experience, including 5+ years leading engineering teams.

  • Deep expertise operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.

  • Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.

  • Experience building and operating large-scale distributed systems with strong reliability, scalability, and performance characteristics.

  • Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.

  • Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.

  • Experience designing and operating observability platforms, including monitoring, logging, alerting, and incident response processes.

  • Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.

  • Demonstrated success leading complex cross-functional technical initiatives and influencing engineering strategy across organizations.

  • Excellent communication skills with the ability to communicate technical concepts clearly to both engineering and executive audiences.

  • A systems-thinking mindset and passion for building simple, reliable, and scalable infrastructure platforms.

Nice to Have

  • Experience operating workflow orchestration platforms such as Temporal.

  • Experience supporting AI/ML or LLM infrastructure at scale.

  • Experience managing GPU fleets, high-performance compute infrastructure, or large-scale capacity planning.

  • Experience with multi-cloud infrastructure or hybrid cloud environments.

Compensation

$260,000 - $340,000 USD

Depending on your location, an Applicant Privacy Notice may apply to you. You can find all of our Applicant Privacy Notices [here].

#LI-AN2

Harvey is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made by emailing accommodations@harvey.ai

What did you think of this job?

Comments 0

Want to leave a comment?
Sign in or create your account in seconds to join the discussion.
Loading comments...

About Engineering

Software Engineering goes beyond traditional development, focusing on scalability, performance, and system architecture. Software engineers are responsible for designing infrastructures that support millions of simultaneous users.

Skills include microservices architecture, DevOps, cloud computing, application security, and performance optimization. Knowledge of containerization (Docker, Kubernetes) and CI/CD is increasingly required.

Senior software engineers are rare and highly compensated professionals, with opportunities at major global tech companies.

About Public Relations

The Public Relations (PR) area focuses on managing the reputation, image, and communication of an organization with its various stakeholders (such as clients, investors, employees, media, and the community). PR professionals develop corporate communication strategies, manage media relations (press relations), organize institutional events, and work in image crisis prevention and management.

Discover Other Areas

Understand the scope of work, key skills, and tools used in different career areas.

About Content Manager

The Content Manager is the professional responsible for leading the entire content strategy, production, and management of an organization. They define the editorial strategy, coordinate writing teams, and ensure content aligns with business goals and brand identity.

Key skills include content strategy, editorial planning, content audit, buyer persona, customer journey, content ops, content governance, performance metrics (ROI, engagement, organic traffic), and team management. Knowledge of WordPress, Contentful, Notion, and analytics tools is a differentiator.

Content Managers in technology companies are highly valued, especially those who can align content with conversion funnels, lead multidisciplinary teams, and use data to optimize editorial strategy. The field offers opportunities from content manager to head of content, with a focus on strategy, quality, and scale.

About People Analyst

The People Analyst is the professional responsible for transforming people data into strategic insights for HR decision-making. They combine data analysis knowledge with people management vision to help organizations understand workforce metrics, turnover, engagement, and diversity.

Key skills include people analytics, workforce analytics, turnover and retention analysis, HR metrics (time-to-hire, cost-per-hire, e-NPS), data visualization (Power BI, Tableau, Visier), workforce planning, and compensation analysis. Knowledge of statistics, SQL, and people analytics tools is a differentiator.

People Analysts in technology companies are highly valued, especially those who can translate complex people data into actionable insights for retention, diversity, and growth strategies. The field offers opportunities from HR analyst to head of people analytics, with a focus on data-driven people management.

About Automation Engineer

The Automation Engineer is the professional responsible for designing, developing, and implementing solutions that automate manual and repetitive processes in IT, infrastructure, testing, and operations. They combine programming knowledge with DevOps and SRE vision to eliminate manual tasks and increase operational efficiency.

Key skills include Infrastructure as Code (Terraform, Ansible, Pulumi), CI/CD (Jenkins, GitHub Actions, GitLab CI), test automation (Selenium, Cypress, Playwright), network automation (Netconf, SDN), RPA (UiPath, Power Automate), and scripting (Python, Bash, PowerShell). Knowledge of Kubernetes, GitOps (ArgoCD, Flux), and automation platforms is a differentiator.

Automation Engineers in technology companies are highly valued, especially those who can create automated deployment pipelines, self-healing infrastructure, and internal developer platforms (IDP). The field offers opportunities from junior automation engineer to automation architect and head of automation.

About Advertising

The Advertising area is aimed at the planning, creation, and delivery of communication campaigns to promote brands, products, ideas, or services. Professionals in the sector work in advertising agencies or in-house marketing departments in creative fields (art direction, copywriting), strategic planning, account management, and media buying.

About Software Development

Software Development is one of the most dynamic and constantly evolving fields in the job market. Professionals in this area are responsible for creating, maintaining, and optimizing web, mobile, and desktop applications that impact millions of users daily.

Key languages and frameworks include JavaScript (React, Node.js, Vue.js), Python (Django, Flask), Java (Spring), PHP (Laravel), and TypeScript. Demand for full-stack developers continues to grow, especially in tech companies and startups.

Salaries range from entry-level to senior positions, with growing opportunities for remote work and international freelancing.

Career Guides

Technology Career Guide

Planning, skills, interviews, and professional growth in IT, Data Science, DevOps, and Product.

Read full guide →

Design Career Guide

UX/UI, Graphic Design, Product Design. Portfolio, tools, interviews, and growth in the Design field.

Read full guide →

Marketing Career Guide

SEO, Paid Media, Growth, Content Marketing. Certifications, tools, and strategies to grow in Digital Marketing.

Read full guide →

Finance Career Guide

Financial market, investments, corporate finance, certifications, and strategies to grow in the financial field.

Read full guide →

Communication Career Guide

Journalism, PR, Corporate Communication, Content Marketing, and Multimedia Production.

Read full guide →

Administration Career Guide

Business Management, HR, Logistics, Consulting, Project Management, and Entrepreneurship.

Read full guide →

Data Career Guide

Data Science, Data Engineering, BI, Machine Learning, and AI. From training to the job market.

Read full guide →

Product Career Guide

Product Management, Product Ownership, Agile, Scrum, and OKRs. From strategy to execution.

Read full guide →

Tech & Remote Glossary

Stop getting lost in interviews and job descriptions

The job market, especially within tech and global companies, has developed its own dialect. Not understanding these acronyms can make you lose valuable opportunities or poorly negotiate your contract. To end this problem, we created the Definitive Glossary for the Remote Professional.

🏢 Work Models & Routine

Async (Asynchronous Work)
A communication model where responses don't need to be immediate. Instead of back-to-back meetings, the team relies on well-structured documents, threads, and messages. It's the gold standard for global companies spanning multiple time zones.
Sync (Synchronous Work)
The opposite of Async. It requires the team to be online and available at the same time for meetings, live chats, and real-time collaboration.
Daily / Stand-up
A quick daily meeting (usually 15 minutes) common in Agile (Scrum) methodologies. The team answers three questions: What did I do yesterday? What will I do today? Are there any blockers?
All-Hands / Town Hall
A company-wide meeting involving all employees. Usually led by the founders (C-Levels) to present results, new goals, and answer team questions.
1:1 (One-on-One)
A recurring individual meeting between a professional and their direct manager. It is used for career alignment, feedback, and problem-solving, not just for project status updates.

💰 Contracts, Benefits & Compensation

PTO (Paid Time Off)
Instead of strict, categorized leave policies, modern US companies usually offer a flexible pool of days (e.g., 20 days of PTO, or even "Unlimited PTO") that you can use for vacations, sick days, or personal matters, while receiving your regular compensation.
Equity / Stock Options
Company ownership. The startup offers you the right to buy shares at a heavily discounted strike price in the future. If the company grows, goes public, or is acquired, these shares can be highly lucrative.
Vesting (Vesting Schedule)
The rule that controls your Equity. It usually lasts 4 years. You don't get all the shares on day one; you "earn" them gradually as you stay with the company. A "1-year Cliff" means you must stay for at least one year to receive your first batch of shares.
RSUs (Restricted Stock Units)
Unlike Stock Options (where you have the right to buy the stock), RSUs are actual shares the company grants you as a bonus or part of your compensation package, following a strict Vesting schedule.
Independent Contractor (1099 / B2B)
The most common international hiring model for global talent working for US companies. You act as a service provider (business-to-business), receiving the gross salary (often six-figure compensation) without standard local payroll tax deductions at the source.

🤖 Recruitment & Hiring Process

ATS (Applicant Tracking System)
The "robot" that reads your resume. Software like Greenhouse, Ashby, and Workday are used to filter candidates by keywords before a human even looks at the document. (Pro tip: this is why your resume must be clean, semantic, and have the right keywords).
JD (Job Description)
The document that lists the responsibilities, technical requirements, and benefits of the open position.
Cultural Fit
The interview stage that evaluates if your core values, communication style, and worldview align with the company's culture. This is the ultimate test of your Soft Skills.
Onboarding
The integration process. It's the period of your first few weeks at the company, where you get your access credentials, learn about the culture, study the internal documentation, and understand how the product works.

🚀 How to use this to your advantage?

The secret isn't just knowing what these acronyms mean, but using them actively. If during an interview for a premium tech role you ask, "How does your PTO policy and Vesting schedule work?", the recruiter will immediately perceive you as a high-level professional, familiar with the global market standards.

The remote job market requires preparation. And having the right vocabulary is the first big step to securing six-figure proposals and standing out among thousands of applicants.

Expert Tip

The Current State of the Cloud & DevOps Market

By Mondywork | The ultimate guide for tech professionals looking to level up their careers, land global remote roles, and master modern infrastructure.

The End of the "SysAdmin" and the Rise of Platform Engineering

If you've been tracking the most sought-after roles at startups and Fortune 500 companies, you've probably noticed a drastic shift. The Cloud Computing and DevOps market is no longer just about "keeping the app online." Today, companies demand extreme resilience, automation, and cost optimization (the famous FinOps).

DevOps culture has evolved. We are now in the era of Platform Engineering, where infrastructure teams build internal platforms to grant developers full autonomy (Self-Service). For those aiming for global remote roles with six-figure compensation, understanding this transition isn't just a plus—it's a strict requirement.

"Elite DevOps teams deploy code 973 times more frequently and boast a Mean Time To Recovery (MTTR) that is 6,570 times faster than low-performing teams."

— State of DevOps Report (DORA / Google Cloud)

What Do Recruiters and ATS Systems Really Look For?

Automated filters (ATS platforms like Greenhouse, Ashby, and Workday) are hardwired to scan for a very specific set of skills in Cloud/DevOps resumes. Simply listing 20 tools won't cut it anymore; you must demonstrate culture and real-world impact.

  • Infrastructure as Code (IaC): Forget clicking around the AWS console. Declarative tools like Terraform, Pulumi, and Ansible are the true industry standards.
  • Container Orchestration: Mastery of Docker and, above all, Kubernetes (K8s) is practically mandatory for premium roles.
  • CI/CD Culture: Building automated pipelines (GitHub Actions, GitLab CI, ArgoCD) that guarantee continuous and secure deployments.
  • Cloud Providers and Architecture: Proficiency in at least one major public cloud provider (AWS, Google Cloud, or Microsoft Azure), with a strong focus on Serverless, microservices, and system resilience.

How to Transition: The Roadmap to Success

Many Back-End developers, support analysts, and network professionals want to pivot to Cloud and DevOps. If that's your goal, follow this market-validated roadmap:

  1. Master the Basics (Linux and Networking): Modern DevOps runs on Linux. Understanding permissions, processes, bash scripting, and the fundamentals of TCP/IP, DNS, and Load Balancers is your foundation.
  2. Learn to Code: A DevOps engineer doesn't need to be a Senior Developer, but you must know how to automate tasks. Python and Go (Golang) are currently the most highly valued languages in this ecosystem.
  3. Are Certifications Worth It? Yes, especially to get past the initial HR screening. Certifications like the AWS Certified Solutions Architect – Associate or the CKA (Certified Kubernetes Administrator) carry massive weight in the global market.
  4. Build a Real Portfolio: Publish a project on GitHub where you provision infrastructure via Terraform, set up a CI/CD pipeline using GitHub Actions, and deploy a simple app to a Kubernetes cluster. Practical projects always beat theoretical resumes.

Essential References and Study Links

To stay up-to-date, we highly recommend following official organizations and annual market reports:


🚀 Want access to the best remote Cloud & DevOps jobs?

Join our mailing list and get the best jobs delivered straight to you, tailored to your areas of interest. Currently, we partner with over 200 companies that advertise their premium, remote-first roles with us. Sign up now and get the best opportunities in the global market right in your inbox!