← Back to jobs

Staff Software Engineer, Production Engineering

harvey

Híbrido San Francisco
Development Public Relations

Job Score

100 pts
Hybrid model (+80) Development (+10) Public Relations (+10)

Why Harvey

At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come.

This is a rare chance to help build a generational company at a true inflection point. We have strong product-market fit and world-class investor support. We’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth — personal, professional, and financial — is unmatched.

Our team moves fast, takes ownership, and is deeply committed to the mission — operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job's Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we're never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we'd love to build with you.

At Harvey, the future of professional services is being written today — and we’re just getting started.

Role Overview

Harvey is building the AI platform trusted by the world’s leading law firms and enterprises. Our infrastructure is the foundation that powers every customer interaction, every model inference, and every production workload.

We’re looking for a Production Engineer to help build and operate Harvey’s core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. You’ll work on the systems that enable engineering teams to move quickly and operate reliable services at scale.

In this role, you’ll improve the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform. You’ll solve complex production challenges across compute fleet management, capacity planning, infrastructure automation, and production operations. You’ll partner closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth.

At Harvey, we value Decisiveness, Simplicity, and the belief that Job’s Not Finished. We move quickly, prioritize clarity, and continuously raise the bar for engineering excellence.

What You'll Do

Infrastructure Engineering & Technical Leadership

  • Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads.

  • Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.

  • Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost.

  • Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions.

  • Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely.

  • Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship.

Infrastructure Foundation & Production Operations

  • Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.

  • Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads.

  • Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth.

  • Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation.

  • Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring.

  • Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.

  • Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.

  • Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform.

  • Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements.

What You Have

  • 10+ years of software, infrastructure, site reliability, or production engineering experience.

  • Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.

  • Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.

  • Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics.

  • Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.

  • Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.

  • Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response.

  • Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.

  • A track record of driving complex, cross-functional technical initiatives and influencing engineering decisions without relying on formal authority.

  • Excellent communication skills and the ability to explain technical concepts clearly to engineering partners and other stakeholders.

  • A systems-thinking mindset and a passion for building simple, reliable, and scalable infrastructure platforms.

Nice to Have

  • Experience supporting AI/ML or LLM infrastructure at scale.

  • Experience operating GPU fleets, high-performance compute infrastructure, or large-scale capacity planning.

  • Experience with multi-cloud infrastructure or hybrid cloud environments.

  • Experience building internal platforms or developer tooling that improves engineering velocity and production safety.

Compensation

$231,000 - $340,000 USD

Depending on your location, an Applicant Privacy Notice may apply to you. You can find all of our Applicant Privacy Notices here.

#LI-AN2

Harvey is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made by emailing accommodations@harvey.ai

What did you think of this job?

Comments 0

Want to leave a comment?
Sign in or create your account in seconds to join the discussion.
Loading comments...

About Software Development

Software Development is one of the most dynamic and constantly evolving fields in the job market. Professionals in this area are responsible for creating, maintaining, and optimizing web, mobile, and desktop applications that impact millions of users daily.

Key languages and frameworks include JavaScript (React, Node.js, Vue.js), Python (Django, Flask), Java (Spring), PHP (Laravel), and TypeScript. Demand for full-stack developers continues to grow, especially in tech companies and startups.

Salaries range from entry-level to senior positions, with growing opportunities for remote work and international freelancing.

About Public Relations

The Public Relations (PR) area focuses on managing the reputation, image, and communication of an organization with its various stakeholders (such as clients, investors, employees, media, and the community). PR professionals develop corporate communication strategies, manage media relations (press relations), organize institutional events, and work in image crisis prevention and management.

Discover Other Areas

Understand the scope of work, key skills, and tools used in different career areas.

About Social Media

The Social Media area is one of the most dynamic and constantly evolving fields in digital marketing. Social media professionals are responsible for creating, managing, and optimizing brand presence on digital platforms, building engagement and community with the target audience.

Key skills include social media management (Instagram, TikTok, LinkedIn, Facebook, YouTube), social media content creation, community management, paid social media (Meta Ads, LinkedIn Ads, TikTok Ads), metrics analysis, and strategic planning. Tools like Hootsuite, Sprout Social, Buffer, Later, and analytics platforms are essential.

Social media professionals in technology companies are highly valued, especially those who master paid social, social media analytics, and content strategies for different platforms. The field offers opportunities from analyst to head of social media, with a focus on growth, engagement, and return on investment.

About Scrum Master

The Scrum Master is the professional responsible for facilitating the adoption of Scrum and agile practices within development teams. They act as servant leaders, removing impediments, promoting continuous improvement, and ensuring Scrum events and ceremonies happen in the best possible way.

Key skills include event facilitation (sprint planning, daily, review, retrospective), backlog management, team coaching, conflict resolution, and agile metrics (velocity, burndown, cycle time). Knowledge of Jira, Trello, Azure DevOps, and frameworks like Kanban, XP, and SAFe is a differentiator.

Scrum Masters in technology companies are highly valued, especially those who can promote team autonomy, create psychologically safe environments, and lead agile transformations at scale. The field offers opportunities from junior scrum master to agile coach, head of agile, and director of agile transformation.

About Information Security

The Information Security area is one of the most strategic and in-demand fields in the technology market. With the rise of cyberattacks, data breaches, and regulations like LGPD and GDPR, companies of all sizes invest heavily in professionals who can protect their digital assets.

Key specializations include Network Security, Cloud Security (AWS, Azure, GCP), Offensive Security (Penetration Testing, Red Team), Defensive Security (SOC, Blue Team), AppSec, and Security Governance. Tools like SIEM (Splunk, QRadar), firewalls, EDR, and Vulnerability Management platforms are essential.

Certifications like CISSP, CEH, OSCP, CompTIA Security+, and AWS Security Specialty are important differentiators. Information security professionals are among the highest-paid in the sector, with growing demand especially in fintechs, healthtechs, and large enterprises.

About Fullstack

Fullstack developers are versatile professionals capable of working on both frontend and backend of web and mobile applications. They master multiple technologies and can build complete products end-to-end, from the user interface to server infrastructure.

Key skills include proficiency in at least one complete stack (React/Vue/Angular + Node.js/PHP/Python/Java), databases (SQL and NoSQL), REST/GraphQL APIs, Git versioning, CI/CD, and basic infrastructure knowledge (Docker, cloud). Clean architecture, DDD, and testing are important differentiators.

Fullstack developers are highly valued in startups and companies that need versatile and autonomous professionals. The field offers opportunities from junior developer to software architect, with a focus on complete delivery, holistic product vision, and ability to work across multiple application layers.

About Customer Success

Customer Success is the area responsible for ensuring clients achieve their goals when using the product or service. It is a strategic function for retention, expansion, and customer satisfaction.

Key skills include account management, churn analysis, NPS, onboarding, upsell, and cross-sell. Knowledge of CS tools like Gainsight, Totango, and ChurnZero is a differentiator.

CS is becoming increasingly strategic in SaaS companies, with professionals directly contributing to recurring revenue growth (MRR/ARR).

Career Guides

Technology Career Guide

Planning, skills, interviews, and professional growth in IT, Data Science, DevOps, and Product.

Read full guide →

Design Career Guide

UX/UI, Graphic Design, Product Design. Portfolio, tools, interviews, and growth in the Design field.

Read full guide →

Marketing Career Guide

SEO, Paid Media, Growth, Content Marketing. Certifications, tools, and strategies to grow in Digital Marketing.

Read full guide →

Finance Career Guide

Financial market, investments, corporate finance, certifications, and strategies to grow in the financial field.

Read full guide →

Communication Career Guide

Journalism, PR, Corporate Communication, Content Marketing, and Multimedia Production.

Read full guide →

Administration Career Guide

Business Management, HR, Logistics, Consulting, Project Management, and Entrepreneurship.

Read full guide →

Data Career Guide

Data Science, Data Engineering, BI, Machine Learning, and AI. From training to the job market.

Read full guide →

Product Career Guide

Product Management, Product Ownership, Agile, Scrum, and OKRs. From strategy to execution.

Read full guide →

Tech & Remote Glossary

Stop getting lost in interviews and job descriptions

The job market, especially within tech and global companies, has developed its own dialect. Not understanding these acronyms can make you lose valuable opportunities or poorly negotiate your contract. To end this problem, we created the Definitive Glossary for the Remote Professional.

🏢 Work Models & Routine

Async (Asynchronous Work)
A communication model where responses don't need to be immediate. Instead of back-to-back meetings, the team relies on well-structured documents, threads, and messages. It's the gold standard for global companies spanning multiple time zones.
Sync (Synchronous Work)
The opposite of Async. It requires the team to be online and available at the same time for meetings, live chats, and real-time collaboration.
Daily / Stand-up
A quick daily meeting (usually 15 minutes) common in Agile (Scrum) methodologies. The team answers three questions: What did I do yesterday? What will I do today? Are there any blockers?
All-Hands / Town Hall
A company-wide meeting involving all employees. Usually led by the founders (C-Levels) to present results, new goals, and answer team questions.
1:1 (One-on-One)
A recurring individual meeting between a professional and their direct manager. It is used for career alignment, feedback, and problem-solving, not just for project status updates.

💰 Contracts, Benefits & Compensation

PTO (Paid Time Off)
Instead of strict, categorized leave policies, modern US companies usually offer a flexible pool of days (e.g., 20 days of PTO, or even "Unlimited PTO") that you can use for vacations, sick days, or personal matters, while receiving your regular compensation.
Equity / Stock Options
Company ownership. The startup offers you the right to buy shares at a heavily discounted strike price in the future. If the company grows, goes public, or is acquired, these shares can be highly lucrative.
Vesting (Vesting Schedule)
The rule that controls your Equity. It usually lasts 4 years. You don't get all the shares on day one; you "earn" them gradually as you stay with the company. A "1-year Cliff" means you must stay for at least one year to receive your first batch of shares.
RSUs (Restricted Stock Units)
Unlike Stock Options (where you have the right to buy the stock), RSUs are actual shares the company grants you as a bonus or part of your compensation package, following a strict Vesting schedule.
Independent Contractor (1099 / B2B)
The most common international hiring model for global talent working for US companies. You act as a service provider (business-to-business), receiving the gross salary (often six-figure compensation) without standard local payroll tax deductions at the source.

🤖 Recruitment & Hiring Process

ATS (Applicant Tracking System)
The "robot" that reads your resume. Software like Greenhouse, Ashby, and Workday are used to filter candidates by keywords before a human even looks at the document. (Pro tip: this is why your resume must be clean, semantic, and have the right keywords).
JD (Job Description)
The document that lists the responsibilities, technical requirements, and benefits of the open position.
Cultural Fit
The interview stage that evaluates if your core values, communication style, and worldview align with the company's culture. This is the ultimate test of your Soft Skills.
Onboarding
The integration process. It's the period of your first few weeks at the company, where you get your access credentials, learn about the culture, study the internal documentation, and understand how the product works.

🚀 How to use this to your advantage?

The secret isn't just knowing what these acronyms mean, but using them actively. If during an interview for a premium tech role you ask, "How does your PTO policy and Vesting schedule work?", the recruiter will immediately perceive you as a high-level professional, familiar with the global market standards.

The remote job market requires preparation. And having the right vocabulary is the first big step to securing six-figure proposals and standing out among thousands of applicants.

Expert Tip

The Current State of the Cloud & DevOps Market

By Mondywork | The ultimate guide for tech professionals looking to level up their careers, land global remote roles, and master modern infrastructure.

The End of the "SysAdmin" and the Rise of Platform Engineering

If you've been tracking the most sought-after roles at startups and Fortune 500 companies, you've probably noticed a drastic shift. The Cloud Computing and DevOps market is no longer just about "keeping the app online." Today, companies demand extreme resilience, automation, and cost optimization (the famous FinOps).

DevOps culture has evolved. We are now in the era of Platform Engineering, where infrastructure teams build internal platforms to grant developers full autonomy (Self-Service). For those aiming for global remote roles with six-figure compensation, understanding this transition isn't just a plus—it's a strict requirement.

"Elite DevOps teams deploy code 973 times more frequently and boast a Mean Time To Recovery (MTTR) that is 6,570 times faster than low-performing teams."

— State of DevOps Report (DORA / Google Cloud)

What Do Recruiters and ATS Systems Really Look For?

Automated filters (ATS platforms like Greenhouse, Ashby, and Workday) are hardwired to scan for a very specific set of skills in Cloud/DevOps resumes. Simply listing 20 tools won't cut it anymore; you must demonstrate culture and real-world impact.

  • Infrastructure as Code (IaC): Forget clicking around the AWS console. Declarative tools like Terraform, Pulumi, and Ansible are the true industry standards.
  • Container Orchestration: Mastery of Docker and, above all, Kubernetes (K8s) is practically mandatory for premium roles.
  • CI/CD Culture: Building automated pipelines (GitHub Actions, GitLab CI, ArgoCD) that guarantee continuous and secure deployments.
  • Cloud Providers and Architecture: Proficiency in at least one major public cloud provider (AWS, Google Cloud, or Microsoft Azure), with a strong focus on Serverless, microservices, and system resilience.

How to Transition: The Roadmap to Success

Many Back-End developers, support analysts, and network professionals want to pivot to Cloud and DevOps. If that's your goal, follow this market-validated roadmap:

  1. Master the Basics (Linux and Networking): Modern DevOps runs on Linux. Understanding permissions, processes, bash scripting, and the fundamentals of TCP/IP, DNS, and Load Balancers is your foundation.
  2. Learn to Code: A DevOps engineer doesn't need to be a Senior Developer, but you must know how to automate tasks. Python and Go (Golang) are currently the most highly valued languages in this ecosystem.
  3. Are Certifications Worth It? Yes, especially to get past the initial HR screening. Certifications like the AWS Certified Solutions Architect – Associate or the CKA (Certified Kubernetes Administrator) carry massive weight in the global market.
  4. Build a Real Portfolio: Publish a project on GitHub where you provision infrastructure via Terraform, set up a CI/CD pipeline using GitHub Actions, and deploy a simple app to a Kubernetes cluster. Practical projects always beat theoretical resumes.

Essential References and Study Links

To stay up-to-date, we highly recommend following official organizations and annual market reports:


🚀 Want access to the best remote Cloud & DevOps jobs?

Join our mailing list and get the best jobs delivered straight to you, tailored to your areas of interest. Currently, we partner with over 200 companies that advertise their premium, remote-first roles with us. Sign up now and get the best opportunities in the global market right in your inbox!