← Back to jobs

Software Engineer, Compute Foundations Systems

openai

San Francisco
Development

Job Score

80 pts
On-site model (+70) Development (+10)

About the Team

Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training.

Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running.

That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets.

About the Role

We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning.

You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate.

You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models.

In this role, you will:

  • Build and maintain the Linux host software stack for large GPU clusters, including Ubuntu and OS images, kernel configuration and modules, drivers, packages, disks and storage configuration, and machine configuration.

  • Design reproducible OS-image builds, package and repository workflows, and system configuration for heterogeneous bare-metal and cloud compute fleets.

  • Integrate, test, and qualify kernels, modules, drivers, packages, and firmware across new and existing hardware platforms; build safe canary, rollback, and recovery paths for system changes.

  • Bring up new hardware platforms and compute SKUs, working with hardware engineers and vendors to resolve firmware, driver, operating-system, and compatibility issues.

  • Debug complex system failures across boot and provisioning, firmware, disks, kernels and drivers, and workload interactions; turn recurring failure modes into durable fixes, tests, and automation.

  • Improve provisioning, repave, maintenance, and recovery correctness by eliminating manual host-by-host intervention and making system behavior predictable at fleet scale.

  • Build focused systems tooling and diagnostics that make the host software stack easier to validate, troubleshoot, and operate across the installed fleet.

You might thrive in this role if you:

  • Have significant experience building, integrating, or operating production Linux systems, particularly Ubuntu, Debian, or another major Linux distribution.

  • Have deep experience in one or more of: Linux kernels, modules, or device drivers; Linux distribution, package, repository, or OS-image engineering; boot, provisioning, disks, or machine configuration; firmware and driver integration, qualification, or rollout.

  • Can write, debug, and maintain production-quality systems software and automation using appropriate systems languages, scripting, and Linux tooling.

  • Understand how to build, test, package, qualify, and safely deliver system-level changes, including compatibility testing, staged rollout, rollback, and recovery.

  • Can systematically debug failures that cross hardware, firmware, boot, operating-system, kernel, and driver boundaries, and collaborate effectively with hardware and software partners.

  • Enjoy investigating difficult system behavior, identifying underlying failure modes, and building pragmatic fixes that make production systems more reliable.

Bonus points if you:

  • Have built bare-metal provisioning, machine-bootstrap, OS-reprovisioning, or system-lifecycle tooling.

  • Have worked with GPU systems, AI infrastructure, HPC clusters, accelerators, or other heterogeneous and performance-sensitive compute environments.

  • Have experience bringing up new server platforms or working directly with hardware and software vendors on firmware, driver, or operating-system compatibility issues.

  • Have contributed to Linux, kernel subsystems or modules, distributions, package ecosystems, drivers, firmware tooling, or other open-source systems software.

  • Have delivered or supported system-software changes across large production compute fleets.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

What did you think of this job?

Comments 0

Want to leave a comment?
Sign in or create your account in seconds to join the discussion.
Loading comments...

About Software Development

Software Development is one of the most dynamic and constantly evolving fields in the job market. Professionals in this area are responsible for creating, maintaining, and optimizing web, mobile, and desktop applications that impact millions of users daily.

Key languages and frameworks include JavaScript (React, Node.js, Vue.js), Python (Django, Flask), Java (Spring), PHP (Laravel), and TypeScript. Demand for full-stack developers continues to grow, especially in tech companies and startups.

Salaries range from entry-level to senior positions, with growing opportunities for remote work and international freelancing.

Discover Other Areas

Understand the scope of work, key skills, and tools used in different career areas.

About Scrum Master

The Scrum Master is the professional responsible for facilitating the adoption of Scrum and agile practices within development teams. They act as servant leaders, removing impediments, promoting continuous improvement, and ensuring Scrum events and ceremonies happen in the best possible way.

Key skills include event facilitation (sprint planning, daily, review, retrospective), backlog management, team coaching, conflict resolution, and agile metrics (velocity, burndown, cycle time). Knowledge of Jira, Trello, Azure DevOps, and frameworks like Kanban, XP, and SAFe is a differentiator.

Scrum Masters in technology companies are highly valued, especially those who can promote team autonomy, create psychologically safe environments, and lead agile transformations at scale. The field offers opportunities from junior scrum master to agile coach, head of agile, and director of agile transformation.

About Agile

The Agile and Digital Transformation area is fundamental for organizations seeking efficiency and rapid adaptation. Agile professionals facilitate processes, eliminate bottlenecks, and promote a culture of continuous improvement.

Key certifications include CSM, PSM, SAFe, ICP, and Kanban. Knowledge of Scrum, Kanban, XP, and agile frameworks is essential, as are leadership and facilitation soft skills.

Senior Agile coaches and Scrum Masters are highly valued, especially in technology companies that adopt agile methodologies at scale.

About Product Management

Product Management is one of the most strategically relevant areas in technology organizations. The Product Manager is responsible for defining product vision, prioritizing features, and coordinating multidisciplinary teams to deliver value to users.

Essential skills include strategic thinking, data analysis, communication, leadership, and technical knowledge. Tools like Jira, Confluence, Miro, and analytics platforms are fundamental in daily work.

Salaries for PMs range from entry-level to senior positions at major tech companies, with growing opportunities for international remote work.

About Project Management

Project Management is essential to ensure strategic initiatives are delivered on time, within scope, and with quality. PM professionals coordinate teams, manage risks, and communicate with stakeholders.

Key methodologies include PMBOK, PRINCE2, Scrum, and Kanban. Tools like Jira, Asana, Monday, and MS Project are widely used in daily work.

Certifications like PMP and PgMP are important differentiators in the market, with growing demand in technology and consulting companies.

About Frontend

The Frontend area is responsible for creating the visual interfaces that users interact with on websites and web applications. Frontend professionals combine technical skills with design to deliver intuitive, responsive, and accessible digital experiences.

Key skills include HTML, CSS, JavaScript/TypeScript, frameworks like React, Angular, and Vue, build tools (Webpack, Vite), CSS (Tailwind, Sass), testing (Jest, Cypress), and knowledge of web performance and accessibility (WCAG). Familiarity with design systems and reusable components is a differentiator.

Frontend developers in technology companies are highly valued, especially those who master React, Next.js, web performance, and accessibility. The field offers opportunities from junior developer to frontend architect, with a focus on user experience, performance, and code quality.

Career Guides

Technology Career Guide

Planning, skills, interviews, and professional growth in IT, Data Science, DevOps, and Product.

Read full guide →

Design Career Guide

UX/UI, Graphic Design, Product Design. Portfolio, tools, interviews, and growth in the Design field.

Read full guide →

Marketing Career Guide

SEO, Paid Media, Growth, Content Marketing. Certifications, tools, and strategies to grow in Digital Marketing.

Read full guide →

Finance Career Guide

Financial market, investments, corporate finance, certifications, and strategies to grow in the financial field.

Read full guide →

Communication Career Guide

Journalism, PR, Corporate Communication, Content Marketing, and Multimedia Production.

Read full guide →

Administration Career Guide

Business Management, HR, Logistics, Consulting, Project Management, and Entrepreneurship.

Read full guide →

Data Career Guide

Data Science, Data Engineering, BI, Machine Learning, and AI. From training to the job market.

Read full guide →

Product Career Guide

Product Management, Product Ownership, Agile, Scrum, and OKRs. From strategy to execution.

Read full guide →

Tech & Remote Glossary

Stop getting lost in interviews and job descriptions

The job market, especially within tech and global companies, has developed its own dialect. Not understanding these acronyms can make you lose valuable opportunities or poorly negotiate your contract. To end this problem, we created the Definitive Glossary for the Remote Professional.

🏢 Work Models & Routine

Async (Asynchronous Work)
A communication model where responses don't need to be immediate. Instead of back-to-back meetings, the team relies on well-structured documents, threads, and messages. It's the gold standard for global companies spanning multiple time zones.
Sync (Synchronous Work)
The opposite of Async. It requires the team to be online and available at the same time for meetings, live chats, and real-time collaboration.
Daily / Stand-up
A quick daily meeting (usually 15 minutes) common in Agile (Scrum) methodologies. The team answers three questions: What did I do yesterday? What will I do today? Are there any blockers?
All-Hands / Town Hall
A company-wide meeting involving all employees. Usually led by the founders (C-Levels) to present results, new goals, and answer team questions.
1:1 (One-on-One)
A recurring individual meeting between a professional and their direct manager. It is used for career alignment, feedback, and problem-solving, not just for project status updates.

💰 Contracts, Benefits & Compensation

PTO (Paid Time Off)
Instead of strict, categorized leave policies, modern US companies usually offer a flexible pool of days (e.g., 20 days of PTO, or even "Unlimited PTO") that you can use for vacations, sick days, or personal matters, while receiving your regular compensation.
Equity / Stock Options
Company ownership. The startup offers you the right to buy shares at a heavily discounted strike price in the future. If the company grows, goes public, or is acquired, these shares can be highly lucrative.
Vesting (Vesting Schedule)
The rule that controls your Equity. It usually lasts 4 years. You don't get all the shares on day one; you "earn" them gradually as you stay with the company. A "1-year Cliff" means you must stay for at least one year to receive your first batch of shares.
RSUs (Restricted Stock Units)
Unlike Stock Options (where you have the right to buy the stock), RSUs are actual shares the company grants you as a bonus or part of your compensation package, following a strict Vesting schedule.
Independent Contractor (1099 / B2B)
The most common international hiring model for global talent working for US companies. You act as a service provider (business-to-business), receiving the gross salary (often six-figure compensation) without standard local payroll tax deductions at the source.

🤖 Recruitment & Hiring Process

ATS (Applicant Tracking System)
The "robot" that reads your resume. Software like Greenhouse, Ashby, and Workday are used to filter candidates by keywords before a human even looks at the document. (Pro tip: this is why your resume must be clean, semantic, and have the right keywords).
JD (Job Description)
The document that lists the responsibilities, technical requirements, and benefits of the open position.
Cultural Fit
The interview stage that evaluates if your core values, communication style, and worldview align with the company's culture. This is the ultimate test of your Soft Skills.
Onboarding
The integration process. It's the period of your first few weeks at the company, where you get your access credentials, learn about the culture, study the internal documentation, and understand how the product works.

🚀 How to use this to your advantage?

The secret isn't just knowing what these acronyms mean, but using them actively. If during an interview for a premium tech role you ask, "How does your PTO policy and Vesting schedule work?", the recruiter will immediately perceive you as a high-level professional, familiar with the global market standards.

The remote job market requires preparation. And having the right vocabulary is the first big step to securing six-figure proposals and standing out among thousands of applicants.

Expert Tip

The 2026 AI Boom: The Most Valuable Tech Careers and How to Land Six-Figure Remote Jobs

We are halfway through 2026, and one thing is crystal clear: the "experimental" phase of Artificial Intelligence is officially over. While 2023 and 2024 were characterized by awe over chatbots drafting emails and generating images, 2026 has solidified AI as the core infrastructure of global enterprises. The transition from standalone "AI tools" to Autonomous Agents and Multi-Agent Systems has radically transformed the job market.

For Tech, Design, and Digital Marketing professionals across the United States, 2026 represents the greatest window of opportunity of the decade to secure top-tier, 100% remote roles with highly lucrative six-figure compensations.

In this article, we will break down the current AI job landscape, backed by recent data, and list the top careers that startups and Fortune 500 companies are desperately trying to fill.

The Current Landscape: 2026 Data and Projections

The market isn't just hiring standard developers anymore; it's hiring intelligence orchestrators. According to recent Future of Work reports:

  • Exponential Growth: The World Economic Forum (WEF) 2026 update highlights that roles focused on AI, Machine Learning, and Big Data have grown by 45% compared to 2024, cementing them as the fastest-growing fields nationwide.
  • Corporate Adoption: Data published by Gartner earlier this year reveals that over 80% of Fortune 500 companies are now running Generative AI applications in production environments. This has created a massive demand for AI maintenance, ethics, and governance.
  • The Remote Premium: An internal analysis from Mondywork's database (which tracks integrations with major ATS platforms like Greenhouse and Ashby) shows that 73% of US-based AI roles are Remote-First. The average salary for senior specialists in these roles currently exceeds the $140,000 to $180,000 annual range, plus equity.

The 5 Hottest AI Opportunities in 2026

If you want to tailor your resume and LinkedIn profile to be easily captured by modern recruiting algorithms, these are the positions with the highest talent deficit in the US market right now:

1. MLOps and LLMOps Engineers (Operations Engineering)

Large Language Models (LLMs) are like Formula 1 engines: they need a full pit crew to avoid crashing on the track. The industry has realized that putting AI into production is vastly different from running a local model.

  • What they do: Manage infrastructure, oversee the model lifecycle, handle fine-tuning with proprietary company data, and ensure the AI does not suffer from large-scale hallucinations.
  • Hot Search Terms: MLOps, LLMOps, Platform Engineering, Data Ops, Kubernetes for AI.

2. Prompt Engineer & AI Interaction Designer

The profession many thought would be a passing fad has heavily evolved. The 2026 Prompt Engineer is not just someone who "talks well to machines"; they are complex logical system designers.

  • What they do: Sitting at the intersection of Software Engineering and UX Design, these professionals design system prompts for Autonomous Agents, build RAG (Retrieval-Augmented Generation) flows, and structure how AI safely interacts with end-users.
  • Hot Search Terms: Prompt Engineering, NLP, AI Behavior Design, UX Writer for AI.

3. Analytics Engineer / Structured Data Specialist

AI is completely useless without clean data. The classic Data Scientist role has yielded massive ground to the Analytics Engineer, the professional who bridges the gap between raw data engineering and business analysis.

  • What they do: Prepare, model, and transform chaotic data lakes into crystal-clear sources so enterprise AI models can consume data and generate real-time insights.
  • Hot Search Terms: Analytics Engineer, dbt, Snowflake, Computer Vision, BigQuery.

4. AI Product Manager (AI PM)

Companies are tired of building AI features "just because." Now, they need these features to drive serious revenue (ROI). The AI-focused Product Manager is the conductor of this orchestra.

  • What they do: Understand the technical limitations of modern LLMs, translate user pain points into viable AI solutions, and manage the product roadmap while ensuring the technology complies with strict privacy laws (like CCPA and GDPR).
  • Hot Search Terms: AI Product Manager, CPO, Product Ops, AI Governance.

5. AI Growth Marketer / High-Performance Media Buyer

In the digital marketing realm, 2026 is the year of autonomous campaign orchestration. Marketers still relying on 100% manual campaign creation are rapidly losing ground to those who can direct predictive AI.

  • What they do: Leverage Machine Learning and advanced AI tools for autonomous Conversion Rate Optimization (CRO), automated A/B testing, mass content generation, and predictive consumer behavior analysis.
  • Hot Search Terms: Growth Marketing, Media Buyer, Programmatic, AI Copywriting, Martech.

How to Prepare and Get Found (Beating the ATS Filters)

US companies utilize incredibly rigorous Applicant Tracking Systems (ATS) like Workday, Greenhouse, and Lever. They configure recruiting bots to filter resumes using fine-mesh keyword grids.

If you want to land these highly competitive roles, the golden rule is to mirror the exact industry jargon:

  • Don't just write "Data Analyst"; use "Data Ops" or "Analytics Engineer".
  • Don't just list "Cloud Support"; highlight "FinOps", "Cloud Architect", or "Platform Engineer".
  • Replace the outdated "Digital Marketer" with "Growth Ops" or "Performance Manager".

Mondywork Does the Heavy Lifting for You

The US market is fiercely competing for top-tier talent. Startups and tech giants are looking for highly skilled professionals ready to collaborate across different time zones in fully remote environments. That is exactly why Mondywork exists. Our proprietary algorithm scans the largest global Job Boards to find verified, high-paying, and 100% remote Tech, Design, and Marketing opportunities.

Don't miss the chance to ride the biggest technological revolution of our generation.

👉 Subscribe now to Mondywork's Job Alerts


Macroeconomic Reference Sources:

  • World Economic Forum - The Future of Jobs Report 2026 Update.
  • Gartner - Hype Cycle for Artificial Intelligence, 2026.
  • McKinsey Global Institute - The Economic Potential of Generative AI (Revisited 2026).