Researcher, Agent Safety, Oversight And System Mitigations
openai
Job Score
80 ptsAbout the Team
The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously.
Our work spans three areas:
Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks.
Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work.
Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful agent autonomy (for example future versions of auto-review).
About the Role
This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. We prioritize building oversight systems that are used in practice today, both internally and externally (see our recent work on action monitoring for codex and former code review). We also study longer-term questions about how increasingly capable agentis systems can be supervised, constrained, and corrected.
We’re looking for a safety&security minded researcher or engineer who can reason rigorously about security boundaries and agent behavior, then build and test practical mitigations. A background in AI control or security is welcome but not required.
This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.
In this role, you will:
Design, build, and evaluate system-level controls for agent actions like agent-based review. Plan how they fit in a broader system including sandboxing with process isolation and permission boundaries.
Work closely with a Codex harness engineering team to productionize the AI controls.
Red-team end-to-end agentic systems to measure whether controls prevent data exfiltration, unsafe tool use, and other harmful outcomes.
Improve the safety–productivity tradeoff by measuring and reducing missed harmful actions, unnecessary blocks, approval burden, and latency.
You might thrive in this role if you:
Have strong systems or security instincts and can reason concretely about isolation boundaries, permissions, attack surfaces, and failure modes in complex systems.
Enjoy turning ambiguous safety questions into concrete threat models, reproducible experiments, and practical mitigations, and revising your approach based on evidence from deployment.
Can build robust experimental infrastructure and design evaluations that distinguish promising mitigations from brittle ones.
Are deeply interested in frontier AI alignment, safety and control.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
OpenAI Global Applicant Privacy Policy
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
About Scrum Master
The Scrum Master is the professional responsible for facilitating the adoption of Scrum and agile practices within development teams. They act as servant leaders, removing impediments, promoting continuous improvement, and ensuring Scrum events and ceremonies happen in the best possible way.
Key skills include event facilitation (sprint planning, daily, review, retrospective), backlog management, team coaching, conflict resolution, and agile metrics (velocity, burndown, cycle time). Knowledge of Jira, Trello, Azure DevOps, and frameworks like Kanban, XP, and SAFe is a differentiator.
Scrum Masters in technology companies are highly valued, especially those who can promote team autonomy, create psychologically safe environments, and lead agile transformations at scale. The field offers opportunities from junior scrum master to agile coach, head of agile, and director of agile transformation.
Discover Other Areas
Understand the scope of work, key skills, and tools used in different career areas.
About Project Management
Project Management is essential to ensure strategic initiatives are delivered on time, within scope, and with quality. PM professionals coordinate teams, manage risks, and communicate with stakeholders.
Key methodologies include PMBOK, PRINCE2, Scrum, and Kanban. Tools like Jira, Asana, Monday, and MS Project are widely used in daily work.
Certifications like PMP and PgMP are important differentiators in the market, with growing demand in technology and consulting companies.
About Blockchain
The Blockchain area involves the development and implementation of secure and distributed transaction ledgers. Professionals in this field work with smart contract development, cryptography, consensus algorithms, and platforms such as Ethereum, Hyperledger, and Solana, ensuring security and decentralization for various types of applications.
About People Analyst
The People Analyst is the professional responsible for transforming people data into strategic insights for HR decision-making. They combine data analysis knowledge with people management vision to help organizations understand workforce metrics, turnover, engagement, and diversity.
Key skills include people analytics, workforce analytics, turnover and retention analysis, HR metrics (time-to-hire, cost-per-hire, e-NPS), data visualization (Power BI, Tableau, Visier), workforce planning, and compensation analysis. Knowledge of statistics, SQL, and people analytics tools is a differentiator.
People Analysts in technology companies are highly valued, especially those who can translate complex people data into actionable insights for retention, diversity, and growth strategies. The field offers opportunities from HR analyst to head of people analytics, with a focus on data-driven people management.
About Ecommerce Manager
The Ecommerce Manager is the professional responsible for the entire strategic and operational management of online stores and marketplaces. They lead teams, define pricing, promotion, and catalog strategies, and monitor online sales performance across multiple platforms.
Key skills include catalog management, dynamic pricing, seasonal campaigns (Black Friday, Cyber Monday), marketplace management (Amazon, Mercado Livre, Shopee, Magalu), paid traffic, CRO, and team management. Knowledge of Shopify, VTEX, WooCommerce, Google Ads, Meta Ads, and performance metrics is a differentiator.
Ecommerce Managers in technology companies are highly valued, especially those who master multi-marketplace management, checkout optimization, and mobile commerce strategies. The field offers opportunities from ecommerce manager to head of ecommerce, with a focus on revenue, customer experience, and growth.
About IT Governance
IT Governance is the area responsible for ensuring that information technology resources are used strategically, efficiently, and in compliance with standards and regulations. IT governance professionals ensure that technology supports business objectives in a secure and reliable manner.
Key skills include IT service management (ITIL), IT audit and compliance, risk management, business continuity, disaster recovery, metrics and indicators (SLAs, KPIs), and strategic alignment between IT and business. Frameworks like COBIT, ITIL, ISO 27001, and compliance standards are essential.
IT Governance professionals in technology companies are highly valued, especially those who master ITSM, IT audit, and risk management. The field offers opportunities from governance analyst to CIO/CTO, with a focus on efficiency, compliance, security, and business value.
Comments 0