Site Reliability Engineering Manager
onebrief
Job Score
100 ptsConsequential Work. Dedicated People.
About Onebrief
Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.
Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.
We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.
Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.
Security Clearance, Location, and Onsite Notice
This role requires regularly working on-site at customer locations.
If you are not currently within commuting distance, you must be willing to relocate. Onebrief provides relocation assistance.
Active Secret clearance required.
About The Role
We're hiring a Site Reliability Engineering Manager to lead our SRE team within Infrastructure & Security. You'll work closely with platform engineering, application engineering, security, and customer success to ensure Onebrief's mission-critical deployments are reliable, secure, and well supported across on-prem DoD and AWS environments.
You'll lead a team whose work spans customer-facing operations, infrastructure, observability, automation, and application reliability. Most of the team focuses on deploying and operating Onebrief in demanding customer environments.
You'll own the team's priorities, planning, execution, and development. A significant part of this role is coordinating work: understanding demand, balancing capacity, sequencing tasks, managing dependencies, and keeping commitments realistic as customer needs change. You'll help the team deliver immediate operational support while making steady progress on improvements that reduce future support demands.
You'll bring the technical grounding to evaluate risks, ask useful questions, and guide decisions. Your engineers will own technical implementation and lead incident response. You'll provide direction, remove blockers, and create the conditions for them to succeed.
About You
You care deeply about reliability and understand the challenges of operating software in environments where connectivity, access, and deployment options can be constrained. You treat infrastructure and operability as products that deserve clear ownership, thoughtful design, and continuous improvement.
You're an effective people manager who sets clear expectations, gives useful feedback, and helps engineers grow. You build accountability through clear priorities and meaningful ownership, and you recognize when your team needs direction, support, or room to solve a problem.
You're comfortable managing a changing workload. You can turn competing requests into an actionable plan, account for operational interruptions, and explain what the team can commit to with its available capacity. You surface tradeoffs early and work with stakeholders to make deliberate decisions about scope and timing.
You bring calm and structure when priorities shift or incidents occur. You support engineers leading the response, help resolve escalations, and coordinate with customer-facing partners. You build a culture where engineers can surface risks early and examine failures honestly.
You have the technical judgment to help the team determine whether a recurring problem needs an infrastructure change, better automation, an application fix, or a clearer process. You bring the right people together to address it and ensure they have time to follow through.
What You'll Do
Lead and develop the SRE team: Hire, coach, and support engineers across infrastructure, operations, and application reliability. Set expectations, manage performance, support career development, and build the skills and coverage the team needs.
Own capacity and work planning: Maintain a clear view of incoming requests, ongoing support needs, and planned engineering work. Break initiatives into manageable tasks with the team, establish ownership, sequence work, and adjust commitments as priorities or capacity change.
Coordinate delivery across teams: Manage dependencies with platform engineering, application engineering, security, and customer success. Identify blockers early, resolve competing priorities, and communicate progress, risks, and decisions to stakeholders.
Set the reliability roadmap: Translate customer needs, production data, incident patterns, and operational risks into a prioritized improvement plan. Protect capacity for work that reduces recurring failures and makes deployments easier to operate.
Establish operational ownership: Ensure production deployments have clear support responsibilities, escalation paths, and readiness criteria. Plan with partner teams for new deployments, releases, and ongoing customer support.
Support team-led incident response: Establish sustainable on-call coverage and clear incident response expectations. Coach engineers who serve as incident commanders and lead blameless postmortems / After Action Reviews (AARs). Help the team assess corrective actions, assign ownership, and schedule follow-through.
Guide technical priorities: Work with engineers and technical leads to evaluate approaches to infrastructure, automation, observability, and application reliability. Ensure plans account for operability, security requirements, and the constraints of on-prem and air-gapped environments.
Make reliability and workload visible: Guide the team's use of SLIs, SLOs, and operational metrics. Use service health, support demand, and delivery progress to explain where investment is needed and whether improvements are working.
Reduce operational toil: Give engineers time and support to automate repetitive deployment, maintenance, troubleshooting, and recovery work. Help turn lessons from individual customer environments into reusable improvements.
What We Look For
An active Secret clearance
5+ years in Site Reliability Engineering, Platform Engineering, DevOps, or a related role, with substantial infrastructure and operations experience
Experience directly managing engineers, including coaching, performance management, career development, and hiring
Experience planning team capacity, prioritizing competing requests, and coordinating engineering work across teams
A track record of delivering reliability improvements while managing ongoing operational responsibilities
Experience with incident response and post-incident review practices, including helping engineers develop leadership and ownership
Technical judgment sufficient to evaluate engineering proposals, understand operational risks, and guide prioritization
Clear communication, including the ability to explain constraints and negotiate scope, timing, and commitments with stakeholders
Technical background:
You should have practical experience operating production systems and enough breadth to guide engineers working across these areas:
Infrastructure and automation: Infrastructure as Code, configuration management, and scripting, using tools such as Terraform, Ansible, Python, Go, or Bash
Containers and orchestration: Kubernetes deployment, troubleshooting, and operations
Delivery practices: CI/CD pipelines, release safety, and repeatable deployments
Cloud and on-prem environments: Operating software across AWS or AWS GovCloud and customer-managed infrastructure
Observability: Monitoring, logging, and actionable alerting using tools such as the Grafana stack, ELK, or Datadog
Networking and security: Core protocols, secure configuration, and connectivity troubleshooting
Application reliability: Working with software engineers to diagnose application failures and evaluate infrastructure or code changes
We value depth in relevant areas and the ability to guide specialists across the rest.
Bonus points (nice to have):
Experience leading teams supporting mission-critical customer deployments
Experience managing engineering programs or coordinating delivery across multiple customer environments
Experience operating in DoD, classified, or air-gapped environments
Familiarity with RMF, STIGs, and ICD 503
Experience implementing SLIs, SLOs, and error budgets for distributed systems
GitOps practices and toolchains
On-prem virtualization experience with VMware, Proxmox, Nutanix, Hyper-V, or similar platforms
Application development experience, especially TypeScript or Node.js
Relevant certifications, such as AWS DevOps Engineer or CKA/CKAD
Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain one within 3 months of employment
Notice to Third Party Recruitment Agencies
Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.
About Engineering
Software Engineering goes beyond traditional development, focusing on scalability, performance, and system architecture. Software engineers are responsible for designing infrastructures that support millions of simultaneous users.
Skills include microservices architecture, DevOps, cloud computing, application security, and performance optimization. Knowledge of containerization (Docker, Kubernetes) and CI/CD is increasingly required.
Senior software engineers are rare and highly compensated professionals, with opportunities at major global tech companies.
About Web Master
The Web Master is the professional responsible for maintaining, securing, and ensuring the technical performance of websites and web applications. They manage servers, hosting infrastructure, uptime monitoring, and ensure everything runs fast and reliably.
Key skills include server management (Apache, Nginx), hosting (AWS, Google Cloud, Azure), CDN (Cloudflare), SSL, DNS, web security (WAF, firewall), performance (Core Web Vitals, cache, compression), and versioning (Git, CI/CD). Knowledge of Docker, WordPress, cPanel, and monitoring (Sentry, New Relic) is a differentiator.
Web Masters in technology companies are highly valued, especially those who master DevOps, SRE, and can guarantee uptime and performance at scale. The field offers opportunities from junior webmaster to SRE and infrastructure engineer, with a focus on reliability, security, and speed.
Discover Other Areas
Understand the scope of work, key skills, and tools used in different career areas.
About Cloud Solutions
The Cloud Solutions area is responsible for designing, implementing, and managing cloud infrastructure and services (AWS, Azure, GCP) for companies. Cloud professionals architect scalable, secure, and cost-optimized solutions, from data center migrations to serverless and multi-cloud architectures.
Key skills include IaC (Terraform, CloudFormation), containers (Docker, Kubernetes), serverless (Lambda, Cloud Functions), managed databases (RDS, DynamoDB, BigQuery), cloud networking (VPC, CDN, load balancer), and security (IAM, WAF, KMS). Knowledge of FinOps, cloud governance, and AWS/Azure/GCP certifications is a differentiator.
Cloud Solutions professionals in technology companies are highly valued, especially those who master multi-cloud architectures, FinOps, and can optimize costs while maintaining performance and security. The field offers opportunities from cloud engineer to cloud solutions architect, head of cloud, and chief cloud architect.
About Project Manager
The Project Manager is the professional responsible for planning, executing, and controlling projects end-to-end, ensuring they are delivered on time, within budget, and with the expected quality. With the growing complexity of businesses, project management professionals are fundamental to organizational success.
Key skills include planning and scheduling, scope, cost, risk, quality, and resource management, stakeholder communication, cross-functional team leadership, and use of agile and traditional methodologies. Certifications like PMP, PRINCE2, and Six Sigma are important differentiators.
Project Managers in technology companies are highly valued, especially those who master agile methodologies (Scrum, Kanban), tools like Jira and MS Project, and can deliver complex projects efficiently. The field offers opportunities from project analyst to head of PMO, with a focus on execution, governance, and business value.
About Product Management
Product Management is one of the most strategically relevant areas in technology organizations. The Product Manager is responsible for defining product vision, prioritizing features, and coordinating multidisciplinary teams to deliver value to users.
Essential skills include strategic thinking, data analysis, communication, leadership, and technical knowledge. Tools like Jira, Confluence, Miro, and analytics platforms are fundamental in daily work.
Salaries for PMs range from entry-level to senior positions at major tech companies, with growing opportunities for international remote work.
About Customer Success
Customer Success is the area responsible for ensuring clients achieve their goals when using the product or service. It is a strategic function for retention, expansion, and customer satisfaction.
Key skills include account management, churn analysis, NPS, onboarding, upsell, and cross-sell. Knowledge of CS tools like Gainsight, Totango, and ChurnZero is a differentiator.
CS is becoming increasingly strategic in SaaS companies, with professionals directly contributing to recurring revenue growth (MRR/ARR).
About Traffic Manager
The Traffic Manager is the professional responsible for planning, executing, and optimizing paid media campaigns across various digital platforms. With the competitiveness of the digital market, paid traffic professionals are essential for generating qualified leads and maximizing return on advertising investment.
Key skills include campaign management on Google Ads, Meta Ads, LinkedIn Ads, and TikTok Ads, media planning, metrics analysis (ROAS, CPA, CPC, CTR), A/B testing, remarketing, and landing page creation. Tools like Google Analytics, Google Tag Manager, Hotjar, and automation platforms are essential.
Traffic managers in technology companies are highly valued, especially those who master performance marketing, conversion funnel optimization, and scaling strategies. The field offers opportunities from media analyst to head of performance, with a focus on growth, budget efficiency, and return on investment.
Comments 0