← Back to jobs

Site Reliability Engineering Manager

onebrief

Híbrido Colorado Springs, CO
Engineering Uncategorized Web Master

Job Score

100 pts
Hybrid model (+80) Engineering (+10) Web Master (+10)

Consequential Work. Dedicated People.


About Onebrief

Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.

Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.

We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.

Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.

Security Clearance, Location, and Onsite Notice

This role requires regularly working on-site at customer locations.

If you are not currently within commuting distance, you must be willing to relocate. Onebrief provides relocation assistance.

Active Secret clearance required.

About The Role

We're hiring a Site Reliability Engineering Manager to lead our SRE team within Infrastructure & Security. You'll work closely with platform engineering, application engineering, security, and customer success to ensure Onebrief's mission-critical deployments are reliable, secure, and well supported across on-prem DoD and AWS environments.

You'll lead a team whose work spans customer-facing operations, infrastructure, observability, automation, and application reliability. Most of the team focuses on deploying and operating Onebrief in demanding customer environments.

You'll own the team's priorities, planning, execution, and development. A significant part of this role is coordinating work: understanding demand, balancing capacity, sequencing tasks, managing dependencies, and keeping commitments realistic as customer needs change. You'll help the team deliver immediate operational support while making steady progress on improvements that reduce future support demands.

You'll bring the technical grounding to evaluate risks, ask useful questions, and guide decisions. Your engineers will own technical implementation and lead incident response. You'll provide direction, remove blockers, and create the conditions for them to succeed.

About You

You care deeply about reliability and understand the challenges of operating software in environments where connectivity, access, and deployment options can be constrained. You treat infrastructure and operability as products that deserve clear ownership, thoughtful design, and continuous improvement.

You're an effective people manager who sets clear expectations, gives useful feedback, and helps engineers grow. You build accountability through clear priorities and meaningful ownership, and you recognize when your team needs direction, support, or room to solve a problem.

You're comfortable managing a changing workload. You can turn competing requests into an actionable plan, account for operational interruptions, and explain what the team can commit to with its available capacity. You surface tradeoffs early and work with stakeholders to make deliberate decisions about scope and timing.

You bring calm and structure when priorities shift or incidents occur. You support engineers leading the response, help resolve escalations, and coordinate with customer-facing partners. You build a culture where engineers can surface risks early and examine failures honestly.

You have the technical judgment to help the team determine whether a recurring problem needs an infrastructure change, better automation, an application fix, or a clearer process. You bring the right people together to address it and ensure they have time to follow through.

What You'll Do

  • Lead and develop the SRE team: Hire, coach, and support engineers across infrastructure, operations, and application reliability. Set expectations, manage performance, support career development, and build the skills and coverage the team needs.

  • Own capacity and work planning: Maintain a clear view of incoming requests, ongoing support needs, and planned engineering work. Break initiatives into manageable tasks with the team, establish ownership, sequence work, and adjust commitments as priorities or capacity change.

  • Coordinate delivery across teams: Manage dependencies with platform engineering, application engineering, security, and customer success. Identify blockers early, resolve competing priorities, and communicate progress, risks, and decisions to stakeholders.

  • Set the reliability roadmap: Translate customer needs, production data, incident patterns, and operational risks into a prioritized improvement plan. Protect capacity for work that reduces recurring failures and makes deployments easier to operate.

  • Establish operational ownership: Ensure production deployments have clear support responsibilities, escalation paths, and readiness criteria. Plan with partner teams for new deployments, releases, and ongoing customer support.

  • Support team-led incident response: Establish sustainable on-call coverage and clear incident response expectations. Coach engineers who serve as incident commanders and lead blameless postmortems / After Action Reviews (AARs). Help the team assess corrective actions, assign ownership, and schedule follow-through.

  • Guide technical priorities: Work with engineers and technical leads to evaluate approaches to infrastructure, automation, observability, and application reliability. Ensure plans account for operability, security requirements, and the constraints of on-prem and air-gapped environments.

  • Make reliability and workload visible: Guide the team's use of SLIs, SLOs, and operational metrics. Use service health, support demand, and delivery progress to explain where investment is needed and whether improvements are working.

  • Reduce operational toil: Give engineers time and support to automate repetitive deployment, maintenance, troubleshooting, and recovery work. Help turn lessons from individual customer environments into reusable improvements.

What We Look For

  • An active Secret clearance

  • 5+ years in Site Reliability Engineering, Platform Engineering, DevOps, or a related role, with substantial infrastructure and operations experience

  • Experience directly managing engineers, including coaching, performance management, career development, and hiring

  • Experience planning team capacity, prioritizing competing requests, and coordinating engineering work across teams

  • A track record of delivering reliability improvements while managing ongoing operational responsibilities

  • Experience with incident response and post-incident review practices, including helping engineers develop leadership and ownership

  • Technical judgment sufficient to evaluate engineering proposals, understand operational risks, and guide prioritization

  • Clear communication, including the ability to explain constraints and negotiate scope, timing, and commitments with stakeholders

Technical background:

You should have practical experience operating production systems and enough breadth to guide engineers working across these areas:

  • Infrastructure and automation: Infrastructure as Code, configuration management, and scripting, using tools such as Terraform, Ansible, Python, Go, or Bash

  • Containers and orchestration: Kubernetes deployment, troubleshooting, and operations

  • Delivery practices: CI/CD pipelines, release safety, and repeatable deployments

  • Cloud and on-prem environments: Operating software across AWS or AWS GovCloud and customer-managed infrastructure

  • Observability: Monitoring, logging, and actionable alerting using tools such as the Grafana stack, ELK, or Datadog

  • Networking and security: Core protocols, secure configuration, and connectivity troubleshooting

  • Application reliability: Working with software engineers to diagnose application failures and evaluate infrastructure or code changes

We value depth in relevant areas and the ability to guide specialists across the rest.

Bonus points (nice to have):

  • Experience leading teams supporting mission-critical customer deployments

  • Experience managing engineering programs or coordinating delivery across multiple customer environments

  • Experience operating in DoD, classified, or air-gapped environments

  • Familiarity with RMF, STIGs, and ICD 503

  • Experience implementing SLIs, SLOs, and error budgets for distributed systems

  • GitOps practices and toolchains

  • On-prem virtualization experience with VMware, Proxmox, Nutanix, Hyper-V, or similar platforms

  • Application development experience, especially TypeScript or Node.js

  • Relevant certifications, such as AWS DevOps Engineer or CKA/CKAD

  • Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain one within 3 months of employment


Notice to Third Party Recruitment Agencies

Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.

What did you think of this job?

Comments 0

Want to leave a comment?
Sign in or create your account in seconds to join the discussion.
Loading comments...

About Engineering

Software Engineering goes beyond traditional development, focusing on scalability, performance, and system architecture. Software engineers are responsible for designing infrastructures that support millions of simultaneous users.

Skills include microservices architecture, DevOps, cloud computing, application security, and performance optimization. Knowledge of containerization (Docker, Kubernetes) and CI/CD is increasingly required.

Senior software engineers are rare and highly compensated professionals, with opportunities at major global tech companies.

About Web Master

The Web Master is the professional responsible for maintaining, securing, and ensuring the technical performance of websites and web applications. They manage servers, hosting infrastructure, uptime monitoring, and ensure everything runs fast and reliably.

Key skills include server management (Apache, Nginx), hosting (AWS, Google Cloud, Azure), CDN (Cloudflare), SSL, DNS, web security (WAF, firewall), performance (Core Web Vitals, cache, compression), and versioning (Git, CI/CD). Knowledge of Docker, WordPress, cPanel, and monitoring (Sentry, New Relic) is a differentiator.

Web Masters in technology companies are highly valued, especially those who master DevOps, SRE, and can guarantee uptime and performance at scale. The field offers opportunities from junior webmaster to SRE and infrastructure engineer, with a focus on reliability, security, and speed.

Discover Other Areas

Understand the scope of work, key skills, and tools used in different career areas.

About Systems Analyst

The Systems Analyst is the professional responsible for analyzing, designing, and implementing technology solutions that meet business needs. They act as a bridge between business areas and the development team, ensuring that systems deliver real value to the organization.

Key skills include requirements gathering and analysis, process modeling (BPMN), data modeling, technical and functional documentation, system integration (APIs, microservices), and knowledge of ERPs and CRMs. Tools like Jira, Confluence, Visio, and project management platforms are essential.

Systems Analysts in technology companies are highly valued, especially those who master agile requirements analysis (user stories, backlog), system integration, and solution architecture. The field offers opportunities from junior analyst to solution architect, with a focus on efficiency, quality, and technological innovation.

About Graphic Designer

The Graphic Designer is the professional responsible for creating visual pieces for print and digital communication, from visual identity and logos to marketing materials and packaging. They combine creativity with technique to convey messages visually and impactfully.

Key skills include Adobe Photoshop, Illustrator, and InDesign, CorelDRAW, visual identity design, typography, color theory, packaging design, and motion graphics. Knowledge of vector illustration, offset/digital printing, and print production is a differentiator.

Graphic Designers in technology companies are highly valued, especially those who master social media design, infographics, and can create materials that strengthen brand visual identity. The field offers opportunities from junior graphic designer to art director and design director.

About Mobile Development

Mobile Development is one of the most dynamic and constantly evolving fields in the technology market. With billions of smartphones worldwide, the demand for qualified mobile developers continues to grow exponentially.

Key stacks include Flutter (Dart), React Native (JavaScript/TypeScript), Kotlin (Android native), Swift (iOS native), and hybrid frameworks like Ionic and Capacitor. Knowledge of mobile architecture (MVVM, Clean Architecture), mobile CI/CD (Fastlane, Bitrise, Codemagic), and App Store/Google Play publishing are essential.

Senior mobile developers are highly valued professionals, with competitive salaries and many remote work opportunities at international companies. Specializing in cross-platform or native is a strategic career decision.

About Tech Recruiter

The Tech Recruiter is a professional specialized in recruiting technology talent, from developers to AI engineers and DevOps professionals. They combine technical knowledge with recruitment skills to evaluate and attract highly qualified candidates.

Key skills include technical screening, analysis of technical profiles (GitHub, portfolios, blogs), knowledge of software stacks and architectures, networking in tech communities and events. Proficiency with tools like LinkedIn Recruiter, Gem, Ashby, and technical assessment platforms is a differentiator.

Tech Recruiters are scarce and highly paid professionals, especially those who can map and access passive talent in competitive markets like AI, data engineering, and cloud computing.

About Sales

The Sales area is responsible for generating revenue and expanding the customer base. B2B and B2C sales professionals are fundamental for sustainable growth of any organization.

Key skills include prospecting, negotiation, CRM (Salesforce, HubSpot), sales enablement, and value consulting. The consultative and data-driven approach is increasingly valued.

Consultative sellers and senior Sales Managers have very high earning potential, with OTE (On-Target Earnings) that can exceed monthly salaries in technology companies.

Career Guides

Technology Career Guide

Planning, skills, interviews, and professional growth in IT, Data Science, DevOps, and Product.

Read full guide →

Design Career Guide

UX/UI, Graphic Design, Product Design. Portfolio, tools, interviews, and growth in the Design field.

Read full guide →

Marketing Career Guide

SEO, Paid Media, Growth, Content Marketing. Certifications, tools, and strategies to grow in Digital Marketing.

Read full guide →

Finance Career Guide

Financial market, investments, corporate finance, certifications, and strategies to grow in the financial field.

Read full guide →

Communication Career Guide

Journalism, PR, Corporate Communication, Content Marketing, and Multimedia Production.

Read full guide →

Administration Career Guide

Business Management, HR, Logistics, Consulting, Project Management, and Entrepreneurship.

Read full guide →

Data Career Guide

Data Science, Data Engineering, BI, Machine Learning, and AI. From training to the job market.

Read full guide →

Product Career Guide

Product Management, Product Ownership, Agile, Scrum, and OKRs. From strategy to execution.

Read full guide →

Tech & Remote Glossary

Stop getting lost in interviews and job descriptions

The job market, especially within tech and global companies, has developed its own dialect. Not understanding these acronyms can make you lose valuable opportunities or poorly negotiate your contract. To end this problem, we created the Definitive Glossary for the Remote Professional.

🏢 Work Models & Routine

Async (Asynchronous Work)
A communication model where responses don't need to be immediate. Instead of back-to-back meetings, the team relies on well-structured documents, threads, and messages. It's the gold standard for global companies spanning multiple time zones.
Sync (Synchronous Work)
The opposite of Async. It requires the team to be online and available at the same time for meetings, live chats, and real-time collaboration.
Daily / Stand-up
A quick daily meeting (usually 15 minutes) common in Agile (Scrum) methodologies. The team answers three questions: What did I do yesterday? What will I do today? Are there any blockers?
All-Hands / Town Hall
A company-wide meeting involving all employees. Usually led by the founders (C-Levels) to present results, new goals, and answer team questions.
1:1 (One-on-One)
A recurring individual meeting between a professional and their direct manager. It is used for career alignment, feedback, and problem-solving, not just for project status updates.

💰 Contracts, Benefits & Compensation

PTO (Paid Time Off)
Instead of strict, categorized leave policies, modern US companies usually offer a flexible pool of days (e.g., 20 days of PTO, or even "Unlimited PTO") that you can use for vacations, sick days, or personal matters, while receiving your regular compensation.
Equity / Stock Options
Company ownership. The startup offers you the right to buy shares at a heavily discounted strike price in the future. If the company grows, goes public, or is acquired, these shares can be highly lucrative.
Vesting (Vesting Schedule)
The rule that controls your Equity. It usually lasts 4 years. You don't get all the shares on day one; you "earn" them gradually as you stay with the company. A "1-year Cliff" means you must stay for at least one year to receive your first batch of shares.
RSUs (Restricted Stock Units)
Unlike Stock Options (where you have the right to buy the stock), RSUs are actual shares the company grants you as a bonus or part of your compensation package, following a strict Vesting schedule.
Independent Contractor (1099 / B2B)
The most common international hiring model for global talent working for US companies. You act as a service provider (business-to-business), receiving the gross salary (often six-figure compensation) without standard local payroll tax deductions at the source.

🤖 Recruitment & Hiring Process

ATS (Applicant Tracking System)
The "robot" that reads your resume. Software like Greenhouse, Ashby, and Workday are used to filter candidates by keywords before a human even looks at the document. (Pro tip: this is why your resume must be clean, semantic, and have the right keywords).
JD (Job Description)
The document that lists the responsibilities, technical requirements, and benefits of the open position.
Cultural Fit
The interview stage that evaluates if your core values, communication style, and worldview align with the company's culture. This is the ultimate test of your Soft Skills.
Onboarding
The integration process. It's the period of your first few weeks at the company, where you get your access credentials, learn about the culture, study the internal documentation, and understand how the product works.

🚀 How to use this to your advantage?

The secret isn't just knowing what these acronyms mean, but using them actively. If during an interview for a premium tech role you ask, "How does your PTO policy and Vesting schedule work?", the recruiter will immediately perceive you as a high-level professional, familiar with the global market standards.

The remote job market requires preparation. And having the right vocabulary is the first big step to securing six-figure proposals and standing out among thousands of applicants.

Expert Tip

The New Paradigm of Digital Marketing and Social Media

The New Paradigm of Digital Marketing and Social Media: Impact, Opportunities, and the Future of the Market

By Mondywork Team | 6-minute read

The Evolution: From "Just Posting" to True Revenue Engines

There was a time when digital marketing was viewed as a secondary department, heavily focused on vanity metrics like likes and follower counts. That era is officially over. In today's global market, Digital Marketing and Social Media Management are the beating heart of customer acquisition, acting as literal Revenue Engines.

Top-tier US startups, unicorns, and global B2B/B2C enterprises have realized that consumer attention is the most expensive asset on the internet. It's no longer just about creativity; it's about merging authentic communication with ruthless data analysis and Artificial Intelligence integrations.

How Does This Market Impact Global Companies?

The impact of modern marketing on the corporate world is absolute. According to the HubSpot State of Marketing report, companies that closely align their digital strategies with sales see a 20% higher revenue growth.

  • CAC Reduction and LTV Focus: Companies rely on marketing professionals to lower Customer Acquisition Cost (CAC) through organic traffic and social media, while simultaneously increasing Lifetime Value (LTV) through retention strategies.
  • Community-Led Growth: Brands don't just want consumers anymore; they want raving fans. Social media has evolved into a community-building channel, where engagement breeds brand advocates who buy repeatedly.
  • Global and Borderless Presence: Digital marketing has shattered geographical borders. A startup in Silicon Valley can easily scale its sales in Europe using campaigns managed by a strategist based in Brazil, Portugal, or anywhere else, working 100% remote.

The 2026 Landscape: What Do Companies Expect?

The digital marketing job market has never been more technical and demanding. The era of the "jack-of-all-trades" generalist is making way for hyper-focused experts and strategists with a systemic vision—the highly sought-after "T-Shaped" professionals.

"Marketing has transitioned from an art based purely on intuition to a science driven by data, predictability, and hyper-segmentation."
— eMarketer Digital Trends

The barrier to entry has risen significantly due to the integration of Artificial Intelligence. Tools that generate copy or images in seconds (like ChatGPT and Midjourney) haven't replaced the professional; they've raised the bar. The market now seeks talent who can "pilot" these AIs to craft strategies, automate workflows using tools like n8n or Zapier, and analyze Business Intelligence (BI) dashboards in Looker Studio.

Golden Opportunities: The Hottest Roles

Highly qualified professionals are being fiercely fought over, especially on platforms like Mondywork, where international companies seek remote talent offering six-figure compensation and global remote flexibility. The areas with the highest demand include:

  1. Growth Marketer / Hacker: The professional focused on rapid experimentation across the marketing funnel (Pirate Metrics - AARRR). They blend marketing, data analysis, and a bit of product engineering to find scalable growth levers.
  2. Performance Manager (Paid Media): The traditional traffic manager has evolved. Today, this professional manages multi-million dollar budgets across Google Ads, Meta Ads, and LinkedIn Ads (essential for B2B), with an absolute focus on ROAS (Return on Ad Spend).
  3. Social Media Strategist & Community Manager: Going far beyond scheduling posts, this talent understands internet culture, dominates short-form formats (TikTok, Reels, Shorts), and knows how to convert a social audience into qualified leads within the company's sales funnel.
  4. SEO & Content Architect: With shifting search algorithms and the impact of AI (Search Generative Experience), the SEO professional who masters search intent, information architecture, and Core Web Vitals is indispensable for long-term organic acquisition.
  5. Marketing Operations (Marketing Ops): The architect behind the curtain. They manage CRMs (Salesforce, HubSpot), email automations, and ensure the entire marketing data infrastructure is flawless.

Is It Worth Entering or Transitioning into This Field?

Absolutely. If you are an analytical, curious professional who understands human behavior and isn't afraid of data spreadsheets, digital marketing offers one of the most promising and flexible career paths in the modern world.

To stand out in rigorous international recruitment systems (ATS like Greenhouse and Ashby), it is crucial that your resume and portfolio showcase real results. Forget generic descriptions; US tech recruiters and Heads of Marketing want to read hard metrics: "Increased organic traffic by 150% in 6 months" or "Managed $500k in ad spend while maintaining CAC 20% below target".

The global market is vast, decentralized, and pays incredibly well for those who know how to turn clicks and views into predictable revenue.