← Back to jobs

Senior Site Reliability Engineer, Colorado Springs (Top Secret Clearance Required, Relocation Provided)

onebrief

Colorado Springs, CO
Uncategorized Web Master Public Relations Advertising

Job Score

100 pts
On-site model (+70) Web Master (+10) Public Relations (+10) Advertising (+10)

Consequential Work. Dedicated People.


About Onebrief

Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.

Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.

We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.

Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.

Security Clearance, Location, and Onsite Notice:

This role requires regularly working on-site at customer locations in Colorado Springs, Colorado.

If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).

Active Top Secret Clearance required; SCI eligibility is a plus.

About The Role

We are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with fellow SREs, security, and customer success.

You will be the first line of support for our mission critical deployments, and responsible for ensuring best-in-class service quality and issue resolution. You will work in both on-premise DoD environments and AWS cloud environments. Your lessons from the field will shape how our team works, from policy to implementation.

In addition to working at the customer, you will contribute directly to solutions that increase stability, performance, and security of our deployments, and improve the overall experience of deploying and managing Onebrief on premise.

About You

You care deeply about reliability and treat it as a core feature of any application or platform, with a bias toward “reliability over novelty.” You think about infrastructure and operability as products to be automated, well-documented, and continuously improved, and you aim to leave systems easier to operate than you found them.

You are equally comfortable leading a post-incident review, or diving into a kubectl shell to triage a complex production issue. You don't just fix problems; you translate constraints and failure modes into clear, automated guardrails and scalable, resilient architecture. For you, robust monitoring, actionable alerting, and insightful runbooks are core parts of the engineering process, not afterthoughts.

You mentor others, fostering a culture of blameless postmortems and proactive reliability. You collaborate naturally with application and platform teams, helping them move quickly but safely by building the tools, processes, and observability that make "fast recovery" a reality.

What You'll Do

You'll own the reliability, scalability, and security of the production application and/or platform. You will do this by:

  • Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users.

  • Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it.

  • Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence.

  • Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation.

  • Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.

What We Look For

  • An active Top Secret clearance

  • 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.

  • Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly.

  • A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.

Technical expertise

  • Infrastructure as Code: Terraform (or CloudFormation), Ansible.

  • Containers and orchestration: Kubernetes design, deployment, and operations.

  • CI/CD: experience building and maintaining pipelines (GitLab CI/CD, Jenkins, GitHub Actions).

  • Scripting: proficiency with at least one of Python, Go, or Bash.

  • Cloud: Familiarity with AWS or AWS GovCloud.

  • Observability: Grafana stack, ELK stack, or Datadog.

  • Networking fundamentals: core protocols and secure configurations.

Bonus points (nice to have)

  • Experience in DoD environments and compliance frameworks (RMF, STIGs, ICD 503).

  • GitOps practices and toolchains.

  • Security‑minded design for sensitive environments.

  • Experience designing and implementing meaningful SLIs/SLOs (including error budgets) for complex, distributed systems.

  • Familiarity with on‑prem virtualization(VMware, Proxmox, Nutanix, Hyper-V, etc).

  • Service mesh exposure (Istio, Linkerd).

  • Relevant certifications (e.g., AWS DevOps Engineer, CKA/CKAD).

  • Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain the valid credentials within 3 months of employment.


Notice to Third Party Recruitment Agencies

Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.

What did you think of this job?

Comments 0

Want to leave a comment?
Sign in or create your account in seconds to join the discussion.
Loading comments...

About Web Master

The Web Master is the professional responsible for maintaining, securing, and ensuring the technical performance of websites and web applications. They manage servers, hosting infrastructure, uptime monitoring, and ensure everything runs fast and reliably.

Key skills include server management (Apache, Nginx), hosting (AWS, Google Cloud, Azure), CDN (Cloudflare), SSL, DNS, web security (WAF, firewall), performance (Core Web Vitals, cache, compression), and versioning (Git, CI/CD). Knowledge of Docker, WordPress, cPanel, and monitoring (Sentry, New Relic) is a differentiator.

Web Masters in technology companies are highly valued, especially those who master DevOps, SRE, and can guarantee uptime and performance at scale. The field offers opportunities from junior webmaster to SRE and infrastructure engineer, with a focus on reliability, security, and speed.

About Public Relations

The Public Relations (PR) area focuses on managing the reputation, image, and communication of an organization with its various stakeholders (such as clients, investors, employees, media, and the community). PR professionals develop corporate communication strategies, manage media relations (press relations), organize institutional events, and work in image crisis prevention and management.

About Advertising

The Advertising area is aimed at the planning, creation, and delivery of communication campaigns to promote brands, products, ideas, or services. Professionals in the sector work in advertising agencies or in-house marketing departments in creative fields (art direction, copywriting), strategic planning, account management, and media buying.

Discover Other Areas

Understand the scope of work, key skills, and tools used in different career areas.

About Systems Analyst

The Systems Analyst is the professional responsible for analyzing, designing, and implementing technology solutions that meet business needs. They act as a bridge between business areas and the development team, ensuring that systems deliver real value to the organization.

Key skills include requirements gathering and analysis, process modeling (BPMN), data modeling, technical and functional documentation, system integration (APIs, microservices), and knowledge of ERPs and CRMs. Tools like Jira, Confluence, Visio, and project management platforms are essential.

Systems Analysts in technology companies are highly valued, especially those who master agile requirements analysis (user stories, backlog), system integration, and solution architecture. The field offers opportunities from junior analyst to solution architect, with a focus on efficiency, quality, and technological innovation.

About IT Governance

IT Governance is the area responsible for ensuring that information technology resources are used strategically, efficiently, and in compliance with standards and regulations. IT governance professionals ensure that technology supports business objectives in a secure and reliable manner.

Key skills include IT service management (ITIL), IT audit and compliance, risk management, business continuity, disaster recovery, metrics and indicators (SLAs, KPIs), and strategic alignment between IT and business. Frameworks like COBIT, ITIL, ISO 27001, and compliance standards are essential.

IT Governance professionals in technology companies are highly valued, especially those who master ITSM, IT audit, and risk management. The field offers opportunities from governance analyst to CIO/CTO, with a focus on efficiency, compliance, security, and business value.

About Social Media

The Social Media area is one of the most dynamic and constantly evolving fields in digital marketing. Social media professionals are responsible for creating, managing, and optimizing brand presence on digital platforms, building engagement and community with the target audience.

Key skills include social media management (Instagram, TikTok, LinkedIn, Facebook, YouTube), social media content creation, community management, paid social media (Meta Ads, LinkedIn Ads, TikTok Ads), metrics analysis, and strategic planning. Tools like Hootsuite, Sprout Social, Buffer, Later, and analytics platforms are essential.

Social media professionals in technology companies are highly valued, especially those who master paid social, social media analytics, and content strategies for different platforms. The field offers opportunities from analyst to head of social media, with a focus on growth, engagement, and return on investment.

About Data

The Data field has undergone a radical transformation with the rise of Generative AI. Data professionals are fundamental for evidence-based decision-making across all industries.

Key specializations include Data Engineering, Data Science, Business Intelligence, Machine Learning Engineering, and Analytics. Tools like SQL, Python, Spark, dbt, and cloud platforms (AWS, GCP, Azure) are essential.

The data market continues with high demand and salaries among the most competitive in the technology sector, with many remote work opportunities.

About Account Manager

The Account Manager is the professional responsible for managing and expanding the relationship with clients after the sale. They act as a strategic partner, ensuring satisfaction, retention, and account growth, connecting client needs with company solutions.

Key skills include relationship management, negotiation, upsell and cross-sell, contract renewal, account planning, business reviews, metrics analysis (NPS, churn, LTV), and CRM knowledge (Salesforce, HubSpot). Communication, empathy, and business vision are fundamental differentiators.

Account Managers in technology and SaaS companies are highly valued, especially those who can increase recurring revenue (MRR/ARR) through account expansion and churn prevention. The field offers opportunities from account executive to director of accounts, with a focus on strategic relationship, revenue growth, and customer success.

Career Guides

Technology Career Guide

Planning, skills, interviews, and professional growth in IT, Data Science, DevOps, and Product.

Read full guide →

Design Career Guide

UX/UI, Graphic Design, Product Design. Portfolio, tools, interviews, and growth in the Design field.

Read full guide →

Marketing Career Guide

SEO, Paid Media, Growth, Content Marketing. Certifications, tools, and strategies to grow in Digital Marketing.

Read full guide →

Finance Career Guide

Financial market, investments, corporate finance, certifications, and strategies to grow in the financial field.

Read full guide →

Communication Career Guide

Journalism, PR, Corporate Communication, Content Marketing, and Multimedia Production.

Read full guide →

Administration Career Guide

Business Management, HR, Logistics, Consulting, Project Management, and Entrepreneurship.

Read full guide →

Data Career Guide

Data Science, Data Engineering, BI, Machine Learning, and AI. From training to the job market.

Read full guide →

Product Career Guide

Product Management, Product Ownership, Agile, Scrum, and OKRs. From strategy to execution.

Read full guide →

Tech & Remote Glossary

Stop getting lost in interviews and job descriptions

The job market, especially within tech and global companies, has developed its own dialect. Not understanding these acronyms can make you lose valuable opportunities or poorly negotiate your contract. To end this problem, we created the Definitive Glossary for the Remote Professional.

🏢 Work Models & Routine

Async (Asynchronous Work)
A communication model where responses don't need to be immediate. Instead of back-to-back meetings, the team relies on well-structured documents, threads, and messages. It's the gold standard for global companies spanning multiple time zones.
Sync (Synchronous Work)
The opposite of Async. It requires the team to be online and available at the same time for meetings, live chats, and real-time collaboration.
Daily / Stand-up
A quick daily meeting (usually 15 minutes) common in Agile (Scrum) methodologies. The team answers three questions: What did I do yesterday? What will I do today? Are there any blockers?
All-Hands / Town Hall
A company-wide meeting involving all employees. Usually led by the founders (C-Levels) to present results, new goals, and answer team questions.
1:1 (One-on-One)
A recurring individual meeting between a professional and their direct manager. It is used for career alignment, feedback, and problem-solving, not just for project status updates.

💰 Contracts, Benefits & Compensation

PTO (Paid Time Off)
Instead of strict, categorized leave policies, modern US companies usually offer a flexible pool of days (e.g., 20 days of PTO, or even "Unlimited PTO") that you can use for vacations, sick days, or personal matters, while receiving your regular compensation.
Equity / Stock Options
Company ownership. The startup offers you the right to buy shares at a heavily discounted strike price in the future. If the company grows, goes public, or is acquired, these shares can be highly lucrative.
Vesting (Vesting Schedule)
The rule that controls your Equity. It usually lasts 4 years. You don't get all the shares on day one; you "earn" them gradually as you stay with the company. A "1-year Cliff" means you must stay for at least one year to receive your first batch of shares.
RSUs (Restricted Stock Units)
Unlike Stock Options (where you have the right to buy the stock), RSUs are actual shares the company grants you as a bonus or part of your compensation package, following a strict Vesting schedule.
Independent Contractor (1099 / B2B)
The most common international hiring model for global talent working for US companies. You act as a service provider (business-to-business), receiving the gross salary (often six-figure compensation) without standard local payroll tax deductions at the source.

🤖 Recruitment & Hiring Process

ATS (Applicant Tracking System)
The "robot" that reads your resume. Software like Greenhouse, Ashby, and Workday are used to filter candidates by keywords before a human even looks at the document. (Pro tip: this is why your resume must be clean, semantic, and have the right keywords).
JD (Job Description)
The document that lists the responsibilities, technical requirements, and benefits of the open position.
Cultural Fit
The interview stage that evaluates if your core values, communication style, and worldview align with the company's culture. This is the ultimate test of your Soft Skills.
Onboarding
The integration process. It's the period of your first few weeks at the company, where you get your access credentials, learn about the culture, study the internal documentation, and understand how the product works.

🚀 How to use this to your advantage?

The secret isn't just knowing what these acronyms mean, but using them actively. If during an interview for a premium tech role you ask, "How does your PTO policy and Vesting schedule work?", the recruiter will immediately perceive you as a high-level professional, familiar with the global market standards.

The remote job market requires preparation. And having the right vocabulary is the first big step to securing six-figure proposals and standing out among thousands of applicants.

Expert Tip

Agile and Project Management in Today's Market

1. Move Beyond "Doing Agile": Deliver Real Value

One of the biggest mistakes candidates make in interviews for Agile roles is focusing too much on ceremonies (Dailies, Plannings, Retrospectives) and forgetting their actual purpose. Recruiters from global tech companies want to know how you eliminated bottlenecks and boosted team efficiency.

"True Agile success isn't measured by how many practices you implement, but by how much value you deliver to the customer and how quickly you do it."

Instead of writing "facilitated Scrum ceremonies" on your resume, use a results-driven approach: "Reduced team Cycle Time by 20% by identifying and removing bottlenecks in the value stream."

2. Certifications That Actually Bypass the ATS

While practical experience is the ultimate deciding factor, Applicant Tracking Systems (ATS) used by top US companies (like Greenhouse, Ashby, and Workday) filter resumes strictly through keywords. Holding the right certifications ensures you pass this first robotic barrier and get your profile in front of a human recruiter.

  • Professional Scrum Master (PSM I, II) by Scrum.org: Highly respected because it doesn't require renewal and features a rigorous, Scrum Guide-based exam.
  • Project Management Professional (PMP) by PMI: The gold standard for traditional, predictive, and hybrid Project Management.
  • PMI Agile Certified Practitioner (PMI-ACP): Excellent for demonstrating broad knowledge across multiple Agile frameworks (Scrum, Kanban, Lean).
  • SAFe (Scaled Agile Framework): Absolutely essential if you are targeting Enterprise-level corporations that have scaled Agile across thousands of employees.

3. Mastering Flow Metrics and Tools (Hard Skills)

The modern Agile professional must be data-driven. Companies offering six-figure compensation expect you to extract actionable insights from tools like Jira, Confluence, or Azure DevOps.

During technical interviews, be prepared to discuss how you utilize flow metrics, such as:

  • Lead Time and Cycle Time: To measure efficiency and the actual time it takes to deliver requests.
  • Throughput: The number of work items completed in a given timeframe.
  • Cumulative Flow Diagram (CFD): To visually identify bottlenecks in the workflow.
  • OKRs (Objectives and Key Results): To align the development team's daily work with the company's high-level strategic goals.

4. Soft Skills and Conflict Resolution in Async Teams

In global, remote-first, and asynchronous environments, communication is the biggest challenge. A top-tier Project Manager or Scrum Master acts as a shield for the tech team and a bridge for stakeholders.

Demonstrate emotional intelligence, empathy, and negotiation skills. You will often face situational questions like: "How would you handle a Product Owner who constantly changes the Sprint scope mid-cycle?" Your answer should highlight peaceful, data-driven negotiation focused on business impact and team predictability.

Conclusion: Your Competitive Edge

To land premium Project Management and Agile roles, you must position yourself as a strategic business partner, not just a process enforcer. Understand the product, master the metrics, and know how to guide people towards continuous improvement.

References & Useful Links