Senior Site Reliability Engineer (Arlington, Va) - Secret Clearance Required - Relocation Provided
onebrief
Job Score
100 ptsConsequential Work. Dedicated People.
About Onebrief
Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.
Today, many critical planning workflows still rely on fragmented systems, static documents, and disconnected tools that make collaboration and decision-making unnecessarily difficult. Onebrief brings modern software, AI, and real-time collaboration into those environments, helping teams operate with greater clarity, coordination, and adaptability in situations where decisions carry real-world consequences.
We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.
Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.
Security Clearance, Location, and Onsite Notice:
This is a hybrid role, requiring regular work on-site at customer locations in Arlington, VA - about 50/50 on-site vs remote.
If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).
Active Secret Clearance required.
About The Role
We're hiring a Site Reliability Engineer to join our Infrastructure & Security team. You'll work closely with product engineers, fellow SREs, security, and customer success.
This is an SRE role for someone who's comfortable in application code. Much of the reliability and performance work happens in the codebase (primarily TypeScript), so you'll fix problems at the source rather than working around them in the infrastructure. You'll be a first line of support for our mission-critical deployments across on-prem DoD and AWS environments, and what you learn in the field will feed directly back into the product.
You'll ship code that makes Onebrief more stable, faster, and easier to deploy and operate. The work sits at the seam between engineering and operations, and it's weighted toward engineering.
About You
You treat reliability as a feature, not an afterthought, and you'd rather fix a problem in the code than route around it. You understand the full software development lifecycle (design, review, testing, release) and you know where reliability fits into each step.
You're comfortable reading and writing application code, and you're just as happy dropping into a kubectl shell to triage a production issue. You turn failure modes into guardrails, and you think monitoring, alerting, and clear runbooks are part of building software, not extra credit.
You mentor others and push a culture of blameless postmortems. You work naturally with product and platform teams, helping them move fast without breaking things by giving them the tools, tests, and observability that make quick recovery real.
What You'll Do
You'll help make our production application reliable, scalable, and secure by improving the software itself, not just the systems it runs on. Day to day that looks like:
Improving the application: Work directly in the codebase (primarily TypeScript) to fix reliability and performance problems at the source. You'll partner with product engineers on design decisions, review code with reliability and security in mind, and treat "make the app better" as a first-class part of the job rather than something you hand off.
Building observability that developers actually use: Design and run our monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana). The goal is alerts and dashboards tied to real application behavior, so teams catch issues before users do.
Owning reliability targets: Define and measure SLIs and SLOs, wire up alerting that feeds them, and be the person who can say what "reliable" means for our systems and prove it with data.
Leading incident response: Act as incident responder, and incident commander when needed. Run blameless post-mortems (AARs) that find the actual root cause and turn it into a code or process fix so it doesn't happen again.
Automating away toil: Spot the repetitive operational work and write software to kill it. Share what works with other teams, including those running in air-gapped environments, and help them get production-ready.
What We Look For
An active Secret clearance
5+ years in software engineering, SRE, or a related role, with real time spent writing and shipping application code
Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript)
Solid grasp of the full SDLC: design, code review, testing, release, and how reliability fits into each stage
Experience with incident response, root cause analysis, and turning findings into lasting fixes
A collaborator who works well across product, platform, and DevOps teams and shares context openly
Technical expertise
Application development in TypeScript (Node and/or a modern front-end framework)
CI/CD: building and maintaining pipelines (GitHub Actions, GitLab CI/CD, Jenkins)
Testing and quality practices as part of the delivery process
Comfort with at least one of Python, Go, or Bash for tooling and automation
Working knowledge of containers and Kubernetes (enough to debug and deploy, not necessarily to stand up clusters from scratch)
Networking fundamentals and secure configuration basics
Bonus points (nice to have)
Observability: Grafana stack, ELK, or Datadog
Infrastructure as Code (Terraform, Ansible) and cloud experience (AWS or AWS GovCloud)
Kubernetes cluster design and operations
Designing meaningful SLIs/SLOs with error budgets for distributed systems
GitOps practices and toolchains
DoD environments and compliance frameworks (RMF, STIGs, ICD 503)
Service mesh (Istio, Linkerd)
On-prem virtualization (VMware, Proxmox, Nutanix, Hyper-V)
Relevant certs (AWS DevOps Engineer, CKA/CKAD)
Notice to Third Party Recruitment Agencies
Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.
About Web Master
The Web Master is the professional responsible for maintaining, securing, and ensuring the technical performance of websites and web applications. They manage servers, hosting infrastructure, uptime monitoring, and ensure everything runs fast and reliably.
Key skills include server management (Apache, Nginx), hosting (AWS, Google Cloud, Azure), CDN (Cloudflare), SSL, DNS, web security (WAF, firewall), performance (Core Web Vitals, cache, compression), and versioning (Git, CI/CD). Knowledge of Docker, WordPress, cPanel, and monitoring (Sentry, New Relic) is a differentiator.
Web Masters in technology companies are highly valued, especially those who master DevOps, SRE, and can guarantee uptime and performance at scale. The field offers opportunities from junior webmaster to SRE and infrastructure engineer, with a focus on reliability, security, and speed.
About Public Relations
The Public Relations (PR) area focuses on managing the reputation, image, and communication of an organization with its various stakeholders (such as clients, investors, employees, media, and the community). PR professionals develop corporate communication strategies, manage media relations (press relations), organize institutional events, and work in image crisis prevention and management.
Discover Other Areas
Understand the scope of work, key skills, and tools used in different career areas.
About Frontend
The Frontend area is responsible for creating the visual interfaces that users interact with on websites and web applications. Frontend professionals combine technical skills with design to deliver intuitive, responsive, and accessible digital experiences.
Key skills include HTML, CSS, JavaScript/TypeScript, frameworks like React, Angular, and Vue, build tools (Webpack, Vite), CSS (Tailwind, Sass), testing (Jest, Cypress), and knowledge of web performance and accessibility (WCAG). Familiarity with design systems and reusable components is a differentiator.
Frontend developers in technology companies are highly valued, especially those who master React, Next.js, web performance, and accessibility. The field offers opportunities from junior developer to frontend architect, with a focus on user experience, performance, and code quality.
About Audiovisual
The Audiovisual area is responsible for producing, editing, and creating video and audio content for various platforms. With the exponential growth of digital content, audiovisual professionals are fundamental for brands that want to communicate visually and impactfully.
Key skills include video production and editing (Premiere Pro, DaVinci Resolve, Final Cut), motion graphics (After Effects), animation (Blender, Cinema 4D), sound design, podcast production, live streaming (OBS Studio), and photography. Knowledge of visual storytelling, rhythm, and art direction is a differentiator.
Audiovisual professionals in technology companies are highly valued, especially those who master motion graphics, social media videos, and content for digital platforms. The field offers opportunities from videomaker to head of audiovisual, with a focus on creativity, technical quality, and storytelling.
About Content
The Content and Social Media area is essential for building digital presence and audience engagement. Professionals create content strategies, manage social networks, and develop impactful brand narratives.
Key skills include copywriting, storytelling, community management, metrics analysis, audiovisual production, and knowledge of each platform algorithms.
With the growth of influencer marketing and social commerce, this area continues to generate new career opportunities.
About Software Development
Software Development is one of the most dynamic and constantly evolving fields in the job market. Professionals in this area are responsible for creating, maintaining, and optimizing web, mobile, and desktop applications that impact millions of users daily.
Key languages and frameworks include JavaScript (React, Node.js, Vue.js), Python (Django, Flask), Java (Spring), PHP (Laravel), and TypeScript. Demand for full-stack developers continues to grow, especially in tech companies and startups.
Salaries range from entry-level to senior positions, with growing opportunities for remote work and international freelancing.
About People Analyst
The People Analyst is the professional responsible for transforming people data into strategic insights for HR decision-making. They combine data analysis knowledge with people management vision to help organizations understand workforce metrics, turnover, engagement, and diversity.
Key skills include people analytics, workforce analytics, turnover and retention analysis, HR metrics (time-to-hire, cost-per-hire, e-NPS), data visualization (Power BI, Tableau, Visier), workforce planning, and compensation analysis. Knowledge of statistics, SQL, and people analytics tools is a differentiator.
People Analysts in technology companies are highly valued, especially those who can translate complex people data into actionable insights for retention, diversity, and growth strategies. The field offers opportunities from HR analyst to head of people analytics, with a focus on data-driven people management.
Comments 0