As a Senior Site Reliability Engineer (SE III) at Ophelia, you will be our dedicated reliability, availability, and performance hire, playing a key role in making the systems that support our mission of treating opioid use disorder through telehealth stable, observable, and fast. Our stack is TypeScript on Node, React, and Firebase (e.g., Firestore, Authentication, Hosting) running on Cloud Run in Google Cloud Platform (GCP). Your work will have a direct impact on patients, clinicians, and our ability to scale.
You will own and drive our system stability and availability initiative, which is already underway: Terraform-driven uptime checks feed our engineering key performance indicators (KPIs), we have an initial service level objective (SLO) plan, and we've recently revamped our on-call rotation and runbook to keep the on-call engineer focused on severity incidents (SEVs). You will take this from a good start to a mature practice: defining and operationalizing SLOs, building out logging, monitoring, and alerting grounded in sound measurement (e.g., percentile-based latency indicators rather than averages), moving incident response into PagerDuty with clear categorization and escalation paths, expanding our infrastructure as code (IaC) footprint, and working down known pain points such as long-latency endpoints and Cloud Run cold starts. You will also partner closely with our head of IT on IAM hardening, secrets management, audit logging, and HIPAA/SOC 2 infrastructure controls. You will be the directly responsible individual (DRI) for our mean time to recovery (MTTR) KPI and for the measurement and reporting of our availability KPIs.
As a senior engineer, you will be a force multiplier for our team of fullstack engineers who are focused on product work. That means leveling up the team's GCP and DevOps skills through documentation, runbooks, pairing, and code review, and bringing reliability recommendations to new features as they are designed, while primarily owning the infrastructure work yourself. This role is entirely reliability-focused for the first six months; after that, you may contribute to product work as needed, though product work will be at most about 25% of your time.
Ophelia's Technology team actively encourages and invests in AI-augmented ways of working: using tools like Claude and Gemini to accelerate infrastructure code, runbooks, log analysis, and incident triage, and staying current as the AI landscape evolves. Consistent with Ophelia's AI Position Statement, we treat AI as a force multiplier for our team, not a replacement for human judgment: it's there to help you move faster from alert to root cause and spend more time on the reliability and architecture decisions that actually require a person. You'll be expected to use AI fluently in your own workflow, to build it into our operational tooling, and to help shape how the larger technology team uses it responsibly and effectively as it evolves.
Together, we will help hard-to-reach individuals treat their opioid dependence. While direct experience in this treatment area is not mandatory, knowledge of the healthcare space, including understanding health outcomes, benchmarks, systems, and regulatory compliance (HIPAA), is highly beneficial.
#LI-Remote
Interested in learning more about Ophelia and this role? Apply to work with us!
A Note Before Applying
We believe a good hiring process is really just about getting to know each other. That's why we care about getting to know the real you, not a polished, AI-generated version of you. Every application is read by a real person who's genuinely curious about your experiences, your voice, and why this role speaks to you.
So as you put together your application, we'd love for it to sound like you. It's the best way for us to get to know each other, and the best way for you to find out if we're the right fit.
This role is hosted on Ophelia's own careers site. You will complete your application there.
Apply on company site