Build and own a critical agent safety layer for Kubernetes at massive scale, protecting production stability for millions of users. Design and implement distributed systems with phased rollouts and canary releases following Google SRE principles. Lead end-to-end reliability for services from inception to production.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
At BairesDev®, we've been leading the way in technology projects for over 15 years. We deliver cutting-edge solutions to giants like Google and the most innovative startups in Silicon Valley.
Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide.
When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success.
Senior Site Reliability Engineer (SRE) at BairesDev
In this role, you'll build and own an agent safety layer that sits in front of Kubernetes, reasoning about system architecture at massive scale, from a handful of users up to hundreds of millions. Following the classic Google SRE model, you'll build and own a service end-to-end and run it reliably in production, not just administer infrastructure someone else designed. This is your opportunity to work on critical infrastructure where phased rollouts and canary releases are the standard, and where your architectural decisions directly protect production stability at scale.
What You'll Do
- Build and own an agent safety layer sitting in front of Kubernetes, likely written in Go.
- Design systems that scale from a handful of users to hundreds of millions.
- Ensure changes reach production safely through phased rollouts and canary releases.
- Reason about distributed systems architecture and production reliability end-to-end.
Interested in remote work opportunities in Development & Programming? Discover Development & Programming Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
- 5+ years of experience in Site Reliability Engineering.
- Strong experience with distributed systems design.
- Proficiency in Python or Golang.
- Experience with Kubernetes.
- Background in production reliability, including SLAs and SLOs.
- Experience with safe rollout practices such as canary or phased deployment.
- Advanced proficiency in English.
- Experience with observability or metrics tooling.
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
- 100% remote work (from anywhere).
- Excellent compensation in USD or your local currency if preferred
- Hardware and software setup for you to work from home.
- Flexible hours: create your own schedule.
- Paid parental leaves, vacations, and national holidays.
- Innovative and multicultural work environment: collaborate and learn from the global Top 1% of talent.
- Supportive environment with mentorship, promotions, skill development, and diverse growth opportunities.
Apply now!
#BD-PRIO-2026
Similar Jobs
Explore other opportunities that match your interests