Job Description

We are looking for an Senior Site Relability Engineer to join our growing engineering team.

We are a company that values SRE principles and practices. We believe in empowering our SREs to make data-driven decisions, automate operational tasks, and continuously improve the reliability of our systems. We foster a blameless culture where everyone is encouraged to learn from mistakes and share knowledge. If you are passionate about building and maintaining highly reliable systems, we would love to hear from you!

What you'll do:

Lead the design of scalable, fault-tolerant and self-healing systems in a multi-region AWS environment.

Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to drive architectural decisions and error budget policies.

Conduct blameless post-incident reviews to uncover systemic root causes and implement long-term preventive measures.

Identify patterns of manual work and lead the development of internal tools/automation to permanently eliminate them.

Develop and maintain automated runbooks and playbooks for common operational tasks and complex incident response.

Shift from simple monitoring to deep observability, ensuring high cardinality data leads to proactive actionable insights.

Proactively identify and mitigate operational risks through chaos engineering and architecture reviews.

Work with software engineers to design systems for reliability, scalability, and maintainability from the early stages of the SDLC.

Continuously evaluate and optimize system performance, capacity, and cost efficiency.

Beyond just participating, you will refine the on-call experience to reduce alert fatigue, improve MTTR, and ensure sustainable rotation health.

Must Haves:

Bachelor’s degree in Computer Engineering or a similar discipline.

5+ years of experience as a Site Reliability Engineer or in a similar role.

3+ years of experience with AWS services including strong knowledge of container orchestration.

2+ years of Kubernetes experience

Deep understanding of observability principles and tools like (Prometheus, Datadog, OpenTelemetry).

Experience with leading incident management and complex postmortem analysis.

Experience and interest in managing infrastructure as code (Terraform).

Experience with chaos engineering and other techniques for testing system resilience.

Experience with CI/CD tools such as GitHub Actions ****for automated delivery.

Proficiency in at least one programming language (Python, Go, Java, etc.) for building automation and internal tooling.

Event-driven architecture experience (SNS, SQS etc)

Ability to work independently and collaboratively in a fast-paced environment.

Team player and open to new ideas.

Good communication skills and fluency in English.

Good to have:

Prior experience with Scrum and other agile methods.

Certification in relevant areas such as AWS Certified DevOps Engineer, Certified Kubernetes Administrator (CKA), or similar.

Prior experience with Telco Core Networks (e.g., 5G/LTE Packet Core, IMS, Signaling) and low-latency networking.

Experience with AI-driven SRE tools for anomaly detection and improvements

Contributions to open-source SRE projects or communities.

Prior work experience in telecommunications.

Deep understanding of eSIM and GSMA related technologies and services.

Apply Now

Senior Site Reliability Engineer

Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interes

admin
engineer

Director, Software Engineering (Site Reliability Engineering)

Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interes

engineer
dev
exec

Senior Site Reliability Engineer

TextNow is looking for motivated Senior Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!What You'll DoE

Senior
admin
engineer

Site Reliability Engineer II

Role SummaryAs an Intermediate Site Reliability Engineer, you will support the reliability, performance, and scalability of cloud-hosted services and database platforms.

admin
engineer

Senior Site Reliability Engineer

Job Description

Remote United Kingdom

SRE Engineer

2 months ago

Senior

Senior Site Reliability Engineer

Director, Software Engineering (Site Reliability Engineering)

Senior Site Reliability Engineer

Site Reliability Engineer II

Find Remote Jobs

About us

Additional

Senior Site Reliability Engineer

Job Description

Remote United Kingdom

SRE Engineer

2 months ago

Senior

Senior Site Reliability Engineer

Director, Software Engineering (Site Reliability Engineering)

Senior Site Reliability Engineer

Site Reliability Engineer II

Subscribe to Job Alerts

Find Remote Jobs

About us

Additional