# Senior Site Reliability Engineer (Performance and Scalability)

**Company:** [Digital Zone](http://jobs.workable.com/companies/mrSqoYoFT1qi3JjDrAyD4e.md)
**Location:** Remote
**Workplace:** remote
**Employment type:** Full-time
**Department:** Technology

[Apply for this job](http://jobs.workable.com/view/34a90ad4-1521-4e94-9bdb-d9cf24466f8a)

## Description

Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.

What you'll do

-   Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load.
-   Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results.
-   Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale.
-   Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.
-   Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on.
-   Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.

## Requirements

What you'll bring

-   5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute.
-   A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted.
-   Deep AWS experience and a solid grasp of Postgres performance and scaling.
-   Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar.
-   A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep.

## Benefits

-   Immediate, large-scale impact on a high-growth business
-   Top-of-the-market compensation packages
-   Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more
