# Senior DevOps Engineer, Infrastructure & Reliability

**Company:** [Worth AI](http://jobs.workable.com/companies/k2E8uRLJaqV8qQDdQHzTkn.md)
**Location:** Remote
**Workplace:** remote
**Employment type:** Full-time
**Department:** Engineering

[Apply for this job](http://jobs.workable.com/view/0e424823-768b-47de-9ebc-edfb30b77cdc)

## Description

Worth AI, a leader in the computer software industry, is looking for a Senior DevOps Engineer to join our Infrastructure team with a singular mission: to make our systems faster, more reliable, and more resilient while making life dramatically easier for engineers shipping software. 

This is a hands-on build role. You will spend most of your time writing Terraform, tuning Kubernetes workloads, automating things that are currently manual, and shipping infrastructure changes to production. You'll join a small platform team with an established roadmap and existing patterns, and a strong voice in how the work gets built.

-   Implement scalable Infrastructure-as-Code patterns using tools like Terraform to standardize cloud provisioning and reduce configuration drift.
-   Own and evolve our Kubernetes platform (EKS or self-managed), ensuring workloads are secure, scalable, and resilient by default.
-   Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase confidence in releases.
-   Design and enforce secure networking, IAM, and secrets management strategies across environments.
-   Improve observability by refining metrics, logs, and tracing using tools like DataDog, ensuring actionable insight into system health.
-   Optimize cloud cost efficiency through rightsizing, autoscaling strategies, and architectural improvements.
-   Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives.
-   Refactor brittle or manually managed infrastructure into automated, testable, and reproducible systems.
-   Introduce new infrastructure tooling or architectural shifts and drive adoption through documentation, workshops, and hands-on support.
-   Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments.
-   Communicate technical trade-offs clearly across engineering and product stakeholders, balancing speed with safety.

### **Technology Stack**

-   Cloud & Infrastructure: AWS (EKS, RDS, MSK, S3, Lambda, IAM, VPC)  
    Containerization & Orchestration: Kubernetes, ArgoCD  
    Infrastructure-as-Code: Terraform  
    CI/CD: GitHub Actions  
    Monitoring & Observability: DataDog  
    Data & Messaging: PostgreSQL, Kafka, Redis  
    Languages (as needed): Bash, Python, TypeScript, JavaScript

## Requirements

-   5+ years in DevOps, SRE, or infrastructure engineering.
-   Proven experience designing and operating production Kubernetes environments at scale.
-   Deep hands-on expertise with AWS infrastructure and cloud networking.
-   Strong experience building and maintaining Terraform modules across large cloud environments.
-   Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics.
-   Experience leading incident response processes and driving meaningful postmortem outcomes.
-   Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL).
-   Proven ability to modernize legacy infrastructure and eliminate manual operational toil.
-   Track record of taking a scoped infrastructure project from an ambiguous starting point to production without needing daily direction.
-   Demonstrated ability to build trust across teams while raising the reliability bar.

### **Success Metrics**

-   System Reliability: Maintain or exceed defined SLO/SLA targets with reduced incident frequency and duration.

-   Infrastructure Stability: Reduce production incidents caused by misconfiguration, manual processes, or infrastructure drift.
-   Operational Efficiency: Increase the percentage of infrastructure managed through code and automation.
-   Cost Optimization: Improve cloud cost efficiency without sacrificing reliability or performance.

### Bonus Points (Nice to Have)

-   Experience coding applications
-   Experience operating high-throughput Kafka clusters (MSK or self-managed).
-   Strong background in database performance tuning (PostgreSQL, Redis).
-   Experience implementing autoscaling strategies for high-traffic systems.
-   Familiarity with service mesh technologies.
-   Experience building internal developer platforms (IDP).
-   Background in security best practices (zero-trust networking, policy-as-code).
-   Experience with multi-region or globally distributed systems.
-   Experience introducing platform-wide reliability frameworks (SLOs, error budgets, chaos testing).

**All Remote Hires will be required to travel to Orlando, Florida at least twice per year for Town Halls and team collaboration, in addition to orientation in Orlando.**

## Benefits

-   Health Care Plan (Medical, Dental & Vision)
-   Retirement Plan (401k)
-   Life Insurance
-   Flexible Paid Time Off
-   9 paid Holidays
-   Family Leave
-   Remote
-   Hybrid work (for Orlando Associates)
-   Free Food & Snacks (Orlando)
-   Wellness Resources
