# Grafana & Observability Engineer - Dallas, Tampa & Jersey City

**Company:** [StradIT](http://jobs.workable.com/companies/gUMYmaPRJguCYquNxEjRoL.md)
**Location:** Jersey City, United States
**Workplace:** hybrid

[Apply for this job](http://jobs.workable.com/view/5445e810-7d0a-4d18-8abc-639a91e15665)

## Description

**Role: Grafana & Observability Engineer**

**Experience: 5 to 10 years**

**Employment: W2**

**Location: Jersey City NJ, Tampa FL & Dallas TX**

**Key Responsibilities**

**Observability Platform Engineering**

-   Administer and support Grafana Cloud and on-premises Grafana deployments.
-   Design and implement enterprise observability solutions for metrics, logs, traces, synthetic monitoring, and alerting.
-   Establish and maintain observability standards, best practices, and governance processes.
-   Configure and manage Grafana data sources, alerting, RBAC, folders, teams, and integrations.
-   Ensure platform scalability, reliability, resiliency, and operational excellence.

**Automation & Infrastructure as Code**

-   Develop and maintain Terraform modules for Grafana infrastructure and configuration management.
-   Automate onboarding of applications, infrastructure, dashboards, alerts, and data sources.
-   Build self-service capabilities that reduce manual operational effort and improve adoption.
-   Integrate observability capabilities into CI/CD and infrastructure provisioning workflows.

**Monitoring, Alerting & Incident Management**

-   Design meaningful monitoring and alerting strategies based on service health and business-critical workflows.
-   Implement and optimize alerting standards to reduce noise and improve signal quality.
-   Support incident response, troubleshooting, root cause analysis, and post-incident reviews.
-   Drive continuous improvement of operational visibility and platform health.

**OpenTelemetry & Telemetry Engineering**

-   Implement and support OpenTelemetry instrumentation across applications and infrastructure.
-   Establish standards for logs, metrics, traces, and telemetry collection.
-   Support telemetry pipelines, agent deployments, and data collection strategies.
-   Assist teams with instrumentation design and observability adoption.

**Migration & Modernization**

-   Support migration initiatives from legacy observability platforms to Grafana.
-   Analyze existing monitoring, alerting, logging, and tracing implementations and recommend modernization approaches.
-   Develop reusable migration patterns, automation, and engineering standards.
-   Partner with application teams to accelerate adoption of enterprise observability capabilities.

**Collaboration & Leadership**

-   Work closely with application development, infrastructure, cloud, and SRE teams.
-   Provide technical leadership and mentoring to engineers across the organization.
-   Contribute to observability architecture, strategy, and roadmap development.
-   Promote observability as a core engineering practice across the enterprise.

**Required Qualifications**

-   Bachelor's degree in Computer Science, Engineering, Information Systems, or related field.
-   5+ years of experience in observability, monitoring, operations, or platform engineering.
-   Hands-on experience administering Grafana in large-scale enterprise environments.
-   Strong experience with Terraform and Infrastructure as Code practices.
-   Experience implementing monitoring, alerting, logging, and distributed tracing solutions.
-   Experience with OpenTelemetry concepts, instrumentation, and telemetry pipelines.
-   Strong Linux and cloud platform administration skills.
-   Experience with scripting and automation using Python, PowerShell, Bash, or similar languages.
-   Knowledge of operational excellence, reliability engineering, and incident management practices.

**Preferred Qualifications**

-   Experience migrating from tools such as Splunk, Dynatrace, AppDynamics, New Relic, OpenText OBM, or similar platforms.
-   Experience with Grafana Alloy, Tempo, Loki, Mimir, or Prometheus.
-   Experience operating observability platforms in AWS environments.
-   Knowledge of Kubernetes, containers, and cloud-native observability.
-   Experience designing enterprise observability strategies and governance models.
-   Familiarity with CI/CD platforms and DevOps practices.

**Desired Skills**

-   Grafana Administration
-   Terraform
-   OpenTelemetry (OTEL)
-   Monitoring & Alerting
-   Observability Engineering
-   Platform Engineering
-   Linux Administration
-   AWS Cloud Services
-   Automation & Scripting
-   Incident Management
-   Infrastructure as Code
-   Telemetry Pipelines
-   Reliability Engineering
-   Root Cause Analysis
-   Enterprise Monitoring Architecture
