# MLOps / Cloud Deployment Engineer

**Company:** [Xenon7](http://jobs.workable.com/companies/khsAik1aNHDNR3tsuc1Jr7.md)
**Location:** Hyderabad, India
**Workplace:** hybrid
**Employment type:** Contract
**Department:** Delivery and Solutions

[Apply for this job](http://jobs.workable.com/view/d2b0ff58-227b-40c2-9769-5eebc759ac32)

## Description

Our Client's Digital Finance IT is scaling AI and agentic systems in production. We need an **MLOps / Cloud Deployment Engineer** to own the deployment, reliability, observability, and operational scale of these systems in a regulated enterprise environment.

This is a **cloud and platform engineering role** with deep MLOps/LLMOps focus, not a model-building role. You will operate the runway that ML and GenAI systems run on, not build the models themselves.

**What You'll Do**

-   Own **CI/CD pipelines** for ML models, RAG applications, and agentic AI systems — from experiment to production
-   Deploy and operate AI workloads on **cloud-native ML/AI platforms** — AWS Bedrock/SageMaker, Azure AI Foundry / Azure Machine Learning, or equivalent
-   Build and maintain **observability, tracing, and monitoring** for LLM and agentic systems — latency, cost, hallucination rates, tool-call success, drift detection
-   Implement **model governance and guardrails** — approval gates, kill-switches, escalation paths, audit trails
-   Manage **infrastructure-as-code** (Terraform, Bicep, or equivalent) for reproducible AI/ML environments
-   Design **cost and performance optimization** strategies — token usage tracking, caching, model routing, autoscaling, warehouse/cluster right-sizing
-   Own **security posture** — RBAC, secret management (Key Vault / Secrets Manager), prompt-injection risk mitigation, auditability for regulated pharma
-   Partner with data engineers, AI engineers, and Finance business stakeholders to move systems from prototype to reliable production
-   Implement **evaluation frameworks** for AI systems in production — regression testing, adversarial testing, accuracy tracking, hallucination monitoring

## Requirements

**Must-Have Experience**

-   **5+ years in cloud/DevOps/MLOps engineering** on AWS, Azure, or GCP
-   **Production deployment of ML or GenAI systems** — CI/CD, containerization (Docker/Kubernetes), infrastructure-as-code (Terraform)
-   **MLOps tooling** — MLflow, SageMaker Pipelines, Azure ML Pipelines, or equivalent
-   **LLM/GenAI operational experience** — observability tools (LangSmith, Weights & Biases, or equivalent), cost monitoring, latency optimization, prompt/model versioning
-   **Cloud-native AI platforms** — hands-on with at least one of: AWS Bedrock, SageMaker, Azure AI Foundry, Azure OpenAI, Vertex AI
-   **Python, Bash, and infrastructure scripting** — strong
-   **Security and governance in regulated environments** — RBAC, secrets, audit, compliance

**Nice to Have**

-   Pharma, life sciences, or regulated financial services domain
-   Experience operating **agentic AI systems** in production — multi-agent orchestration, tool-calling, human-in-the-loop workflows
-   LangChain, LangGraph, CrewAI, AutoGen, or Semantic Kernel operational experience
-   Kubernetes-native ML platforms (Kubeflow, Ray)
-   Snowflake or Databricks operational experience (compute governance, cost management)
-   Certifications: AWS/Azure ML Engineer, Kubernetes CKA/CKAD, Terraform Associate

**What We're NOT Looking For**

-   **Data Scientists** or research engineers — this is a production platform role
-   **Application developers with light DevOps exposure** — need real MLOps/cloud engineering depth
-   **Pure infra engineers with no AI/ML operational experience** — need to understand what makes LLM systems different (evals, hallucinations, prompt versioning, RAG grounding)
