# AI Application Engineer (Part-Time)

**Company:** [Workana](http://jobs.workable.com/companies/wSyiKiaa7abTa8N4S2cpd7.md)
**Location:** Remote
**Workplace:** remote
**Employment type:** Full-time
**Department:** Workana Premium

[Apply for this job](http://jobs.workable.com/view/79d48785-50d1-4ef2-8b24-d6a4b0dc3dcb)

## Description

**Client:** medxprts.ai  
**Location:** Remote  
**Type:** Part-time, with potential to transition to Full-time  
**Schedule:** U.S. timezone overlap required

### **Description**

Medxprts.ai is building an AI-powered platform for the legal and healthcare space, using LLMs, agentic workflows, and automation to create production-grade applications.

We are seeking an AI Engineer - LLM Fine-Tuning: a hands-on engineer who has personally trained and fine-tuned open-weight models, built training infrastructure, engineered datasets from messy documents, and established rigorous evaluation and preference/feedback training pipelines. This role works closely with engineering teams to ship models and integrated features into production.

### **Responsibilities**

-   Design, implement, and run LLM fine-tuning experiments (LoRA/QLoRA and full SFT) on open-weight models (e.g., Llama, Mistral, Qwen) and ship trained models into product workflows.
-   Build and maintain training infrastructure using PyTorch and Hugging Face tooling (Transformers, PEFT, TRL/Axolotl), including multi-GPU training orchestration (DeepSpeed/FSDP) on AWS or GCP.
-   Engineer datasets from real-world unstructured sources (long PDFs, medical/legal records), performing deduplication, filtering, contamination checks, and train/eval splits.
-   Create evaluation harnesses tailored to domain needs: held-out test sets, LLM-as-judge with human calibration, regression tests across model versions, and automated monitoring for model drift.
-   Implement preference/feedback training workflows (DPO/RLHF-style or similar) to learn from expert corrections (doctor-in-the-loop), and integrate feedback loops into model retraining pipelines.
-   Collaborate with backend/frontend engineers to integrate fine-tuned models into services, optimize inference latency/cost, and support production deployments.
-   Participate in PR reviews, release processes, incident debugging, and continuous improvement of training and deployment tooling.
-   Ensure secure, compliant handling of sensitive data (HIPAA-awareness is highly preferred) during dataset preparation and model training.

## Requirements

-   Hands-on experience fine-tuning LLMs: personally trained or fine-tuned open-weight models using LoRA/QLoRA and full SFT; able to explain trade-offs and provide at least one shipped example.
-   Training infrastructure experience: PyTorch + Hugging Face ecosystem (Transformers, PEFT, TRL/Axolotl), multi-GPU training knowledge (DeepSpeed or FSDP), and running training workloads on AWS or GCP.
-   Dataset engineering expertise: built instruction/preference datasets from messy, unstructured documents; practical knowledge of deduplication, filtering, train/eval splits, and contamination prevention.
-   Strong evaluation discipline: designed domain-specific evaluation harnesses beyond standard benchmarks, including human-calibrated judge setups and regression testing.
-   Practical experience with preference/feedback learning methods (DPO, RLHF-style workflows, or equivalent) and integrating expert feedback into model updates.
-   Solid software engineering fundamentals: production workflows (Git, PRs), debugging, testing, deployment experience, and maintainable code.
-   Experience with APIs, databases, and service integration for model inference.
-   Ability to work independently, learn quickly, and follow technical direction.
-   Good English communication skills and availability to overlap with U.S. working hours.  
    

### **Nice to Have**

-   Experience with long-context handling strategies for very large documents (retrieval-aware training, context extension, RAG for multi-thousand-page sources).
-   Model deployment and inference optimization: quantization (GPTQ/AWQ), vLLM/TGI serving, batching/throughput tuning, latency and cost optimization.
-   Familiarity with vector databases, advanced RAG pipelines, MCPs, or n8n-style workflow automation tools.
-   Knowledge of Docker, CI/CD, Kubernetes/EKS, or serverless infrastructure.
-   Prior experience in healthcare, legaltech, HIPAA-aware processes, or other regulated/data-sensitive environments.
-   Public portfolio, GitHub, or examples of shipped LLM/agentic applications and fine-tuning projects.

## Benefits

-   Fully remote role.
-   Opportunity to work on real AI products in the legal and healthcare domain.
-   High ownership and autonomy.
-   Performance-based incentives and outcome-driven bonuses.
-   Potential to grow into a long-term, full-time role.
