# Senior Data & Document Ingestion Engineer (OCR / RAG) - REMOTE

**Company:** [Gramian Consulting Group](null/companies/kANY7hHLXDH7fUqyifRLmf.md)
**Location:** Remote
**Workplace:** remote
**Employment type:** Contract
**Department:** Talent Solutions

[Apply for this job](null/view/8e64393d-9076-436a-9a20-1f5bfa69fa16)

## Description

### About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

### About the Role

Our client is a Big4 Consultancy group that works with leading financial institutions on **AI-driven transformation, automation, advanced analytics, and financial crime prevention**. Their work spans intelligent fraud detection, AML/KYC modernization, autonomous workflows, enterprise AI platforms, and the secure industrialization of AI in highly regulated environments.

We are looking for a **Senior Data & Document Ingestion Engineer** to build robust pipelines for processing high-volume unstructured insurance content. The role focuses on **OCR, document parsing, ingestion pipelines, text normalization, semantic chunking, metadata extraction, and retrieval-ready data preparation** for downstream AI systems.

**CONTRACT:** Contractor assignment, expected October 2026 – July 2027, with extension to other projects (and retention rate)

**COMMITMENT:** Full-time

**LOCATIONS:** REMOTE 100%, Europe-based

**PROCESS:** Initial qualification followed by technical and client interviews

**NOTES: Fluent English is required. Must be able to work in EU.**

### Responsibilities

-   Design and build scalable **document ingestion pipelines** for PDFs, scans, emails, and office documents.
-   Integrate and optimize **OCR and document extraction technologies** for high-accuracy text and layout extraction.
-   Build workflows for text cleaning, normalization, semantic chunking, and metadata tagging.
-   Process unstructured formats including PDF, Word, Excel, and PowerPoint.
-   Develop connectors for enterprise sources such as **SharePoint and email systems**.
-   Design data schemas and retrieval mechanisms for downstream AI and **RAG** use cases.
-   Build validation and monitoring loops to detect low-confidence OCR or extraction results.
-   Ensure ingestion pipelines meet enterprise security, reliability, and latency requirements.
-   Implement logging, testing, and operational monitoring across data-processing workflows.
-   Apply Git, CI/CD, and software-engineering best practices to pipeline development.

## Requirements

-   Approximately **5–10 years of professional data engineering or backend/data-platform experience**.
-   Strong hands-on experience with **Python and SQL**.
-   Proven experience building **data ingestion and document-processing pipelines**.
-   Hands-on experience processing unstructured documents such as PDF, Word, Excel, PPT, scans, or emails.
-   Experience with **OCR/document extraction tools** such as AWS Textract or equivalent.
-   Professional experience building data-processing pipelines on **public cloud platforms**.
-   Experience with AWS services such as **S3, Step Functions, and CloudWatch**, or comparable cloud services.
-   Strong development practices including **Git, CI/CD, and automated testing**.

**Preferred Qualifications**

-   Experience with **Azure, AWS, or Databricks** in enterprise data environments.
-   Experience with **vector databases, embeddings, or RAG architectures**.
-   Experience designing connectors to SharePoint, email, or other enterprise content systems.
-   Background in insurance, financial services, or regulated-data environments.
