# Quantexa Data Engineer

**Company:** [Unison Group](https://jobs.workable.com/companies/deKpPhPMtQZga7XoPG1tSo.md)
**Location:** Singapore, Singapore
**Workplace:** on site
**Department:** Bharat

[Apply for this job](https://jobs.workable.com/view/c61fd3b4-4333-49ea-8a7a-839c8763a7f6)

## Description

### Role Overview

We are seeking a talented and experienced **Data Engineer** with strong expertise in **Quantexa, Hadoop, Scala, Apache Spark, Elasticsearch, OpenShift Container Platform (OCP), and DevOps practices**.

The successful candidate will be responsible for designing, developing, optimizing, and maintaining scalable big data solutions using **Apache Spark, Scala, Hadoop, and Elasticsearch**. The role involves collaborating with cross-functional teams to build efficient data processing pipelines and search applications.

**Knowledge and experience in the Compliance / AML domain will be an added advantage.**

Key Responsibilities

-   Design, develop, and implement **Spark/Scala applications** and data processing pipelines for large volumes of structured and unstructured data.
-   Implement **data transformation, aggregation, enrichment, and computation** processes to support analytics and machine learning initiatives.
-   Collaborate with cross-functional teams to understand data requirements and translate them into effective data engineering solutions.
-   Integrate **Elasticsearch with Spark** for efficient data indexing, querying, and retrieval.
-   Implement transformations and aggregations using **Spark RDDs, DataFrames, Datasets, and Spark SQL**.
-   Develop scalable, reliable, and fault-tolerant Spark applications following industry best practices and coding standards.
-   Optimize and tune **Spark jobs and Elasticsearch queries** to improve performance, scalability, and resource utilization.
-   Monitor job performance, identify bottlenecks, troubleshoot issues, and implement appropriate optimizations.
-   Troubleshoot and resolve issues related to **data processing, data quality, Spark performance, and Elasticsearch integration**.
-   Ensure data quality, consistency, accuracy, and integrity throughout the data processing lifecycle.
-   Design and deploy data engineering solutions on **OpenShift Container Platform (OCP)** using containerization and orchestration technologies.
-   Optimize data engineering workflows for containerized environments and efficient resource utilization.
-   Collaborate with DevOps teams to streamline deployments and implement **CI/CD pipelines**.
-   Implement **data governance, data lineage, and metadata management** practices to ensure data accuracy, traceability, and compliance.
-   Implement monitoring and logging mechanisms to ensure the health, availability, and performance of data infrastructure.
-   Monitor and optimize end-to-end data pipeline performance and implement required enhancements.
-   Document data engineering processes, workflows, architecture, and infrastructure configurations for knowledge sharing and future reference.

Requirements

-   Bachelor’s or Master’s degree in **Computer Science, Software Engineering, Information Technology, or a related field**.
-   **Quantexa Certified Data Engineer / Data Architect** with hands-on experience and strong proficiency in the Quantexa platform.
-   Proven experience as a **Data Engineer** working with Hadoop, Spark, and large-scale data processing technologies.
-   Strong proficiency in **Scala** and familiarity with functional programming concepts.
-   In-depth understanding of **Apache Spark architecture, RDDs, DataFrames, Datasets, and Spark SQL**.
-   Strong expertise in **Hadoop ecosystem technologies**, including HDFS, Hive, Pig, and related tools.
-   Hands-on experience with **Elasticsearch**, including data indexing, search applications, data modeling, indexing strategies, and query optimization.
-   Experience with **OpenShift Container Platform (OCP)** and Kubernetes-based container orchestration.
-   Strong programming skills in **Scala, Python, Java, and/or Spark**.
-   Good understanding of **DevOps practices, CI/CD pipelines, and infrastructure automation**.
-   Experience with tools such as **Docker, Jenkins, Ansible, and Bitbucket**.
-   Experience with **distributed computing, parallel processing, and large-scale datasets**.
-   Strong experience in **performance tuning and optimization** of Spark applications and Elasticsearch queries.
-   Experience with **Git** and collaborative software development workflows.
-   Strong analytical and problem-solving skills with the ability to troubleshoot complex technical issues.
-   Excellent communication and collaboration skills with the ability to work effectively with cross-functional teams.
-   Experience with **Grafana, Prometheus, and Splunk** will be an added advantage.
-   Exposure to cloud platforms such as **AWS, Azure, or GCP** and their data services will be a plus.
-   Knowledge or experience in the **Compliance / AML domain** will be an added advantage.
