# Data Engineer

**Company:** [Ocean Infinity](https://jobs.workable.com/companies/aRQ2q8yapu851mu6dpxJ42.md)
**Location:** London, United Kingdom
**Workplace:** hybrid
**Department:** Ocean Infinity

[Apply for this job](https://jobs.workable.com/view/0357ac0d-3245-404a-b1ae-46aa16db4c0b)

## Description

At **Ocean Infinity**, we're on a bold mission to subsea data using innovative technology.

Using cutting-edge robotics, autonomous technology and world-class software, we're transforming how complex operations are carried out at sea. Our uncrewed systems are making ocean exploration and operations **safer, smarter and more sustainable**, helping unlock the vast potential of our oceans while reducing environmental impact. 🌍

This isn't a vision for the future. **It's happening right now.  
  
**As we continue to push the boundaries of innovation and redefine what's possible in the maritime industry, we're looking for exceptional talent to join our fast-growing team.

Ocean Infinity is seeking a motivated and technically capable **Data Engineer** to join our team. The successful candidate will focus on building and implementing scalable data pipelines and automation solutions that transform complex manual workflows into reliable, production-grade data systems. You will contribute to delivering robust, production-ready data infrastructure, working across data, product, and engineering teams to apply engineering standards and provide high-quality curated datasets that power analytics, operational systems, and AI initiatives across the company.

## Requirements

**What you will do! 🚀**

-   Develop and implement scalable data pipelines that automate manual data processes and improve reliability and efficiency across the organization.
-   Build and maintain end-to-end data transformation workflows (Bronze → Silver → Gold), focusing on implementing business logic, data validation, and reusable transformation components.
-   Develop robust Python-based data processing solutions, including scripts, services, and automation tools that support ingestion, transformation, and data delivery.
-   Implement data integration pipelines across a variety of sources, including operational databases, external APIs, and event/streaming data, using both batch and near-real-time approaches.
-   Develop and maintain data quality checks, validation rules, and testing frameworks to ensure accuracy, consistency, and reliability of data products.
-   Build backend data services and lightweight APIs that expose curated datasets to analytics tools, reporting systems, and downstream applications.
-   Collaborate closely with data analysts, product teams, and AI engineers to translate data requirements into efficient, production-ready implementations.
-   Optimize and refactor existing pipelines and processes to improve performance, reliability, and maintainability, with a strong focus on code quality and automation.
-   Follow and contribute to engineering best practices, including version control standards, modular pipeline design, CI/CD for data workflows, and documentation of implemented solutions.
-   Participate in code reviews and pairing, learning from more experienced engineers and sharing implementation knowledge and reusable patterns with the team.

**What we look for!🔍**

-   A degree in Computer Science, Mathematics, Engineering, or a related field, or equivalent practical experience.
-   Minimum 2 years of experience as a Data Engineer, Backend Engineer, or Software Engineer working with data-intensive systems and data pipelines.
-   Solid software engineering skills in Python, with experience writing reliable, maintainable data processing code and automation workflows.
-   Good SQL skills, with an understanding of query optimization and working with analytical datasets.
-   Experience building and maintaining data pipelines and transformation workflows, ideally in cloud-based environments.
-   Hands-on experience with at least one cloud platform (e.g., AWS, GCP, or Azure), particularly for building and running data pipelines, storage systems, and compute workloads.
-   Familiarity with data lake and lakehouse environments, including object storage systems (e.g., S3 or equivalent) and structured transformation layers (e.g., Bronze/Silver/Gold patterns).
-   Understanding of unstructured and semi-structured data storage patterns (e.g., JSON, logs, event data, files in object storage) and how to process them efficiently.
-   Experience with relational and/or NoSQL databases (e.g., Postgres, MongoDB, Redis), including schema design and indexing.
-   Working knowledge of containerization technologies such as Docker.
-   Good understanding of data modelling principles and how to structure datasets for analytics and downstream consumption.
-   Exposure to distributed systems concepts and large-scale data processing patterns, with emphasis on practical implementation and debugging.
-   Strong problem-solving skills with a focus on reliability, performance, and automation in production systems.
-   Ability to collaborate effectively with analysts, product teams, and engineers, translating requirements into robust, testable implementations.
-   Interest in improving and refactoring existing pipelines and workflows to increase reliability, maintainability, and automation.
-   Comfortable working in modern engineering practices including Git-based workflows, code reviews, CI/CD pipelines, and automated testing for data systems.
-   Exposure to workflow orchestration tools such as Airflow, Prefect, or similar, including scheduling, dependency management, retries, and observability.

**Nice to have! ✨**

-   Experience working with telemetry, sensor, location, and operational technology (OT) data workloads.
-   Experience processing and managing large-scale video, image, and unstructured media datasets, including streaming ingestion and analytics workflows.
-   Understanding of feature engineering patterns and data preparation workflows supporting machine learning systems.
-   Experience working with maritime, fleet, vessel operations, logistics, or industrial operational domains, including telematics, tracking, asset monitoring, or operational analytics use cases.

## Benefits

_**\*Salary: London £65,000 / Porto up to 55,000 EUR Per Annum\***_
