We operate a production data lakehouse built on Apache Iceberg, Trino, dbt, and Dagster. The core architecture is established and governed through a documented review process; we are hiring a Senior Data Engineer to build, operate, and optimize the platform — and to help evolve it through well-argued design contributions. You will design data models and pipeline architectures within the platform, own production reliability, and deliver trustworthy datasets for AI, analytics, and backend applications.
Technology Stack
- Storage & table format: S3-compatible object storage, Apache Iceberg
- Catalog: Lakekeeper (Iceberg REST catalog)
- Query engine: Trino
- Transformation: dbt
- Orchestration: Dagster
- Batch ingestion: Airbyte
- CDC / streaming: Debezium, Apache Kafka
- Metadata, lineage & governance: OpenMetadata
- BI & visualization: Superset
- Vector storage (AI workloads): pgvector, Qdrant
- Identity & access control: Keycloak, OPA
- Primary source systems: PostgreSQL (NestJS backend services), Redis, Elasticsearch
- Runtime: Kubernetes
Key Responsibilities
-
Design data models, dataset layouts, and pipeline architectures for new data domains within the established platform architecture.
-
Build and maintain ELT pipelines (Airbyte ingestion → dbt transformations → curated marts) orchestrated in Dagster.
-
Operate Apache Iceberg tables in production: partitioning strategy, file compaction, snapshot expiration and retention, schema evolution, and time-travel-based reprocessing.
-
Operate and extend CDC ingestion with Debezium and Kafka: connector configuration, PostgreSQL logical replication (WAL / replication slots), schema registry, idempotent sinks, backfill and replay procedures.
-
Tune Trino performance: storage layout, table statistics, query plans, resource groups.
-
Engineer data quality: dbt tests, data contracts, freshness / volume / schema-drift monitoring, duplicate detection, missing-value handling; maintain lineage and metadata in OpenMetadata.
-
Ensure reproducibility and traceability through Iceberg snapshots and versioned dbt models.
-
Build feature and embedding pipelines serving AI services (entity resolution, deduplication, vector stores).
-
Implement governed pipelines: multi-tenant isolation, data classification, access-control integration (Keycloak / OPA), audit logging, and compliance with Vietnamese data-residency requirements.
-
Contribute to architectural evolution: evaluate alternatives, write ADRs, participate in design reviews (e.g., streaming expansion, catalog or vector-store migration).
-
Own production reliability: SLOs, monitoring, runbooks, incident triage, and root-cause analysis for data pipelines.
-
Collaborate with Backend, AI, and Analytics teams to deliver reliable data products.
Requirements
-
Bachelor's degree in Computer Science, Information Technology, or a related field — or equivalent practical experience.
-
5+ years as a Data Engineer with end-to-end production ownership.
-
Strong SQL and Python.
-
Hands-on experience designing Data Lake / Lakehouse and Data Warehouse architectures and data models — and the judgment to work effectively within an established architecture.
-
Proven ELT/ETL pipelines across heterogeneous sources (databases, APIs, files, event streams).
-
Production experience with at least one open table format — Apache Iceberg strongly preferred; Delta Lake or Hudi acceptable with commitment to transition.
-
Hands-on experience with a distributed SQL engine (Trino / Presto or comparable) and dbt or an equivalent transformation framework.
-
Production orchestration experience — Dagster preferred; Airflow / Prefect acceptable with commitment to transition.
-
CDC experience (Debezium or equivalent) and working knowledge of PostgreSQL logical replication.
-
Experience building feature pipelines, embedding pipelines, or ML-serving datasets.
-
Data quality engineering: validation, duplicate detection, missing-value handling.
-
English reading proficiency for technical documentation.
Preferred Skills (strong plus)
-
Iceberg REST catalogs (Lakekeeper, Polaris, or Nessie); OpenMetadata; Superset.
-
Kafka operations and schema registry management.
-
Running data workloads on Kubernetes.
-
Formal feature stores (e.g., Feast) or NLP data preparation for AI/ML.
-
Spark for batch processing.
-
Experience in regulated or data-residency-constrained environments (e.g., Vietnam PDPL 91/2025, Decree 53/2022).
How We Work
-
Architecture evolves through documented ADRs and a change-classification process: standard changes ship fast, architectural changes get review.
-
Documentation-first culture: conclusions up front, self-contained specifications, explicit assumptions.