inDrive · Almaty, Kazakhstan · по договорённости
We are looking for a Data Engineer to join one of the Data Platform teams that works with the Marketing, Growth, partner, and financial data domains. You will be working with cutting edge cloud technologies (GCP, AWS, BigQuery, Databricks, K8s) and building a large scale data infrastructure for analytics, machine learning, and streaming/CDC data delivery.
Build and operate batch and streaming ingestion into a layered BigQuery DWH (raw → ODS → data marts) using Airflow, Debezium CDC over Kafka with protobuf, Pub/Sub, and Dataflow
Integrate external data sources end-to-end — marketing platforms (GA4, AppsFlyer, TikTok/Meta/Google Ads), payment providers, S3 buckets, and third-party APIs — including schema contracts, backfills, and reconciliation
Engineer the data platform itself in Python: custom Airflow operators and connectors in a shared ETL framework, Kafka Connect on Strimzi (K8s), Cloud Functions, and API integrations with external providers
Build CI/CD and change-management tooling for BigQuery: GitHub-based test-and-approval flows, SQL migration engines (Liquibase/Flyway/Bytebase), sandbox validation, backup and rollback
Own reliability and correctness of pipelines: idempotency, deduplication, late-data handling, backfill and replay, freshness monitoring and alerting; write integration and unit tests
Drive data governance and compliance: ITGC-compliant change management for BigQuery, IAM and least-privilege access, PII policy tags and DLP, Unity Catalog on Databricks, column-level lineage (OpenMetadata/Dataplex), and disaster-recovery planning
Build internal data tools and platform services for agentic workflows with data — Streamlit apps, Slack bots, LLM-based agents and MCP servers that help teams find and use data
Support analysts and business teams with data requests, fostering data-driven decision-making across the company
Contribute to system design and architecture with the development team
Strong practical Python: clean, well-structured, and tested code for services, tooling, and data pipelines
Solid software design skills (OOP, modularity, design patterns) — we build platform tools for agentic workflows with data and plan to develop data-related backend services, so well-designed code is highly valued
Experience building and operating services in a cloud environment (GCP, AWS or similar): CI/CD, containerization, monitoring and alerting
Familiarity with Kubernetes and Terraform — our infrastructure runs on GCP/K8s
Hands-on experience with DWH-related tasks (BigQuery or another cloud warehouse) and confident working SQL
Clear communication with non-engineering stakeholders — a meaningful share of the work is data requests from analysts and business teams
Demonstrated ability to take ownership of technologies or services and proactively contribute ideas to the team
Nice to have
Advanced SQL: complex queries, window functions, partitioning, clustering, and cost optimization
Experience building reliable pipelines around CDC (e.g., Debezium): idempotency, schema evolution, backfills, and reconciliation
Analytical data modeling skills: table grain, facts vs dimensions, slowly changing dimensions, and metric definitions
Experience with stream processing frameworks such as Flink or Apache Beam/Dataflow
Exposure to data governance and audit compliance (ITGC/SOX), Databricks Unity Catalog, or lineage/catalog tooling (OpenMetadata, Dataplex)
Interest in building LLM-based agents and AI tooling for data
Отклик ведёт на сайт работодателя. Бесплатная регистрация открывает отклик и разбор резюме.