Senior Data Engineer (Databricks)

Cluepoints
Cluepoints

Data Science

Belgium

Posted on Sep 21, 2026
At CluePoints, we're redefining how clinical trials are run. As a leading providerof Risk-Based Quality Management and Data Quality Oversight software, we useadvanced statistics, artificial intelligence and machine learning to improve thequality, accuracy and integrity of clinical trial data — turning artificialintelligence into human intelligence.We're an ambitious, fast-growing technology company with a dynamic anddiverse international team. Collaboration, flexibility and continuous learning arepart of our DNA. Guided by our values of Care, Passion and Smart Disruption,we're united by a shared mission: to create smarter ways to run efficient clinicaltrials and deliver AI-powered insights that improve human outcomes worldwide.

The Role
We're looking for a hands-on Senior Data Engineer to own the technical design, reliability and evolution of our Databricks data foundation. You'll work alongside Clinical Data Insights Analysts — colleagues with strong business, clinical and BI expertise who also write SQL and Python — complementing their skills with deeper engineering, automation and production-grade rigour.
  • 5+ years in data engineering or data-platform engineering.
  • Strong hands-on Databricks experience (or comparable cloud lakehouse platform).
  • Advanced SQL and strong Python / PySpark skills.
  • Production-grade pipeline experience — design, build, test, deploy and maintain.
  • Solid understanding of Delta Lake, medallion architecture, and OLAP/OLTP transformation.
  • Experience with data governance, Unity Catalog (or equivalent), automated quality controls and CI/CD.
  • At least one major cloud platform: Azure, AWS or GCP.
  • Practical use of generative AI in engineering workflows, with solid critical judgement on its outputs.
  • High autonomy — you take ownership from design through deployment, monitoring and maintenance.
  • Collaborative by nature; comfortable working directly alongside analysts who also write code.

Nice to Have:
  • Lakeflow Jobs / Lakeflow Declarative Pipelines; Databricks Asset Bundles.
  • Databricks AI/BI, Genie, conversational analytics or AI agent experience.
  • Databricks Data Engineer certification.
  • Experience in a regulated environment (clinical trials, life sciences or similar).

Tech Stack:
Databricks · Delta Lake · Unity Catalog · Lakeflow Jobs & Declarative Pipelines · Medallion architecture · SQL / Python / PySpark · Microsoft Azure · Git / CI/CD · Databricks Asset Bundles · Databricks AI/BI · Genie
  • Own the architecture, reliability and governance of the team's Databricks lakehouse (Delta Lake, medallion architecture, Unity Catalog).
  • Design and operate batch and streaming ingestion pipelines from the CluePoints platform and other business systems (e.g. Zendesk).
  • Co-develop SQL, Python and PySpark transformations with the analyst team; review and optimize for performance, scalability and cost.
  • Productionize analytical and AI prototypes — adding testing, orchestration, monitoring, error handling and deployment controls.
  • Build automated data-quality controls, monitoring and alerting; investigate and resolve pipeline incidents.
  • Implement technical governance: naming conventions, metadata, lineage, tagging and access controls via Unity Catalog.
  • Apply AI and automation across the data lifecycle — code development, documentation, metadata classification and inconsistency detection.
  • Collaborate with Engineering on source-system access and with Product when Databricks insights are candidates for platform integration.