Databricks Lakehouse & Data Engineering
Service description
I build production data platforms on Databricks, and I do it end to end so your analysts and data scientists can trust every table they open. My focus is the lakehouse: I design Spark pipelines that ingest from databases, APIs, files and streaming sources, then land the data as Delta tables that are versioned, ACID-safe and fast to query. I follow the medallion architecture, moving raw data through bronze, silver and gold layers so that cleansing, deduplication and business logic each live in a clear, testable stage instead of one tangled script nobody wants to touch.
My engagements usually cover the full lifecycle. I set up workspaces, clusters and job compute with sensible autoscaling so you are not paying for idle machines, and I orchestrate everything with Databricks Workflows or your existing scheduler, complete with retries, alerting and data-quality checks. On the governance side I implement Unity Catalog: a single catalog for tables, volumes and models, row and column level access, lineage and audit so security and compliance are handled by design rather than bolted on later. I also tune performance where it matters, using partitioning, Z-ordering, liquid clustering and file compaction to keep both cost and query latency down.
What you get is a maintainable foundation, not a one-off demo. I write modular, tested PySpark and SQL, document the model and the run books, and hand over pipelines your own team can extend. Whether you are migrating a legacy warehouse, standing up your first lakehouse, or preparing clean feature tables for machine learning, I bring the pragmatism to ship something reliable on schedule.
— Spark and Delta Lake pipelines with the bronze/silver/gold model
— Workflows orchestration, CI/CD and data-quality gates
— Unity Catalog governance, lineage and access control
— Performance and cost tuning for clusters and queries
My engagements usually cover the full lifecycle. I set up workspaces, clusters and job compute with sensible autoscaling so you are not paying for idle machines, and I orchestrate everything with Databricks Workflows or your existing scheduler, complete with retries, alerting and data-quality checks. On the governance side I implement Unity Catalog: a single catalog for tables, volumes and models, row and column level access, lineage and audit so security and compliance are handled by design rather than bolted on later. I also tune performance where it matters, using partitioning, Z-ordering, liquid clustering and file compaction to keep both cost and query latency down.
What you get is a maintainable foundation, not a one-off demo. I write modular, tested PySpark and SQL, document the model and the run books, and hand over pipelines your own team can extend. Whether you are migrating a legacy warehouse, standing up your first lakehouse, or preparing clean feature tables for machine learning, I bring the pragmatism to ship something reliable on schedule.
— Spark and Delta Lake pipelines with the bronze/silver/gold model
— Workflows orchestration, CI/CD and data-quality gates
— Unity Catalog governance, lineage and access control
— Performance and cost tuning for clusters and queries
Contact the freelancer
Order the service or ask the freelancer a question.
Freelancer contacts
E-mailShow
