• Jobs
  • >
  • Senior Data Engineer (Databricks)

Senior Data Engineer (Databricks)

  • Permanent
  • Full-time
  • Hybrid (Santa Clara, CA, United States)

About the Role

We are looking for a Senior Data Engineer with a minimum of 5+ years of professional experience to design, build, and maintain scalable, high-performance data platforms on Databricks that power our analytics and machine learning initiatives. You will work in a distributed, data-intensive environment and collaborate with cross-functional teams to ensure data accuracy, reliability, and security.

Key Responsibilities

  • Design, develop, and maintain scalable data processing pipelines and workflows on the Databricks Lakehouse Platform using Apache Spark and PySpark.

  • Build and maintain microservices in Python that serve data-driven features in production.

  • Develop internal tools to support CI/CD pipelines, experiment tracking, and data versioning (e.g., Databricks Asset Bundles, MLflow, Delta Lake time travel).

  • Collect, process, and integrate large datasets from multiple sources — databases, file systems, and APIs — into the Lakehouse using Auto Loader, CDC, and streaming ingestion patterns.

  • Ensure data integrity, consistency, and quality through robust validation and monitoring, leveraging Delta Live Tables expectations and Lakehouse monitoring capabilities.

  • Optimize data systems for performance, scalability, and high availability, including Spark job tuning, Delta table optimization (OPTIMIZE, Z-ORDER/liquid clustering), and efficient cluster/compute management.

  • Implement best practices for data security, access control, and privacy using Unity Catalog for centralized governance, fine-grained permissions, and lineage.

  • Collaborate with data scientists, analysts, and engineers to support analytics and ML workflows across notebooks, Databricks SQL, and MLflow.

  • Lead complex migration initiatives involving the transition from on-premise Cloudera (HDFS, Hive, Impala, HBase) environments to the Databricks Lakehouse, ensuring zero data loss and minimal downtime.

  • Architect and build a large-scale, greenfield data platform on Databricks from the ground up, including end-to-end infrastructure setup, data ingestion layers, transformation frameworks (medallion architecture), and high-throughput integrations.

Requirements

  • Strong understanding of distributed systems and modern Lakehouse data architectures.

  • Proven, hands-on experience with large-scale data platform migrations, specifically transitioning from Cloudera (CDH/HDP) to Databricks. This includes re-architecting Hive/Impala workloads to Spark SQL and Delta Lake, migrating HDFS data to cloud object storage, and refactoring legacy ETL processes to leverage Databricks-native features (e.g., Delta Live Tables, Auto Loader, Databricks Workflows).

  • Deep technical expertise in building and orchestrating a high-performance data platform from scratch on Databricks, including complex integrations (Reverse ETL, CDC, API ingestions) using Auto Loader, Structured Streaming, Delta Sharing, and Unity Catalog.

  • 5+ years of professional experience in software engineering or data engineering.

  • Strong software engineering skills with Python in large-scale, high-performance production environments.

  • Hands-on experience with Spark/PySpark and the broader big data ecosystem.

Nice to Have

  • Databricks certifications (e.g., Databricks Certified Data Engineer Professional).

  • Experience with Databricks Asset Bundles or Terraform for infrastructure-as-code deployment of Databricks workspaces and jobs.

  • Familiarity with Photon, Databricks SQL Warehouses, and serverless compute.

About Opplane

Opplane specializes in providing advanced data-focused solutions for financial services, telecommunication, and reg-tech to accelerate their digital transformation journey. Opplane's leadership team is comprised of Silicon Valley serial entrepreneurs and experienced executives. Its expertise comes from years of specific industry experience at some of the world's top companies, such as PayPal, Xerox Parc, Amazon, Wells Fargo, SoFi in the areas of product management, data technology, data governance, data privacy, security, machine learning, and risk management.

🌍 Global & Multicultural – Diverse perspectives, global collaboration (US, Portugal, India and Singapore offices)

Startup Energy – Fast-moving, impact-driven environment

💪 Ownership Mindset – Engineers own what they build

🤝 Collaborative & Friendly – Open, curious, and supportive culture

|
|
Powered by Factorial
Build my own jobs page