Back to jobs

Senior Data Engineer (Databricks / Spark Streaming)

AptonetAnywherePosted 2w ago
Senior Data Engineer
Apply on company site

Job Description

Senior Data Engineer (Databricks / Spark Streaming) Location: Remote (Mexico) About the Role We are looking for a Senior Data Engineer to design, build, and scale our data infrastructure with a focus on real-time and batch processing pipelines. You will work extensively with Databricks and Apache Spark Structured Streaming to deliver reliable, high-throughput data platforms that power analytics, machine learning, and business-critical applications. This is a fully remote position open to candidates based in Mexico. What You'll Do Design, build, and maintain scalable ETL/ELT pipelines using Databricks and Apache Spark Develop and optimize real-time streaming pipelines using Spark Structured Streaming (Kafka, Kinesis, Event Hubs, or similar) Architect and manage Delta Lake tables, implementing best practices for schema evolution, partitioning, and performance tuning Build and maintain data pipelines across the medallion architecture (bronze/silver/gold layers) Collaborate with data scientists, analysts, and software engineers to understand data needs and deliver production-grade solutions Implement CI/CD workflows for data pipelines using tools such as Databricks Asset Bundles, GitHub Actions, or Azure DevOps Monitor, troubleshoot, and optimize pipeline performance, cost, and reliability Establish and enforce data quality, governance, and observability standards (e.g., using Unity Catalog, Great Expectations, or similar) Mentor junior engineers and contribute to engineering best practices and documentation Participate in architecture discussions and provide technical guidance on data platform strategy Required Qualifications 8+ years of experience in data engineering, with a strong focus on distributed data processing Hands-on production experience with Databricks (clusters, jobs, workflows, Unity Catalog) Strong expertise in Apache Spark , including Spark Structured Streaming for real-time data processing Proficiency in Python and/or Scala for data pipeline development Advanced SQL skills for data transformation and optimization Experience with Delta Lake and lakehouse architecture concepts Experience with AWS — Databricks-supported environments Familiarity with message streaming systems such as Kafka , Kinesis , or Event Hubs Experience with orchestration tools (e.g., Databricks Workflows, Airflow) Solid understanding of data modeling, partitioning strategies, and performance optimization Experience with version control (Git) and CI/CD practices for data engineering Nice to Have Databricks certifications (e.g., Databricks Certified Data Engineer Professional) Experience with infrastructure-as-code (Terraform) Experience with dbt for transformation workflows Familiarity with MLOps or feature store concepts Experience working in a fully remote, distributed team environment Prior experience in a regulated industry (finance, healthcare, etc.) What We're Looking For Strong communication skills and comfort working async with distributed, cross-functional teams A proactive, ownership-driven mindset — comfortable taking projects from design through production Ability to balance pragmatic delivery with long-term architectural quality Fluent English and Spanish (written and verbal)