Back to jobs
Senior Data Engineer (Databricks / Spark Streaming)
AptonetAnywherePosted 2w ago
Senior Data Engineer
Job Description
Senior Data Engineer (Databricks / Spark Streaming)
Location:
Remote (Mexico)
About the Role
We are looking for a Senior Data Engineer to design, build, and scale our data infrastructure with a focus on real-time and batch processing pipelines. You will work extensively with Databricks and Apache Spark Structured Streaming to deliver reliable, high-throughput data platforms that power analytics, machine learning, and business-critical applications. This is a fully remote position open to candidates based in Mexico.
What You'll Do
Design, build, and maintain scalable ETL/ELT pipelines using Databricks and Apache Spark
Develop and optimize real-time streaming pipelines using Spark Structured Streaming (Kafka, Kinesis, Event Hubs, or similar)
Architect and manage Delta Lake tables, implementing best practices for schema evolution, partitioning, and performance tuning
Build and maintain data pipelines across the medallion architecture (bronze/silver/gold layers)
Collaborate with data scientists, analysts, and software engineers to understand data needs and deliver production-grade solutions
Implement CI/CD workflows for data pipelines using tools such as Databricks Asset Bundles, GitHub Actions, or Azure DevOps
Monitor, troubleshoot, and optimize pipeline performance, cost, and reliability
Establish and enforce data quality, governance, and observability standards (e.g., using Unity Catalog, Great Expectations, or similar)
Mentor junior engineers and contribute to engineering best practices and documentation
Participate in architecture discussions and provide technical guidance on data platform strategy
Required Qualifications
8+ years of experience in data engineering, with a strong focus on distributed data processing
Hands-on production experience with
Databricks
(clusters, jobs, workflows, Unity Catalog)
Strong expertise in
Apache Spark
, including
Spark Structured Streaming
for real-time data processing
Proficiency in
Python
and/or
Scala
for data pipeline development
Advanced
SQL
skills for data transformation and optimization
Experience with
Delta Lake
and lakehouse architecture concepts
Experience with AWS — Databricks-supported environments
Familiarity with message streaming systems such as
Kafka
,
Kinesis
, or
Event Hubs
Experience with orchestration tools (e.g., Databricks Workflows, Airflow)
Solid understanding of data modeling, partitioning strategies, and performance optimization
Experience with version control (Git) and CI/CD practices for data engineering
Nice to Have
Databricks certifications (e.g., Databricks Certified Data Engineer Professional)
Experience with infrastructure-as-code (Terraform)
Experience with dbt for transformation workflows
Familiarity with MLOps or feature store concepts
Experience working in a fully remote, distributed team environment
Prior experience in a regulated industry (finance, healthcare, etc.)
What We're Looking For
Strong communication skills and comfort working async with distributed, cross-functional teams
A proactive, ownership-driven mindset — comfortable taking projects from design through production
Ability to balance pragmatic delivery with long-term architectural quality
Fluent English and Spanish (written and verbal)