Back to jobs
Remote Senior Software Engineer-Data Platform
AccruetalentNewmarket, ON (+1 other)Posted 1w ago
Senior Software Engineer — TypeScript/PostgreSQL (remote)
Job Description
Hiring the first dedicated engineer for its data platform. Today the platform is nascent, solid backend infrastructure exists, but nobody owns it, and the perception/ML team is being handed data that isn't in the shape they need. This is a founding-level, zero-to-one seat: you'll define how raw maritime signals become clean, consistent, labeled datasets for perception and foundation models.
You'll architect high-throughput pipelines that ingest real-time acoustic and telemetry data, then align, resample, calibrate, and restructure it, not necessarily in the moment (it can be matched up every 30 minutes, hour, or day) but into exactly what the ML team needs. Expect to build backfill/reprocessing frameworks, lineage and versioning for reproducibility, dataset discovery APIs, data-quality instrumentation, and, likely, a labeling tool from scratch. There's real backend infra to build on, but the shape of the platform is yours to define.
What You'll Own
Post-processing pipelines that align, resample, and calibrate multi-sensor data
Backfill/reprocessing frameworks for new filters, syncs, label corrections, and metadata enrichment across historical data
Lineage and versioning to guarantee experiment reproducibility
Dataset discovery + access APIs/SDKs (query by time, region, modality, labels, quality flags)
Data-quality metrics, dashboards, and alerts; canary dataset builds
Storage-layout optimization (columnar formats, compression, chunking, sharding, prefetching)
Likely build a labeling tool and pre-labeling workflows for the perception team
Ramp: month 2, a first working pipeline in place; months 3–6 — iterating and honing it to exactly what the ML/perception team needs, plus the tooling around it
Requirements
Strong data-pipeline architecture, high-throughput, with the ability to architect the system from a blank page
Solid knowledge of at least one cloud provider, preferably AWS
Comfortable deploying pipelines to the cloud
Python + data tooling (PyArrow/Polars/Pandas, NumPy/SciPy), plus one of Go/Rust/TypeScript for services
Bonus: full-stack/generalist range — able to build the tools the ML team needs end-to-end
Execution and Ownership
Comes in and builds day one with minimal hand-holding
Low ego; takes criticism without taking it personally
High autonomy; startup-native
Background
~5–8 years; startup time counts double
Architected data pipelines / greenfield data-platform work; dataset-as-a-product ownership is a strong signal
Exposure to edge/sensor data (audio/sonar, video, telemetry), time sync, and geospatial context is a plus
Location and Visa
Remote OK; LA strongly preferred, with occasional on-site visits to Torrance
US citizenship required
Nice-to-Have
Labeling workflows (interfaces, ontologies, consensus, QA) and label-store integrations
Splitting/sampling strategy design (by time, platform, geography, class, SNR) to avoid leakage
Orchestration (Airflow, Prefect) and metadata/lineage (MLflow, W&B)
Dataset-as-a-product track record with strong lineage and documentation
Athletic or competitive background (e.g., competitive chess) — reads as dynamic and startup-fit
Who Will Thrive Here
The data engineer who wants to own an entire platform employee-early
Architect-operators who can go from block diagram to shipped pipeline without hand-holding
Zero-to-one builders who've stood up data platforms from scratch and like greenfield
Low-ego, autonomous, comfortable in crunch