Back to jobs
Remote | Machine Learning & NLP Research Specialist — $75–$105/hour
24-MAGNew YorkPosted 3d ago
Data Scientist (Contract)
Job Description
We are sharing a specialised part-time consulting opportunity for US-based machine learning and natural language processing professionals with hands-on experience in Python, model training and evaluation, transformers, large language models, retrieval systems, and applied ML pipelines.
This role supports an advanced AI initiative focused on identifying reasoning and capability gaps in frontier models. Selected professionals will design challenging real-world ML and NLP tasks, develop executable reference solutions, evaluate model performance, and analyse failures across language understanding, generation, retrieval, training workflows, and applied machine learning systems.
Key Responsibilities
ML & NLP Task Development
Design challenging machine learning and natural language processing problems based on practical research or industry experienceCreate tasks involving model training, evaluation, language understanding, generation, retrieval, or applied ML pipelinesTarget specific reasoning, implementation, and capability gaps in advanced AI modelsDefine clear specifications, expected behaviour, constraints, datasets, and evaluation criteria
Reference Solutions & Python Development
Develop accurate reference solutions and supporting materials using PythonIntegrate tasks into agent-based development and evaluation environmentsCreate executable tests, validation scripts, scoring logic, or model pipelines where appropriateEnsure reference implementations are technically sound, reproducible, and appropriately challenging
Model Evaluation & Failure Analysis
Evaluate model and agent performance across assigned ML and NLP tasksCompare generated outputs with reference solutions and expected resultsIdentify tasks where models demonstrate meaningful limitations or inconsistent behaviourClassify failures involving reasoning, implementation, retrieval, language understanding, generation, or instruction adherenceDocument findings through clear and technically detailed written analysis
Quality Calibration & Collaboration
Review tasks and evaluation methods developed by other machine learning specialistsMaintain consistent standards for difficulty, accuracy, realism, and technical qualityParticipate in calibration and peer-review activitiesCollaborate with other subject matter experts to improve evaluation coverage and reliability
Ideal Profile
Strong candidates may have:
Deep hands-on experience in machine learning, natural language processing, or bothPractical proficiency in Python demonstrated through professional, academic, or open-source workStrong understanding of modern ML and NLP methods, including transformers and large language modelsExperience with model training, fine-tuning, evaluation, retrieval, or production ML pipelinesFamiliarity with PyTorch, TensorFlow, JAX, Hugging Face, or comparable frameworks and toolingAbility to design realistic technical problems and develop complete reference solutionsStrong written communication and the ability to explain complex model behaviour clearlyReliable availability for approximately 20 hours per week
Educational Background
A degree in computer science, machine learning, artificial intelligence, computational linguistics, data science, or a related technical field is highly relevantGraduate or doctoral research in machine learning, NLP, language modelling, information retrieval, or related areas may be especially valuableEquivalent professional experience in applied ML or NLP may also be consideredPublished research, open-source contributions, or production ML work may strengthen an application
Nice to Have
Experience with transformer architectures, large language models, fine-tuning, or alignment methodsBackground in information retrieval, embeddings, reranking, search, or retrieval-augmented generationExperience with text classification, sequence modelling, summarisation, translation, or language generationFamiliarity with benchmark design, model evaluation, error analysis, or adversarial testingExperience building executable technical assessments or automated evaluation pipelinesPrevious involvement in AI training, model evaluation, data annotation, or quality-review programmesFamiliarity with agent-based development environments and structured technical rubrics
Why This Opportunity
Apply advanced machine learning and NLP expertise to frontier AI evaluationDesign realistic tasks grounded in practical research and engineering experienceAnalyse how modern models perform across language understanding, generation, retrieval, and applied ML workflowsHelp identify meaningful capability gaps and opportunities for model improvementWork remotely with a focused part-time commitment and competitive hourly compensation
Contract Details
Part-time W-2 contingent employment arrangementFully remote role available to candidates based in the United StatesExpected commitment of approximately 20 hours per weekCompetitive rates between $75–$105 per hour depending on expertise and project s