Back to jobs

Remote | Machine Learning & NLP Research Specialist — $75–$105/hour

24-MAGNew YorkPosted 3d ago
Data Scientist (Contract)
Apply on company site

Job Description

We are sharing a specialised part-time consulting opportunity for US-based machine learning and natural language processing professionals with hands-on experience in Python, model training and evaluation, transformers, large language models, retrieval systems, and applied ML pipelines. This role supports an advanced AI initiative focused on identifying reasoning and capability gaps in frontier models. Selected professionals will design challenging real-world ML and NLP tasks, develop executable reference solutions, evaluate model performance, and analyse failures across language understanding, generation, retrieval, training workflows, and applied machine learning systems. Key Responsibilities ML & NLP Task Development Design challenging machine learning and natural language processing problems based on practical research or industry experienceCreate tasks involving model training, evaluation, language understanding, generation, retrieval, or applied ML pipelinesTarget specific reasoning, implementation, and capability gaps in advanced AI modelsDefine clear specifications, expected behaviour, constraints, datasets, and evaluation criteria Reference Solutions & Python Development Develop accurate reference solutions and supporting materials using PythonIntegrate tasks into agent-based development and evaluation environmentsCreate executable tests, validation scripts, scoring logic, or model pipelines where appropriateEnsure reference implementations are technically sound, reproducible, and appropriately challenging Model Evaluation & Failure Analysis Evaluate model and agent performance across assigned ML and NLP tasksCompare generated outputs with reference solutions and expected resultsIdentify tasks where models demonstrate meaningful limitations or inconsistent behaviourClassify failures involving reasoning, implementation, retrieval, language understanding, generation, or instruction adherenceDocument findings through clear and technically detailed written analysis Quality Calibration & Collaboration Review tasks and evaluation methods developed by other machine learning specialistsMaintain consistent standards for difficulty, accuracy, realism, and technical qualityParticipate in calibration and peer-review activitiesCollaborate with other subject matter experts to improve evaluation coverage and reliability Ideal Profile Strong candidates may have: Deep hands-on experience in machine learning, natural language processing, or bothPractical proficiency in Python demonstrated through professional, academic, or open-source workStrong understanding of modern ML and NLP methods, including transformers and large language modelsExperience with model training, fine-tuning, evaluation, retrieval, or production ML pipelinesFamiliarity with PyTorch, TensorFlow, JAX, Hugging Face, or comparable frameworks and toolingAbility to design realistic technical problems and develop complete reference solutionsStrong written communication and the ability to explain complex model behaviour clearlyReliable availability for approximately 20 hours per week Educational Background A degree in computer science, machine learning, artificial intelligence, computational linguistics, data science, or a related technical field is highly relevantGraduate or doctoral research in machine learning, NLP, language modelling, information retrieval, or related areas may be especially valuableEquivalent professional experience in applied ML or NLP may also be consideredPublished research, open-source contributions, or production ML work may strengthen an application Nice to Have Experience with transformer architectures, large language models, fine-tuning, or alignment methodsBackground in information retrieval, embeddings, reranking, search, or retrieval-augmented generationExperience with text classification, sequence modelling, summarisation, translation, or language generationFamiliarity with benchmark design, model evaluation, error analysis, or adversarial testingExperience building executable technical assessments or automated evaluation pipelinesPrevious involvement in AI training, model evaluation, data annotation, or quality-review programmesFamiliarity with agent-based development environments and structured technical rubrics Why This Opportunity Apply advanced machine learning and NLP expertise to frontier AI evaluationDesign realistic tasks grounded in practical research and engineering experienceAnalyse how modern models perform across language understanding, generation, retrieval, and applied ML workflowsHelp identify meaningful capability gaps and opportunities for model improvementWork remotely with a focused part-time commitment and competitive hourly compensation Contract Details Part-time W-2 contingent employment arrangementFully remote role available to candidates based in the United StatesExpected commitment of approximately 20 hours per weekCompetitive rates between $75–$105 per hour depending on expertise and project s