Back to jobs

Python AI Model Evaluation Developer
OpenTrain AIAnywhereAdded 1w ago
AI Evaluation Engineer
Part-time and Contractor
Remote
Apply on company siteCraft my tailored resume free
Free to start · No card · 5 credits the moment you sign up
Job description
About OpenTrainOpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting for this role, giving contributors a way to build experience in a fast-growing field where human expertise helps shape advanced AI systems.
Fully remote work available worldwideFree OpenTrain account and profileBuild a portfolio around AI training and data-labeling experienceAbout AI Model Evaluation WorkAI training is the human side of building artificial intelligence. Contributors create examples, review model outputs, write evaluations, and provide feedback that helps language models become more accurate, useful, and aligned.
This role focuses on coding and evaluation rather than building or fine-tuning the models themselves. Your Python expertise and technical judgment will help produce reliable training data and assess the quality of model-generated solutions.
Create supervised fine-tuning examplesEvaluate and rank language-model responsesSupport reinforcement learning from human feedback workflowsHelp improve the quality of AI-generated code and technical explanationsThe RoleOpenTrain is seeking a Python AI Model Evaluation Developer to support large language model improvement through hands-on coding, data generation, and evaluation. You will write Python solutions to code-based questions, create high-quality training examples, compare responses from different models, and provide detailed feedback.
This is a fully remote contractor assignment lasting one month. The role is categorized as entry level, while the stated technical requirements call for at least three years of strong Python programming experience.
Engagement: Contractor and part timeDuration: One monthWorkload options: 20, 30, or 40 hours per weekMinimum commitment: 20 hours per week and at least 4 hours per dayRequired overlap: 4 hours with Pacific TimeLocation: Worldwide and fully remoteWorking language: Fluent written and conversational EnglishCompensation: Not disclosedWhat You’ll DoYou will combine software-development discipline with careful model evaluation. The work includes creating and reviewing technical content, analyzing performance, and delivering feedback that researchers and annotators can use to strengthen AI training processes.
Design, develop, and maintain efficient, high-quality Python code for AI training and evaluation workflows.Conduct evaluations to benchmark model performance and analyze results for continuous improvement.Evaluate and rank AI model responses using quality, relevance, accuracy, and alignment criteria.Create task-specific supervised fine-tuning datasets, model responses, code solutions, and evaluation rationales.Create and refine responses to improve clarity, relevance, and technical accuracy.Contribute to reinforcement learning from human feedback activities and reward-model refinement.Design evaluation strategies, review code and documentation, and provide constructive technical feedback.Collaborate with researchers and annotators while exploring tools and methods that strengthen AI training processes.RequirementsApplicants should bring strong Python programming ability, software-development fundamentals, and the communication skills needed to explain technical judgments clearly. Experience working with model responses and AI training data-generation workflows is also required.
At least three years of strong Python programming experienceFamiliarity with Python frameworks and librariesKnowledge of software-development quality, formatting, architecture, and best practicesExperience with unit, integration, and property-based testing in PythonUnderstanding of multithreading and asynchronous programmingAbility to refactor code safely without introducing regressionsAbility to diagnose memory or concurrency issuesAbility to write clear evaluation rationales and technically accurate responsesWorking knowledge of supervised fine-tuning and reinforcement learning from human feedback data-generation and evaluation workflowsFluent conversational and written English communicationWho Should ApplyThis opportunity is suited to Python developers who enjoy solving technical problems, reviewing code, and making precise quality judgments. It may also appeal to software engineers interested in applying their skills to the rapidly growing field of AI training.
Attention to detail matters because your code, rankings, rationales, and feedback will be used to assess and improve language-model behavior. The assignment requires a consistent weekly commitment and Pacific Time overlap.
Python developers with strong testing and debugging experienceSoftware professionals comfortable with concurrency and asynchronous programmingTechnical communicators who can explain evaluation decisions clearlyContributors interested in coding datasets, model evaluation, and RLHF workHow It WorksApply through OpenTrain to be considered for this contractor assignment. If selected, you will complete remote AI training and evaluation work within the available 20, 30, or 40 hour weekly commitment.
OpenTrain helps contributors discover and grow careers in AI training and data labeling. Your profile can showcase relevant experience and support a longer-term portfolio as you take on additional opportunities in the field.
Create a free OpenTrain accountBuild a profile highlighting Python and evaluation experienceApply in minutes through OpenTrainComplete the assignment remotely according to the required schedule