Back to jobs

AI Engineer for Software Engineering Evaluation

OpenTrain AIAnywherePosted 1w ago
AI Engineer - Python/LLMs
60–120 an hour
Part-time and Contractor
Remote
Apply on company siteCraft my tailored resume free

Free to start · No card

Job description

About OpenTrainOpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps people discover cutting-edge projects, build their AI work profile, and apply in minutes. Creating an OpenTrain account is free. About AI Training and EvaluationAI training is the human side of building modern artificial intelligence. Engineers, researchers, and other contributors create examples, test model behavior, and provide feedback that helps AI systems reason, use tools, and produce more reliable results. This project focuses on evaluating AI agents through realistic software engineering tasks. Your technical judgment will help measure how well models discover information, solve problems, use development tools, and implement dependable software. The AI Engineer RoleOpenTrain AI is seeking experienced AI Engineers and software engineers for next-generation AI training and evaluation projects. You will help create reinforcement learning environments that test AI models on complex software engineering tasks using Model Context Protocol tools. This is a remote contractor opportunity for approximately 15 hours per week, with a flexible schedule that may include weekends. The expected project duration is 1 to 3 months, and urgent hiring is underway. Compensation: $60–$120 per hourContract, part-time engagementWorkload: approximately 15 hours per week and under 20 hours weeklyRemote work with flexible days and hoursOutput-based payment for completed tasks meeting project specifications and quality requirementsMinimum weekly submission requirements may applyWhat You'll DoYou will contribute practical software engineering expertise to development environments and evaluation systems designed for AI agents. Individual assignments may vary according to project needs, experience, and workflow. Fix bugs and debug complex technical issuesImplement new features and develop maintainable software solutionsRefactor existing codebasesOptimize software performance and scalabilityDesign reproducible development environmentsCreate deterministic verification systemsDevelop golden reference solutions for AI evaluationEvaluate software engineering ability and effective MCP tool usageRequired Skills and QualificationsStrong software engineering expertise and practical programming experience are the primary requirements. No prior professional AI experience is required, but you should be comfortable working from technical specifications and producing reliable, maintainable code. The role requires fluent English and experience with computer code programming labeling workflows, AI software engineering training, and evaluation subject matter. Candidates should be available to collaborate effectively in a remote environment. Strong proficiency in C++, Python, Java, Go, TypeScript, or RustPython 3, Java, Rust, C++ fundamentals, and TypeScript experienceDeep understanding of algorithms, data structures, and performance tuningDemonstrated experience debugging complex software issuesStrong experience with feature development and codebase refactoringProven ability to optimize software for performance and scalabilityExcellent written and verbal communication skillsStrong attention to detail and ability to follow technical specificationsIntermediate experience levelFluent English proficiencyPreferred ExperienceThe following experience is helpful but not required. Candidates may be considered based on the strength of their software engineering background and practical programming work. Experience with large-scale or distributed codebasesFamiliarity with AI or machine learning systemsExperience with modern developer tools and APIsParticipation in rigorous code reviewsKnowledge of software engineering best practicesExperience designing automated testing or verification systemsFamiliarity with Model Context Protocol or similar tool-use frameworksLocation, Selection, and Start TimelineCandidates must be based in one of the following countries: the United Kingdom, United States, United Arab Emirates, Ireland, India, South Korea, Japan, Finland, Mexico, Brazil, Australia, Austria, Canada, Belgium, Egypt, or France. The working language is English. The selection process is designed to assess technical capability and project fit. Selected candidates should be prepared to begin their first tasks within 24–48 hours after completing onboarding. Submit an application with your updated resume and availability for the interviewComplete screening questionsComplete an approximately 30-minute AI interviewComplete a technical assessment if requiredProceed through hiring manager reviewComplete onboarding and receive a project assignmentStart as soon as possible if selected