Back to jobs

Fall 2026 Internship- AI/ML Developer

Medlaunch ConceptsAnywherePosted 3w ago
Machine Learning Intern
Internship
Remote
Apply on company siteCraft my tailored resume free

Free to start · No card

Job description

About Us We’re building a transformative healthcare accreditation platform that is revolutionizing how hospitals manage compliance, quality improvement, and regulatory processes. Our platform combines cutting-edge technology with deep healthcare domain expertise to solve real problems for healthcare organizations nationwide. The Opportunity The goal is to have interns turn into full-time employees; therefore, you will be given full-time responsibilities from day one. You will be working in a high-velocity growth startup and will be required to move fast. You’ll work directly with our engineering team on production healthcare ML systems, gaining hands-on experience with enterprise-grade pipelines while making real contributions that impact our product and customers. Compensation Structure Base position is unpaid; however, qualified candidates may receive upfront equity compensation based on their experience level and demonstrated capabilities. We evaluate each applicant individually and offer equity packages commensurate with their potential contribution. About the Role We’re hiring an ML Pipeline & Data Science Developer. This role is focused on building and operating production machine learning pipelines—clustering, classification, recommendation, and outcome analysis—on top of healthcare event and quality data. You’ll design the data workflows that power risk identification, pattern detection, and corrective action effectiveness analysis across the platform. Requirements Deep AWS experience: Lambda, Step Functions, EventBridge, S3, SageMaker, and Bedrock for production ML infrastructureStrong Python ML fundamentals: scikit-learn, pandas, numpy, scipy for building and evaluating modelsPipeline architecture: Designing end-to-end data and ML pipelines—ingestion, transformation, feature engineering, model training, inference, and monitoringClustering and classification: Hands-on experience with unsupervised and supervised learning techniques (topic modeling, text clustering, multi-class classification)Statistical rigor: Hypothesis testing, A/B testing, experimental design, and outcome measurementVersion control: Git for code and model versioningCollaborative mindset: Ability to work within a specialized team structure and move fast in a startup environment Nice to Have Experience with BERTopic, sentence-transformers, or similar NLP clustering librariesFamiliarity with recommendation systems or best-practice identification from historical outcome dataHealthcare or regulated industry experienceTime-series forecasting and anomaly detectionKnowledge of data privacy and compliance frameworksExperience with MongoDB aggregation pipelines or similar document-database analytics What You’ll Build ML Pipelines & Predictive Risk Analytics Event clustering pipelines: Automated workflows that ingest healthcare event data and cluster by similarity to allow AI inference to surface patterns and systemic risksClassification systems: Multi-class models that categorize incoming events by type, severity, and process area using both structured metadata and unstructured textCorrective action recommendation: Models that analyze historical outcomes to identify the most effective actions for a given type of event or similar event, surfacing best practices from past resolutionsRisk scoring and prioritization: Scoring frameworks that rank identified risks by frequency, severity, and trend direction to focus quality improvement effortsOutcome and effectiveness analysis: Statistical models that measure whether implemented corrective actions actually reduce recurrence, with confidence intervals and significance testing Data Pipeline Infrastructure Ingestion and transformation: Lambda-based pipelines that process incoming event data, normalize fields, and route records through feature engineering stagesFeature engineering: Domain-specific feature stores built on healthcare event attributes—process area, finding type, facility context, temporal patterns, and text-derived featuresBatch and real-time inference: Scheduled batch pipelines (EventBridge + Lambda) for nightly AI/ML model runs and on-demand inference endpoints for interactive useModel monitoring and versioning: MLOps workflows for tracking experiments, model performance drift, and A/B comparisons across model versionsVisualization and reporting: Interactive dashboards and data visualizations that surface pipeline outputs to end users and internal stakeholders Key Responsibilities Design and implement end-to-end ML pipelines on AWS—from data ingestion through model inference and output deliveryBuild clustering and topic modeling workflows that group healthcare events by similarity and identify systemic patternsDevelop classification models for automated event categorization across process areas and finding typesAnalyze historical action data to build recommendation models that surface the most effective interventions for a given scenarioConduct statistical analysis on action outcomes to measure effectiveness and identify best practicesEngineer features from healthcare event data—combining structured fields, temporal patterns, and text-derived signalsBuild and maintain Lambda-based data transformation and inference pipelines with EventBridge schedulingImplement model versioning, experiment tracking, and performance monitoring using SageMakerCreate visualizations and reports that communicate risk patterns and model outputs to non-technical stakeholdersCollaborate with product and engineering teams to integrate ML outputs into the broader platform Required Qualifications Candidates must meet all Core Qualifications plus demonstrate depth in the ML Pipeline & Data Science focus area. Core Qualifications Advanced Python programming skills (2+ years)Git for version control and collaborative development workflowsUnderstanding of deployment and production systemsExperience building on AWS (Lambda, S3, EventBridge, Step Functions, SageMaker) ML Pipeline & Data Science 2+ years with Python ML libraries (scikit-learn, pandas, numpy, scipy)Experience designing and operating ML pipelines—data ingestion, feature engineering, training, inferenceHands-on experience with clustering (k-means, DBSCAN, hierarchical, or topic modeling such as BERTopic/LDA)Strong statistical analysis skills: hypothesis testing, A/B testing, experimental design, outcome measurementExperience with text embeddings and NLP feature extraction (sentence-transformers, spaCy, or similar)1+ years with AWS SageMaker or equivalent managed ML platformData visualization expertise (matplotlib, seaborn, plotly, Tableau)SQL proficiency for analytics and data queryingMLflow or equivalent for experiment tracking Nice to Have Experience with recommendation systems or outcome-based ranking modelsHealthcare or regulated industry experienceMongoDB aggregation pipelines for document-database analyticsHugging Face transformers and fine-tuning workflowsTime-series forecasting and anomaly detectionAdvanced statistical modeling (survival analysis, propensity scoring)Knowledge of data privacy and compliance frameworks Technical Stack Category Technologies AWS Infrastructure Lambda, Step Functions, EventBridge, S3, EC2, SageMaker, Bedrock ML & Data Science Python (scikit-learn, pandas, numpy, scipy, statsmodels), MLflow, BERTopic, sentence-transformers NLP & Text Processing spaCy, NLTK, sentence-transformers, text embeddings, topic modeling Data & Analytics SQL, MongoDB aggregation pipelines, Plotly, Tableau, matplotlib, seaborn DevOps & Tooling Git, Docker, CI/CD, CloudWatch, infrastructure-as-code Our Hiring Process We believe in a transparent and thorough selection process that respects your time while ensuring mutual fit: Initial Screening Call We’ll discuss your background, experience, and career goals, while providing an overview of the role and our team culture. Technical Challenge (issued on a case-by-case basis) You’ll receive a real-world technical challenge to complete within a specified timeframe. We encourage you to leverage all available resources—including AI tools, documentation, and libraries—just as you would in a production environment. This reflects how we actually work and allows you to showcase your problem-solving approach. Technical Interview We’ll have an in-depth discussion about your solution and explore related technical concepts. You should be prepared to walk through every aspect of your submission—explaining architectural decisions, code logic, trade-offs, and potential improvements. Whether you wrote specific code sections manually or generated them with AI assistance, you must demonstrate complete ownership and understanding of the entire codebase. This is a production-level assessment: we expect you to discuss, debug, and defend your work as if it were going live tomorrow. We’re looking for engineers who can think critically, adapt their approach, and truly understand the systems they build—not just those who can generate code. Ready to apply? We look forward to hearing from you! MedLaunch is an equal opportunity employer committed to diversity and inclusion.