Back to jobs

LLM / GenAI Engineer

Evlo AIAtlanta, GAPosted 3d ago
AI Engineer
Apply on LinkedIn

Job Description

About The Role

The role is responsible for taking generative AI initiatives from proof-of-concept to production-grade deployment, focusing on reliability, throughput, and system latency. The engineer will design and deploy agentic workflows, advanced RAG systems, and self-hosted model pipelines that power core application features.

This position collaborates closely with backend engineers, product managers, and data engineers to architect highly scalable infrastructure that serves millions of LLM requests daily while maintaining strict cost and quality constraints.

Key Responsibilities

  • Design and deploy production-grade Retrieval-Augmented Generation (RAG) pipelines incorporating hybrid search, reranking, and metadata filtering
  • Develop multi-agent workflows and orchestration layers using frameworks like LangGraph, CrewAI, or custom state machines
  • Optimize LLM inference latency and throughput using quantization techniques, speculative decoding, and engines like vLLM or TensorRT-LLM
  • Implement systematic LLM evaluation and guardrail systems to monitor and mitigate hallucination, prompt injection, and output drift in real time
  • Build automated data curation and preprocessing pipelines in Python for fine-tuning open-source models (e.g., Llama, Mistral) on domain-specific tasks
  • Integrate vector databases such as Pinecone, Qdrant, or pgvector into distributed application architectures with high-throughput synchronization

What We Are Looking For

  • 3-6 years of software engineering experience, with at least 1.5 years dedicated to building and deploying LLM applications in a production environment
  • Expert-level Python programming skills, including asynchronous programming, FastAPI development, and concurrent data processing
  • Hands-on experience with LLM APIs (OpenAI, Anthropic) as well as hosting and running open-source models locally or on cloud VMs
  • Deep understanding of vector embeddings, semantic search, indexing strategies, and advanced retrieval techniques
  • Solid software engineering fundamentals including Docker, CI/CD pipelines, structured logging, and observability tools like Arize Phoenix or LangSmith
  • Bonus: Experience with parameter-efficient fine-tuning (PEFT, LoRA), Triton Inference Server, or Kubernetes deployment for AI workloads