Back to jobs

Artificial Intelligence Engineer (Local to GA)

ThoughtStormAtlanta, GAPosted 1w ago
AI Engineer
Apply on LinkedIn

Job Description

Job Description

Agent development, evals, observability, tool-use, and production AI systems

Role Overview

We are hiring an AI Engineer to design, build, evaluate, and operate production-grade AI agents and AI-powered

product features. This role sits at the intersection of software engineering, model integration, workflow design,

eval-driven development, and product execution.

The right candidate has moved beyond prompt experiments. They know how to build systems that use tools,

maintain state, retrieve context, stream useful UI events, recover from failures, expose traces, and prove

behavior with evals before shipping.

Core Responsibilities

• Build AI agents and AI-powered product features using TypeScript, Node.js, Next.js, and modern agent

runtimes.

• Integrate model providers including Anthropic Claude, OpenAI, Google Gemini, OpenRouter, Ollama, LM

Studio, and Azure OpenAI where appropriate.

• Design tool-using agents that can call APIs, use MCP servers, query databases, browse web content,

execute workflows, and interact with internal product systems.

• Build production AI interfaces with streaming responses, structured outputs, tool-call visibility, approval

states, retries, and user-facing error handling.

• Create eval suites for prompts, tool use, RAG quality, safety boundaries, regression behavior, task

completion, and end-to-end product workflows.

• Instrument agents with traces, logs, cost metrics, latency metrics, tool-call audits, and failure analysis

workflows.

• Implement memory, session persistence, retrieval, workflow state, and audit trails using appropriate

persistence layers.

• Design safety boundaries for tool execution, sandboxed code, permissions, human-in-the-loop review, and

data access.

• Collaborate with product, design, and engineering partners to turn AI capabilities into reliable user

workflows rather than demos.

Preferred Core Stack

Area Preferred Tools

Language / Runtime TypeScript, Node.js

Frontend / Product Next.js, React, Vercel AI SDK

Package Management pnpm

Agent Runtimes Claude Agent SDK, OpenAI Agents SDK, VoltAgent,

LangGraph, Genkit

Tool Protocol MCP

Data / Persistence Postgres, Supabase, SQLite, Redis, pgvector

Testing Vitest, Playwright, provider-mocked tests, golden-task

evals

Deployment Vercel, Docker, Cloud Run, Fly.io, AWS / GCP / Azure

where appropriate

Required Experience

• Strong TypeScript, Node.js, and web application engineering skills.

• Experience building production applications with React or Next.js.

• Hands-on experience with LLM APIs, tool calling, structured outputs, streaming responses, and providerspecific SDKs.

• Experience building agent workflows that use tools, state, retrieval, retries, and human approval, not just

one-shot prompts.

• Experience designing and running evals for prompts, agents, tools, or AI product flows.

• Ability to debug model behavior using traces, logs, transcripts, datasets, and reproducible test cases.

• Practical understanding of safety, permissions, secrets, data access, and tool execution risk.

Nice To Have

• Experience with Claude Code, Claude Agent SDK, OpenAI Agents SDK, VoltAgent, LangGraph, Genkit, or

Vercel AI SDK.

• Experience building or integrating MCP servers.

• Experience with Promptfoo, Braintrust, LangSmith, Langfuse, Ragas, DeepEval, TruLens, Arize Phoenix, or

similar eval and observability tools.

• Experience with RAG systems, embeddings, vector stores, query planning, reranking, and retrieval quality

measurement.

• Experience with browser automation, sandboxed execution, coding agents, developer tools, or internal

copilots.

• Experience with red teaming, guardrails, prompt-injection testing, structured validation, and policy

enforcement.

What Success Looks Like

• AI features are useful, reliable, observable, and testable.

• Agent behavior can be measured against golden tasks and regression datasets.

• Tool calls are safe, visible, auditable, and easy to debug.

• Model/provider choices are pragmatic and swappable where the product needs flexibility.

• The team can understand why an AI workflow failed and verify when it has improve