← Zurück zu den ErgebnissenAI Software Engineer — Python, Document Intelligence & AWS (Remote)
NextGen Coding Company
- Standort
- Anywhere
- Veröffentlicht
- 7. Okt. 2026
Prüfe vor der Bewerbung auf der Website des Arbeitgebers, ob die Stelle noch offen ist und welche Bedingungen gelten.
Stellenbeschreibung
Die Oberfläche ist auf Deutsch. Die Titel und Beschreibungen der Arbeitgeber bleiben in ihrer ursprünglichen Veröffentlichungssprache, die Englisch sein kann.
AI Software Engineer — Python, Document Intelligence & AWS
NextGen Coding Company
Location: New York City — Hybrid / In-Person Required
Compensation: $30/hour
Hours: 40 Hours/Week
Engagement: Contract with potential for long-term engagement
Eligibility: Must be a U.S. citizen and able to complete a W-9. No visa sponsorship is available for the role.
Role Overview
NextGen Coding Company is hiring an AI Software Engineer in New York City to join an active enterprise AI project focused on document intelligence, large-scale data processing, and evidence-backed AI analysis.
The work is hands-on and engineering-heavy. You will help build a cloud platform that ingests thousands of pages of PDFs, scans, images, HTML, and other data; performs OCR and Python-based processing; structures and indexes the resulting information; and makes the data searchable and usable by modern LLMs.
We are looking for a strong builder, not someone whose AI experience consists primarily of calling an LLM API.
You should be comfortable jumping directly into an existing project, understanding the architecture, debugging difficult data-processing problems, and shipping production code.
What You’ll Work On
Build production systems primarily in PythonBuild and improve large-scale document ingestion pipelinesProcess PDFs, scans, images, HTML, and other file formatsBuild OCR and document-intelligence workflowsClassify incoming documents and route them through appropriate processing pipelinesBuild Python post-processing to clean, normalize, validate, and structure extracted informationSolve large-file processing issues using chunking, queues, parallel processing, retries, and recoveryExtract text, tables, entities, metadata, relationships, and structured recordsBuild ETL and asynchronous data-processing pipelinesBuild hybrid search, vector search, embeddings, RAG, and rerankingIntegrate Claude, OpenAI, Gemini, Qwen, and other LLMsBuild evidence and citation systems connecting AI outputs to original source materialBuild and maintain production infrastructure in AWSWork with PostgreSQL, OpenSearch, S3, queues, caching, and APIsBuild integrations including authentication, webhooks, Stripe, and third-party servicesWrite tests for OCR, extraction, retrieval, data processing, APIs, and AI outputsDiagnose production failures and improve system reliability and performanceCore Technical Skills
Python: FastAPI, data processing, ETL, APIs, asynchronous/background jobs
Document Intelligence: OCR, PDF parsing, scanned documents, OpenCV, PyMuPDF, PaddleOCR, Tesseract, Docling, Unstructured, or similar tools
AI: Claude, OpenAI, Gemini, Qwen, RAG, embeddings, structured outputs, reranking, model orchestration
Data: PostgreSQL, SQL, OpenSearch/Elasticsearch, vector databases, Redis
AWS: S3, RDS, OpenSearch, Bedrock, EC2/ECS/EKS, Lambda, SQS, IAM, CloudWatch
Infrastructure: Docker, Linux, CI/CD, production cloud deployments
Frontend experience with React, Next.js, and TypeScript is helpful but is not the primary focus.
Who We Want
We want someone who can be given a difficult engineering problem and figure it out.
For example:
“We have several thousand pages across PDFs, scans, images, and other file formats. Large files are processing inconsistently. Build a reliable AWS pipeline that classifies the files, performs OCR where required, cleans and structures the output with Python, handles failures, indexes the information, and makes the evidence usable by an LLM.”
You should be able to break a problem like the one above into an architecture and then actually build it.
Strong candidates will have experience building real production systems involving Python, data pipelines, OCR/document processing, AWS, or AI.
Requirements
Must be based in New York CityMust be a U.S. citizenMust be available approximately 40 hours per weekStrong Python engineering abilityProduction AWS experienceExperience with OCR, document processing, data engineering, or similar high-volume processing systemsExperience integrating modern LLMsStrong backend/API/database fundamentalsComfortable independently debugging complex engineering problemsAble to meet with our team in person in NYCAble to complete a technical engineering screenInterview Process
We are intentionally keeping the process straightforward:
1. Application Review — Resume, LinkedIn, GitHub, and/or examples of systems you have built.
2. Technical Screen — Python, OCR/document processing, AWS, data architecture, and LLM engineering.
3. In-Person NYC Meeting — Meet the team, walk through the project, and discuss how you would approach real engineering problems from the platform.
We are looking for someone who can join quickly, take ownership, and immediately contribute to a technically ambitious production AI system.
About NextGen Coding Company
NextGen Coding Company is a U.S.-based software engineering firm building custom software, AI systems, automation platforms, data infrastructure, and enterprise applications.
Our engineers work on real production systems across AI/ML, document intelligence, data engineering, financial and compliance technology, and cloud infrastructure.
To Apply
Please send:
Resume or LinkedInGitHub and/or portfolio, if availableA short description of the most technically difficult production system you have builtAny relevant experience with Python, OCR/document processing, AWS, and LLMsNYC candidates only.