Software Engineer – Document AI / OCR & ML Inference
Role Details:
- Location: Remote.
- Schedule: Full time, 40 hours/week.
Overview:
- We are looking for a Software Engineer with deep experience in Document AI, OCR, and production ML inference to own and evolve a document processing pipeline used by enterprise customers at scale.
- This is a hands-on engineering role focused on turning complex document AI systems into fast, reliable, and maintainable production infrastructure.
- You will work across OCR, document understanding, CPU-based ML inference, and Python backend systems, with significant ownership over technical decisions and architecture.
Key Responsibilities:
- Own and evolve the end-to-end OCR and document understanding pipeline in Python.
- Improve OCR performance across accuracy, speed, layout fidelity, and structured data extraction.
- Optimize ML inference for CPU-only environments, focusing on latency, memory usage, and throughput.
- Implement production optimization techniques such as ONNX Runtime, quantization, and batching.
- Redesign and simplify existing architecture, removing brittle heuristics, unnecessary complexity, and performance bottlenecks.
- Scale document processing systems to support higher volumes and multiple enterprise clients reliably.
- Improve model lifecycle and operations, including training pipelines, data sampling, versioning, debugging, and retraining.
- Expand into broader system ownership over time, including backend services, APIs, data pipelines, and architecture decisions.
Requirements:
- 6+ years of software engineering experience, with strong production-level Python.
- Hands-on experience building or owning Document AI, OCR, or document processing systems in production.
- Experience with OCR technologies such as Tesseract, PaddleOCR, EasyOCR, Doctr, or similar.
- Experience with document understanding models such as LayoutLMv3, Donut, DocFormer, or comparable architectures.
- Experience deploying and optimizing ML models in CPU-only environments using tools such as ONNX Runtime, quantization, TorchScript, or alternative inference runtimes.
- Experience designing scalable, maintainable production systems used by external customers.
- Strong English communication skills.
Nice to Have:
- Experience with PyTorch, torchvision, Hugging Face Transformers, or Accelerate.
- Experience with quantized models, including INT8 and dynamic/static quantization.
- Experience designing ML training pipelines, data sampling strategies, and model versioning workflows.
- Experience scaling document processing or ML systems across multiple enterprise clients.
- Experience working with PDF processing and complex document layouts.
Why Join:
- Take ownership of a high-impact production Document AI system, solve challenging OCR and ML inference problems at scale, and grow from deep pipeline ownership into broader backend and architecture responsibility.