Automation Engineer at Globaltize | Torre

Automation Engineer

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: To be defined

Provide your expected compensation while applying (or, if you'd like more context, simply ASK).
location_on
Remote (anywhere)
Posted 3 days ago

Responsibilities


Software Engineer – Document AI / OCR & ML Inference Role Details: - Location: Remote. - Schedule: Full time, 40 hours/week. Overview: - We are looking for a Software Engineer with deep experience in Document AI, OCR, and production ML inference to own and evolve a document processing pipeline used by enterprise customers at scale. - This is a hands-on engineering role focused on turning complex document AI systems into fast, reliable, and maintainable production infrastructure. - You will work across OCR, document understanding, CPU-based ML inference, and Python backend systems, with significant ownership over technical decisions and architecture. Key Responsibilities: - Own and evolve the end-to-end OCR and document understanding pipeline in Python. - Improve OCR performance across accuracy, speed, layout fidelity, and structured data extraction. - Optimize ML inference for CPU-only environments, focusing on latency, memory usage, and throughput. - Implement production optimization techniques such as ONNX Runtime, quantization, and batching. - Redesign and simplify existing architecture, removing brittle heuristics, unnecessary complexity, and performance bottlenecks. - Scale document processing systems to support higher volumes and multiple enterprise clients reliably. - Improve model lifecycle and operations, including training pipelines, data sampling, versioning, debugging, and retraining. - Expand into broader system ownership over time, including backend services, APIs, data pipelines, and architecture decisions. Requirements: - 6+ years of software engineering experience, with strong production-level Python. - Hands-on experience building or owning Document AI, OCR, or document processing systems in production. - Experience with OCR technologies such as Tesseract, PaddleOCR, EasyOCR, Doctr, or similar. - Experience with document understanding models such as LayoutLMv3, Donut, DocFormer, or comparable architectures. - Experience deploying and optimizing ML models in CPU-only environments using tools such as ONNX Runtime, quantization, TorchScript, or alternative inference runtimes. - Experience designing scalable, maintainable production systems used by external customers. - Strong English communication skills. Nice to Have: - Experience with PyTorch, torchvision, Hugging Face Transformers, or Accelerate. - Experience with quantized models, including INT8 and dynamic/static quantization. - Experience designing ML training pipelines, data sampling strategies, and model versioning workflows. - Experience scaling document processing or ML systems across multiple enterprise clients. - Experience working with PDF processing and complex document layouts. Why Join: - Take ownership of a high-impact production Document AI system, solve challenging OCR and ML inference problems at scale, and grow from deep pipeline ownership into broader backend and architecture responsibility.
Closes in:
0
days
0
hours
0
min
0
sec
tune NOT FOR YOU? IMPROVE YOUR RESULTS