Lead AI Data Platform Architecture & Integration Consultant (CKAN/Deep Lake) at Proximal Cloud | Torre

Lead AI Data Platform Architecture & Integration Consultant (CKAN/Deep Lake)

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Freelance
Recurrent
Provide your expected compensation while applying
location_on
Remote (anywhere)
Shared by
Gabriela Enríquez Meléndez
5 days ago

Responsibilities


We’re looking for a Data Platform Architecture & AI Integration Consultant to bridge the gap between our National Data Platform (NDP) and Deep Lake.This isn't a "sit back and maintain" role. You will design the blueprint, build write-back pipelines, and hand over a fully versioned AI retrieval layer to top-tier academic teams.Role: Lead AI Data Platform ConsultantEngagement: 8–10 Weeks (Phased)Format: Remote / Phased Consulting Enablement🚀 What You’ll OwnThe Blueprint: Design identity-aware (CILogon/OIDC), federated integration patterns between CKAN and Deep Lake.The Pipelines: Build automated chunking, embedding, and write-back pipelines with full dataset versioning.The Retrieval Layer: Enable RAG-ready semantic search with >95% provenance back to NDP source metadata.Knowledge Transfer: Deliver Jupyter tutorials and runbooks so internal teams and faculty can run experiments independently.💡 What We’re Looking ForProven experience with CKAN (or distributed data catalogs).Deep expertise in Deep Lake, DVC, or LakeFS dataset versioning.Comfort with multimodal research data (text, tabular, geospatial, images).Hands-on mastery of Kubernetes, Docker, and JupyterHub workflows.Clear communication skills to turn complex infrastructure into student-friendly documentation.⏱️ The TimelineWeeks 1–2: Discovery, Security, & Blueprint DesignWeeks 3–5: Build Ingestion & Write-Back Pipelines (Target: ≥3 core datasets)Weeks 6–8: RAG Workflows, Semantic Search & Provenance APIWeeks 9–10: Handover, Enablement, & Educational Runbooks🔥 Why This Engagement? You won't just write code that gets buried. You are building the reference architecture for national-scale research, enabling faculty and students across the country to run reproducible AI experiments.📩 Interested? DM Pooja Lodwal or drop a comment below with your favorite dataset versioning stack (Deep Lake, DVC, LakeFS)? Let's talk!