AI Data Architect (AWS) Designing Data Lakes, Warehouses, ETL and RAG Systems at Innovative Solutions | Torre

AI Data Architect (AWS) Designing Data Lakes, Warehouses, ETL and RAG Systems

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Provide your expected compensation while applying
location_on
Remote (for United States residents)
Shared by
Emma of Torre.ai
7 days ago

Responsibilities


As a Data/AI Architect, you'll design and build data-driven cloud architectures on AWS — from S3 data lakes and Glue ETL pipelines to data warehouses and RAG-powered AI systems. You'll own the full data stack across a variety of industries and projects: one engagement you're designing a Redshift data warehouse with medallion architecture processing 31M transactions/month, the next you're building a Bedrock Knowledge Base with OpenSearch vector search. Real ownership, real variety.Location: This can be a remote opportunity, with 2 weeks of travel into Rochester, NY per quarter What You'll Do:Design and build S3 data lakes with multi-zone organization, partitioning strategies, lifecycle policies, and encryptionImplement medallion architecture (bronze/silver/gold) for data warehouses on Redshift, Snowflake, or DatabricksBuild AWS Glue ETL pipelines (Python Shell and Spark) with incremental extraction, Data Catalog management, and optimized Parquet outputDesign star/snowflake schemas, materialized views, and gold-layer models optimized for BI consumption (QuickSight, PowerBI)Configure data warehouse platforms — Redshift with Zero-ETL from Aurora, Snowflake with Snowpipe, Databricks with Delta Lake and Auto LoaderDesign RAG systems using Bedrock Knowledge Base with OpenSearch Serverless vector search and Titan EmbeddingsArchitect document AI pipelines using Textract, Comprehend, and Bedrock for entity extractionDesign SageMaker ML pipelines for training, Model Registry, and inferenceLead data discovery sessions with client stakeholders and present architecture recommendations to technical and business audiencesMentor delivery team members on data architecture patterns and AWS data servicesContribute to R&D projects evaluating emerging AWS data and AI capabilitiesRequired Skills:5+ years professional IT experience, 2+ years professional AWS experienceAt least one AWS Professional-level certification (Solutions Architect Professional or Data Engineer Specialty preferred)Python for data pipelines (Glue jobs, Lambda, SageMaker scripts) and PySpark for Glue Spark jobsSQL and NoSQL on AWS — Aurora PostgreSQL, RDS PostgreSQL, DocumentDB, DynamoDB — including schema design and query optimizationData modeling — conceptual, logical, and physical models for AWS data platforms; normalized silver-layer schemas, denormalized star/snowflake gold-layer schemas, data dictionariesDimensional modeling and medallion architecture (bronze/silver/gold) on Redshift, Snowflake, or Databricks, including materialized views and incremental refresh patternsAWS Glue ETL (Python Shell and Spark), Glue Data Catalog, and crawlersS3 data lake architecture with partitioning, lifecycle policies, and encryptionsPreferred:RAG systems with Bedrock Knowledge Base and OpenSearch Serverless vector searchAmazon SageMaker for ML training, Model Registry, and inferenceAWS HealthLake, FHIR R4 transformation, and HIPAA-compliant data pipelinesDocument AI with Amazon Textract and ComprehendAmazon Athena, QuickSight, or PowerBI integrationTerraform or CloudFormation for data infrastructure as codeStep Functions, EventBridge, and Lambda for event-driven pipeline orchestration