Staff Applied AI Engineer, Product & Agent Performance at Arcadia | Torre

Staff Applied AI Engineer, Product & Agent Performance

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Provide your expected compensation while applying
location_on
Remote (for United States residents)
Shared by
Emma of Torre.ai
1 day ago

Responsibilities


Why This Role Is Important to ArcadiaArcadia’s data and analytics platform is used by hundreds of health systems, ACOs, payers, and life sciences organizations, touching tens of millions of patient lives. This role owns how our agentic capabilities perform at that same scale: accurate, transparent about their own confidence, and safe for the clinicians, care teams, and patients who depend on them.As a staff-level individual contributor, you will own the product-layer decisions that shape agent behavior, including prompting, retrieval and context, memory and state, evaluation, and escalation, while partnering with Product and Engineering on the systems that support them. Your work will help Arcadia make evidence-based launch decisions and scale responsible AI that is steerable, trustworthy, and ready for real healthcare workflows.What Success Looks LikeIn 3 monthsYou have established a production-grounded baseline for priority agentic workflows, with documented failure modes, severity-weighted evaluation rubrics, and a clear measurement planYou have mapped the current retrieval, context, memory, and escalation patterns and identified the highest-value opportunities to improve reliability, calibration, and costYou have earned trust across Product and Engineering by turning production evidence into clear, actionable recommendationsIn 6 monthsProduction-representative evaluation suites and regression checks inform model-change decisions for priority agentic workflowsYou have delivered measurable improvements in accuracy, reliability, steerability, latency, or cost for one or more priority workflowsHuman-review and escalation behavior has been validated under adversarial and edge-case conditions, with decision criteria and ownership boundaries clearly documentedIn 12 monthsArcadia has a repeatable product-layer AI performance practice that moves from production failure to diagnosis, experiment, evaluation, and release decisionHigh-severity regressions are caught earlier, and agent behavior is more transparent, calibrated, and trustworthy at scaleModel cards, intended-use guidance, limitations, and performance documentation are current and useful to product and customer-facing teamsWhat You'll Be DoingDesign and iterate on agent behavior across real, live workflows, including long-horizon, multi-turn agentic tasksDesign retrieval and context architecture so the right source data reaches a model in the right structure and agents remain grounded in real data rather than filling gaps with assumptionsDesign memory and state handling across multi-turn and multi-agent flows, determining what is carried forward, summarized, or dropped and whyCreate context and prompt templates that combine few-shot examples, structured formatting, and reasoning scaffolding for consistent agent behaviorImprove performance through prompting, tool-use strategy, and context construction, validated through direct experimentation rather than guessworkBuild and run evaluations against real production conditions to measure performance, regressions, failure modes, and edge casesAuthor evaluation rubrics, quality heuristics, and thresholds that weight failures by severity and cost, not just frequency, and monitor those measures against production behaviorDesign and validate escalation paths that route agents to human review based on confidence and uncertainty while preserving safety and consistency under adversarial and edge-case conditionsDesign for cost-aware performance alongside latency, reliability, and accuracy through efficient context construction and tool-call economyEvaluate and sign off on model changes by baselining current behavior, running comparative evaluations, and making the go/no-go call before a change reaches a customerMaintain product-level AI documentation, including model cards, intended use, limitations, and known failure modes, so customer-facing teams work from actual agent behaviorPartner closely with Product and product managers to ensure agents are not just capable, but steerable, trustworthy, and ready to scaleWhat You'll BringWe value equivalent practical experience that demonstrates the depth required for this staff-level role8+ years of production software engineering experience, including 3+ years of hands-on ownership of ML, LLM, or agentic systems in production, with direct experience in healthcare, finance, or another regulated industryDemonstrated ability to diagnose why an agent failed, correctly attribute the fix to instruction, retrieval, context, or memory design, and weigh failures by severity and cost rather than frequency aloneHands-on experience with RAG architecture, production-grounded evaluation frameworks, and fallback or human-in-the-loop logic for automated systemsWorking familiarity with AWS AI/ML services, including Bedrock and SageMaker, sufficient to build and evaluate effectively in Arcadia’s environmentEvidence-led judgment and the credibility to push back on launch decisions, paired with a builder’s instinct to run the experiment and move from a production failure to a fixWould Love for You to HaveExperience applying AI to healthcare data or workflows where safety, transparency, and calibrated uncertainty directly affect care teams or patientsExperience with long-horizon, multi-turn or multi-agent workflows and product-level AI documentation such as model cardsWhat You'll GetThe opportunity to define how agent performance, safety, and readiness are measured for production healthcare workflowsMeaningful ownership across prompts, context, memory, evaluations, and escalation patterns at product scaleA cross-functional role translating production evidence into AI improvements used across Arcadia’s platformA mission-driven company working to improve how patients receive careA flexible, remote-friendly culture with personality and heartEmployee-driven programs and initiatives for personal and professional developmentMembership in the talented, energized, diverse, and purpose-driven Arcadian communityAbout ArcadiaArcadia.io helps innovative providers and payers across the country transform healthcare to reduce cost while improving patient health. We do this by aggregating large amounts of disparate data, applying algorithms to identify opportunities to provide better patient care, and making those opportunities actionable by physicians at the point of care in near-real time. We are passionate about helping our customers drive meaningful outcomes. We are growing fast and have emerged as a market leader in the highly competitive population health management software market and have been recognized by industry analysts KLAS, IDC, Forrester, and Chilmark for our leadership. For a better sense of our brand and products, please explore our website.