Software Engineer, Agent Intelligence
About this role
About Sage Care
Sage Care is a fast-growing, early-stage healthcare startup founded by exceptional leaders from Apple, Uber, Carbon Health and backed by top-tier venture capital (General Catalyst, Chelsea Clinton). With a strong customer pipeline, Sage Care is transforming healthcare by simplifying care navigation.
Our platform makes it easier for patients to find the right doctor and helps providers focus on those who need them most through harnessing the latest AI innovations.
Building on our successful collaborations with health systems across the U.S., we have expanded internationally to the MENA region. We are now partnering with health systems there to deploy our AI-powered care navigation platform.
About the Role
Every day, our AI agents handle real patient conversations for hospital systems. Those conversations contain everything needed to make the agents better. Today, much of that learning is still manual: humans review calls, identify issues, investigate failures, and work with engineers to improve agent behavior.
Your mission is to build the systems that close that loop.
You will own the intelligence layer of our agent platform: the evaluation pipelines, feedback systems, and ML infrastructure that turn production conversations into measurable, continuous improvement.
This role sits at the intersection of AI evaluation, ML pipelines, and production quality. You will work closely with the engineers who own the agent runtime, with the operations teams who review calls, and with product leaders who decide what good looks like for each hospital partner.
What You'll Do
Build the evaluation and feedback platform
Design systems that continuously analyze production conversations, identify and cluster quality issues
Build evaluation pipelines that measure agent performance across the dimensions that matter for patient care
Develop workflows that transform human feedback into actionable improvements
Design mechanisms for routing issues to the appropriate AI, engineering, or operational owners
Make failures diagnosable
Investigate production failures and identify root causes across transcription, reasoning, retrieval, and orchestration
Build tooling that helps engineers quickly understand why an agent behaved a certain way
Establish quality metrics and reliability standards for production agents
Automate the learning loop
Build ML pipelines that reduce the manual effort required to improve agents
What We're Looking For
Required
5+ years of software engineering experience
Experience building production systems
Experience working with LLMs, AI agents, or conversational AI applications
Experience in one or more of the following:
building AI evaluation platforms or frameworks
developing feedback systems for ML and LLM applications
creating observability, reliability, or quality tooling for AI products
building ML pipelines that improve model or agent performance
Strong backend engineering skills and systems thinking
Experience working with ambiguous problems and defining solutions from first principles
Nice to Have
Experience with evaluation frameworks and model quality measurement at AI-native companies
Experience with voice AI systems
Experience designing human-in-the-loop workflows
Experience turning emerging research or new techniques into tested prototypes
Experience building internal platforms used by engineering teams
Experience in healthcare or other regulated, safety-critical domains
What Success Looks Like
Failures are detected by systems, not discovered by people
Every production issue has a diagnosable root cause and a clear owner
Human feedback measurably changes agent behavior, with the lag between the two shrinking over time
