About the Company :
The company is building autonomous AI systems for one of the most complex and high-stakes industries in the world, US healthcare revenue cycle. It is backed by leading venture capital investors and has built clinically driven AI that generates highly accurate medical codes at scale. What traditionally required large teams and weeks of manual effort can now be done autonomously with speed, precision, and reliability. The company has seen rapid growth alongside strong customer retention. Its customers expand because the impact is measurable, including improved accuracy, reduced coding costs, and faster revenue realization. This is applied AI in live healthcare systems, not experimentation. As the company scales, it is building the foundational data layer that will power its next phase of growth.
Why Explore This Career Opportunity :
Joining this team means working on a hard, real-world AI problem with a small, high-ownership team. You will build systems that make engineering faster, quality stronger, and production more reliable. The company values structured thinkers, high-agency builders, and people who can convert messy workflows into simple, scalable systems.
About the Role :
You will build production-grade internal AI agents that reduce repetitive engineering work and improve reliability. This is a hands-on engineering role focused on building production AI systems that engineers rely on every day. The person must be able to design, ship, evaluate, monitor, and iterate agents used daily by engineering teams for PR review, ticket triage, test generation, go-live validation, customer call summarization, regression checks, incident response, anomaly detection, and post-mortems.
Key Responsibilities :
1. Agent Development :
- Build and maintain internal agents for PR review, issue triage, testing, go-live validation, customer calls, regression checks, incident response, anomaly detection, and post-mortem drafting.
- Design each agent with clear success metrics, feedback loops, monitoring, cost tracking, and kill criteria.
- Create reliable agent workflows with structured outputs, tool use and function calling, retries, fallbacks, and audit trails.
- Measure adoption, quality, false positives, latency, and time saved for every agent.
2. Agent Platform and Integrations :
- Build integrations with GitHub, Slack, Jira or Linear, Notion, PagerDuty, CI/CD systems, and BigQuery.
- Own webhook and event-driven triggers for PRs, issues, deploys, alerts, and transcript availability. Create reusable patterns for prompt management, prompt versioning, evaluation datasets, and model selection.
- Keep LLM usage cost-disciplined through caching, routing by model tier, context management, and per-agent cost visibility.
3. SRE AI Safety Net :
- Build incident response agents that query logs, deploy events, error spikes, and accuracy signals to generate first-line triage within minutes.
- Build post-mortem and anomaly detection agents using BigQuery-backed reliability data.
- Design safe runbook automation where irreversible production actions require human approval. Partner with SDET and DevEx engineers to define the quality bar and data foundation for SRE-AI workflows.
Requirements :
- 3+ years of software engineering experience with at least 1 year building production LLM-powered systems or agents used by real users.
- Strong Python engineering skills, with the ability to write maintainable, tested, production-quality services.
- Hands-on experience with LLM APIs, preferably Anthropic Claude, including tool use and function calling, structured outputs, streaming, prompt caching, and multi-turn context handling.
- Experience with at least one agent or orchestration framework such as LangChain, LlamaIndex, DSPy, CrewAI, or strong custom orchestration experience.
- Strong REST API and webhook integration experience across tools such as GitHub, Slack, Jira or Linear, Notion, PagerDuty, or similar.
- Good understanding of LLM evaluation, including defining success metrics, measuring false positives and negatives, creating feedback loops, and iterating from telemetry.
- Comfortable with Docker, CI/CD, observability, logging, alerts, and operating services in production.
- Strong SQL and BigQuery ability for querying logs, deploy events, incident timelines, accuracy signals, and time-series reliability data.
- Clear written communication, with the ability to write design docs that define scope, non-goals, evaluation methodology, owners, and risks.
Nice to Have :
- Experience with vector databases or retrieval systems such as pgvector, Pinecone, Weaviate, Chroma, or equivalent.
- Experience building engineering productivity, DevEx, or SRE tooling.
- Exposure to healthcare, medical coding, or high-accuracy regulated workflows.
- Experience with incident management systems, runbooks, or production on-call workflows.
- Experience running internal demos, enablement sessions, or playbooks for engineering teams.
Benefits :
- Exceptional total compensation.
- Hybrid setup.
- Work with a world-class team on a very forward-looking AI problem.
- Leading healthcare benefits.
- Be part of a team working on AI applied to healthcare.
- Work in a collaborative environment with opportunities for professional growth.
- Contribute to impactful projects that improve patient care and clinical efficiency.
(ref:hirist.tech)