All work
GenAI evaluation · LLM-as-a-judge

Quality- and cost-aware GenAI evaluation for Signify Health

For Signify Health, a CVS Health company, I built Azure ML foundations, a Snowflake call-center rating application, and an LLM-as-a-judge architecture designed around judge quality per dollar.

  • +8% predictive accuracy. Signify Health, a CVS Health company.
  • +15% conversion rate. Signify Health, a CVS Health company.
  • 25% faster call resolution. Signify Health, a CVS Health company.

Context & problem

Signify Health, a CVS Health company, needed a consistent way to rate call-center interactions while balancing predictive quality, serving cost, and operational usefulness. I delivered the work through Tredence.

What I built

Three layers, in sequence:

  1. Platform foundation: Set up the Azure ML platform supporting the ML work.
  2. The application: Built a Snowflake-based LLM application rating call-center interactions in production.
  3. The evaluation layer: Designed an LLM-as-a-judge architecture on Snowflake with serving economics as a first-class constraint, optimizing judge quality per dollar.

Measured result

+8% predictive accuracy, +15% conversion rate, and 25% faster call resolution versus the prior scoring approach.

Reach out to learn more.