Everyone sells AI.
We tell you if it's worth it.

Independent outcome assurance for AI support agents.
Graded from the outside against your own published policy.

THE RELIABILITY PROBLEM

Vendors are billing you for outcomes.
Are you measuring them independently?

The market has shifted. AI agent vendors — Intercom, Ada, Decagon, Fin — have moved from per-seat pricing to outcome-based billing: charging per resolved conversation. That means your vendor now has a direct financial interest in how "resolution" is defined. And that definition is theirs, not yours.

A conversation marked resolved in your vendor's dashboard is not the same as a customer who received a correct answer. The gap between those two numbers is what goes unmeasured — and unmeasured, it compounds. Wrong answers that close tickets. Escalations that never surface. Warranties voided by bots that don't know their own policy.

5 / 5
brands in our benchmark with at least one high-severity bot failure
80+
policy probes run — every one graded against the company's own published text
4 / 5
brands whose bots omit warranty-void consequences — the single most consistent failure mode
HOW WE WORK

01

Read the policy.
Map the fault lines.

We read your published policy and identify every genuine risk: thresholds, exceptions, version changes, cross-product rules. We engineer every probe so we already know the correct answer before we ask.

02

Ask what a real
customer would ask.

Probes target the questions that matter most — warranty claims, return eligibility, product-specific conditions. No contrived scenarios. Only the questions real customers send when they are about to make an expensive decision.

03

Grade every response
against your own words.

Every failing answer is cited verbatim alongside the policy text that contradicts it. Not our opinion — the exact gap between what your bot said and what your policy says.

We don't sell AI.
We tell you which AI
is worth keeping.

Winnower is independent. We have no vendor relationships, no referral arrangements, no financial interest in your choice of platform. Our only product is an accurate grade — and our neutrality is what makes it worth anything.

Everything we do is grounded in the policy text you have already published. We work from the outside: no system access, no integration, no data sharing required.

If your bot is telling customers something your policy contradicts, we show you exactly where and exactly how.

For design partners, Phase 2 goes further. With a read-only connection to your support data — or a simple CSV export of resolved tickets — we measure your actual resolution rate against what your vendor dashboard reports. That is when the real number surfaces.

Find out what your bot is
actually telling your customers.