All articles

Engineering

Why Our AI Has to Show Its Work Before It Calls Anything a Lead

A short story about how our first attempt at picking out real leads failed in two opposite directions, and the simple rule that fixed it. We are sharing it because the mistakes are ones almost everyone makes.

June 11, 20268 min readLeadPinger Team

Every tool in this space makes the same promise. We will read the flood of posts so you do not have to. The hard part was never the reading. The hard part is earning your trust about what we flag. This is the story of how our first attempt failed, what we learned, and the simple idea that fixed it. We are telling it in the open because the ways it broke are the same ways almost everyone breaks, and most companies will never show you theirs.

We started by grading ourselves

Before touching anything, we sat down and labeled a test set by hand. Around a hundred and twenty real posts pulled from the product, spread across different platforms and agents, each one sorted into four buckets. Not relevant. Worth watching. Worth acting on. Act right now. Labeling them taught us the first hard lesson. Most posts that match a sensible search are bad leads. So any tool that calls almost everything a lead is not being hopeful. It is broken in the one direction you will not notice until it has wasted your time.

The first failure: a model that says yes to everything

Ask one AI a simple question, is this a lead, and it will say yes far more often than it should. These models are trained to be agreeable, and 'yes, this looks promising' feels like the helpful answer. Even scoring each post from one to ten did not save us, because the scores all drifted high too. The model was basically nodding along.

The second failure: a model that says no to everything

So we added a second AI to check the first one, and we swung straight past the target. We told that second model to find the reason this is not a real lead. Give a smart model that job and it will always succeed, because you can find a reason to doubt anything if you look hard enough. The result was a system that approved almost nothing. Not a few good leads. Almost none. Every genuine buyer got argued out of existence by a model doing exactly what we told it to do.

A checker told to find problems will always find one. The fix was not gentler wording. It was forcing it to name a specific thing that actually failed before it could overrule anything.

The setup that finally worked

The version we run now has three steps, and each one has a narrow job.

  • First, a quick and cheap check that only asks one thing. Is this post even about the right topic? Its only job is to throw out the obvious junk before it costs us anything.
  • Second, the real read. This one sorts the post into a bucket and fills in the details, like what the person is struggling with and a suggested way to reach out. The catch is that it has to quote the exact lines from the post that justify its call.
  • Third, a stronger model that only takes a second look at the posts marked worth acting on. It can agree or knock something down a level, but it can never push something up, and to knock something down it has to point at a specific thing that failed. Healthy doubt, not a mission to destroy.

No proof, no lead, and we mean it in the code

If the second step marks a post as a lead but comes back with no quoted evidence, the system automatically drops it. This is not a polite request in a prompt. It is a hard rule in the code. A confident guess with nothing behind it simply cannot reach your queue.

What the numbers did

Measured against that hand labeled test set, the rebuilt system went from roughly a coin flip to being right about nineteen times out of twenty when it says act. We do miss some on the very strictest reading, but almost all of those misses land on the worth watching shelf instead of vanishing. The wider net is still there. It is just one shelf down, labeled honestly, where you can still find it.

That trade was on purpose. A lead queue is really a trust instrument. The first time it cries wolf, the user goes back to reading everything themselves and the product is dead to them. So we chose precision first, a visible warm shelf second, and evidence every single time. If a lead tool will not show you why it flagged something, it is worth asking what it would rather you not see.

Stop guessing who to reach out to.

Start with the people already asking for what you sell.

Free for 7 days · No credit card required