Our methodology
The weighted criteria, the verification process, and the disclosures behind every verdict on this site, plus a running record of the calls we got wrong.
How do we evaluate an AI tool? We run it on real work from our own operation, score it against six weighted criteria, cross-check every vendor claim against a primary source, and re-check the verdict on a schedule. No vendor benchmarks taken at face value, no paid placement, and no "it depends" without the specifics.
TL;DR. Verdicts come from hands-on runs scored on six weighted criteria, verified against primary sources, and re-checked on a schedule. When we get one wrong, we say so.
How we score a tool
Six criteria, weighted. The weights sum to 100% and are the same for every tool in a category. We do not re-weight to flatter a favorite.
30%
20%
15%
15%
10%
10%
How we verify a verdict
A verdict is a claim, so it carries a chain of custody, from a hands-on run to a published, re-checkable trail.
Step 1
Step 2
Step 3
Step 4
What we disclose
Trust is only worth as much as what we admit. Three things we put in writing.
We build competing products
Neomanex builds and sells AI products: ConvOps, Gnosari, and others. When one appears in the directory it is labeled as ours, scored on the same six criteria as everything else, and never ranked above a competitor that beats it on the numbers.
No paid placement
Nobody can buy a listing, a higher rank, or a softer verdict. There are no affiliate-driven rankings and no sponsored entries in the directory. If that ever changes, it will be disclosed on this page first.
Where we have gaps
As of July 2026, some verdicts are lighter than others. Enterprise features locked behind a sales call, and tools we have not yet run in our own production, carry a provisional verdict and say so on the profile. We would rather flag an untested claim than launder it into a confident one.
What we got wrong
A verdict stated as permanent is a verdict waiting to be wrong. We keep the misses in public.
We over-recommended RAG for long-document analysis
For months we recommended a retrieval pipeline (RAG) for long-document analysis, full stop. When Claude 4.6 Opus shipped a 1M-token context window in May 2026, we re-ran our comparison. Below roughly 50 documents, single-pass analysis beat retrieval on both accuracy and setup cost. We had been telling people to build and maintain chunking infrastructure they did not need.
What we learned: Scope a verdict by the variable that actually decides it. We now split this recommendation by corpus size instead of shipping one answer for every case. Under ~50 documents, go single-pass; a living corpus at volume, keep RAG.
Read the full verdict changeQuestions about how we work
The things people ask before they trust a verdict.