Our methodology

The weighted criteria, the verification process, and the disclosures behind every verdict on this site, plus a running record of the calls we got wrong.

How do we evaluate an AI tool? We run it on real work from our own operation, score it against six weighted criteria, cross-check every vendor claim against a primary source, and re-check the verdict on a schedule. No vendor benchmarks taken at face value, no paid placement, and no "it depends" without the specifics.

TL;DR. Verdicts come from hands-on runs scored on six weighted criteria, verified against primary sources, and re-checked on a schedule. When we get one wrong, we say so.

How we score a tool

Six criteria, weighted. The weights sum to 100% and are the same for every tool in a category. We do not re-weight to flatter a favorite.

30%

Real-world performance
How the tool does on actual tasks from our own operation, not the vendor's benchmark deck. The single heaviest input.

20%

Reliability
Does it produce the same quality across repeated runs? We count failed runs, dead-end loops, and silent degradations, not just best-case demos.

15%

Cost transparency
The total cost of real usage, and whether the pricing page tells the truth. Hidden per-seat minimums and token-tier multipliers cost points.

15%

Integration & fit
How cleanly it drops into an existing workflow: API quality, exports, and whether it locks your data behind its own walls.

10%

Documentation & support
Can you actually get unstuck? We test the docs by using them, and we note when support is a dead end.

10%

Data handling
Where your data goes, how long it is retained, and whether the tool trains on it. Vague answers here are treated as a "no".

How we verify a verdict

A verdict is a claim, so it carries a chain of custody, from a hands-on run to a published, re-checkable trail.

Step 1

Hands-on first
We run the tool on real tasks before we write a word. A verdict never rests on a press release or a feature list.

Step 2

Cross-check the claims
Every pricing number, limit, and capability is checked against a primary source: the vendor's own docs, changelog, or pricing page, with the date we read it.

Step 3

Score against the criteria
The run and the sources feed the six weighted criteria above. The score is the verdict; the reasoning is published with it.

Step 4

Re-check on a schedule
Agents re-read each profile on a cadence, material changes are gated behind human review, and every "verified" badge links back to the check behind it.

What we disclose

Trust is only worth as much as what we admit. Three things we put in writing.

We build competing products

Neomanex builds and sells AI products: ConvOps, Gnosari, and others. When one appears in the directory it is labeled as ours, scored on the same six criteria as everything else, and never ranked above a competitor that beats it on the numbers.

No paid placement

Nobody can buy a listing, a higher rank, or a softer verdict. There are no affiliate-driven rankings and no sponsored entries in the directory. If that ever changes, it will be disclosed on this page first.

Where we have gaps

As of July 2026, some verdicts are lighter than others. Enterprise features locked behind a sales call, and tools we have not yet run in our own production, carry a provisional verdict and say so on the profile. We would rather flag an untested claim than launder it into a confident one.

What we got wrong

A verdict stated as permanent is a verdict waiting to be wrong. We keep the misses in public.

Verdict changed · May 2026

We over-recommended RAG for long-document analysis

For months we recommended a retrieval pipeline (RAG) for long-document analysis, full stop. When Claude 4.6 Opus shipped a 1M-token context window in May 2026, we re-ran our comparison. Below roughly 50 documents, single-pass analysis beat retrieval on both accuracy and setup cost. We had been telling people to build and maintain chunking infrastructure they did not need.

What we learned: Scope a verdict by the variable that actually decides it. We now split this recommendation by corpus size instead of shipping one answer for every case. Under ~50 documents, go single-pass; a living corpus at volume, keep RAG.

Read the full verdict change

Questions about how we work

The things people ask before they trust a verdict.