Editorial
How we assess AI agents
33 agents are catalogued here. 25 carry a score; 8 do not. This page explains the difference, because a rating is only worth what it is based on.
What our scores are based on
Every scored review on this site is currently a documentation-and-sentiment assessment, not a hands-on trial. We read vendor documentation, pricing pages, changelogs and public user feedback, then score against the criteria below. We are explicit about this because a review site is worth nothing if it overstates its own evidence — and because you deserve to know whether “4.3/5” came from using a product or from studying it.
Where we have used a product first-hand, it is badged Tested and says so on the page. Nothing carries that badge unless it is true.
The three coverage levels
We have used the product ourselves. Scored, and the review draws on direct experience.
Assessed from vendor documentation, pricing pages and public user sentiment. Scored on that basis, with sources linked in the “what users say” section of each review. This is where most of the site sits today.
Catalogued so you can find it, but not assessed — usually because there is too little public information to judge it fairly. These entries carry no score, emit no rating data to search engines, are excluded from head-to-head comparisons, and never rank above something we have assessed.
How a score is composed
Capability
25%How much of the job the agent finishes without a human stepping in, judged against what the vendor documents it can do and what users report it actually does.
Ease of setup
20%How far it is from signup to first useful output, and how much of that distance needs an engineer — including the steps vendor documentation glosses over.
Integrations
20%Breadth and depth of native connections, and whether an integration genuinely lets the agent act or only read.
Value for money
20%Cost per unit of completed work, not headline price. Consumption pricing is judged on what a realistic month costs, including runs that fail.
Reliability
15%Consistency across repeated runs, quality of error handling, and how the agent behaves when it is out of its depth.
On user reviews
The “what other users say” section of each review is an editorial summary of publicly available feedback, with sources linked at the bottom of that section. We do not write testimonials and attribute them to invented users. Where we have collected no reviews for an agent, the page says so rather than filling the space — and we publish no aggregate rating markup for ratings we did not collect.
Pricing accuracy
Each review records the date its pricing was last verified against the vendor’s own page. Pricing in this category changes often enough that anything older than a quarter should be treated as indicative. Always confirm before buying.
Commercial independence
Scores are set before any commercial discussion. Vendors cannot buy a listing, a position, or the removal of criticism. Several agents we rate highly have no affiliate programme at all — we link to them anyway and earn nothing. Read the full advertising disclosure for how the site is funded.
Corrections
If a review is wrong or out of date, email us. Vendors are welcome to dispute a score; we will re-examine it, and if we were wrong we will change it and note the change.