Methodology
Direct answer
In scope
What we measure
- Parallelism: max concurrent sessions on the vendor's published tiers and the practical ceiling of the underlying runtime.
- Self-heal: whether the tool can survive non-trivial DOM drift without a human edit, scored on vendor docs and named reviews.
- Mobile: first-class support for iOS and Android, real-device cloud quality.
- API coverage: whether the tool can also exercise HTTP / gRPC contracts well.
- Language breadth: number and quality of supported test-authoring languages.
- Maintainability: code organisation, fixture model, and the engineer-hour cost of running the suite at scale.
Out of scope
What we explicitly do not measure
We do not run reproducible benchmarks on identical apps. The numbers in vendor blogs claiming "3x faster than X" are mostly unverifiable, and we will not pretend ours are different.
We also do not score support quality, sales attitude, or partnership ecosystem. Those matter, but we have no fair way to measure them.
Honest treatment
Quote-only vendors
12 of 31 commercial vendors do not publish a price card. For these we display the published "starts at" statement only. We refuse to infer a rate from listicles or from competitor blogs. The TCO calculator treats quote-only licence as $0 with a clear note that the buyer must obtain a real quote.
Freshness
Re-verification cadence
Public price pages are re-checked monthly. The changelog is the audit trail. Every numeric claim on the site carries a per-claim verified-stamp linking to its source URL.