How Search Engine Watch tests and reviews products

The evidence, scoring and commercial-independence standard behind our reviews, comparisons and recommendations.

Search Engine Watch editorial standard

We will not make the claim larger than the test.

Methodology version 1.0 · Last reviewed August 2026

When we call a product tested, we want you to know what we tested, how we tested it and what we could not verify. This page is the standard behind our reviews, comparisons and recommendations.

The publication that became Search Engine Watch did not begin with a feature list or a press release. It began with an experiment.

In 1996, Danny Sullivan changed webpages and watched how different search engines responded. He compared their stated rules with their observable behaviour and reported the differences. That instinct helped establish this publication: look closely, test honestly and do not pretend an opaque system is simpler than it is.

Three decades later, we review a different kind of search product. SEO platforms now estimate traffic, group keywords, monitor rankings, audit websites, generate content, count links and measure visibility inside AI answers. The dashboards look precise. The underlying data often is not.

That makes our responsibility straightforward. If we describe a tool as the best, assign it a rating or ask you to spend money through one of our links, we must be able to explain the evidence behind that judgement.

Search Engine Watch’s history is not a reason to trust us without question. It is a reason to expect more from us. We cannot ask Google, Microsoft, AI platforms and software companies to be transparent while hiding how our own conclusions were reached.

This methodology applies to every review and recommendation carrying an SEW evidence label or numerical rating. Older articles that predate this system may be marked as using our legacy methodology until they are retested. Updating a price, feature or screenshot does not turn an older article into a newly tested review.

On this page

Our four evidence labels

Not every useful article requires the same kind of evidence. A year of daily use can reveal problems that a controlled afternoon cannot. A standardised benchmark can make products comparable in a way that personal familiarity cannot. Careful research can map a market, but it cannot tell us how a product feels to use.

We show those differences instead of flattening them into the word “reviewed.”

Label What it means What it does not mean
Production Tested We used the product in genuine editorial, SEO, analytics or business work for at least 30 calendar days or one complete natural work cycle, with at least five substantive sessions. The article states the actual period and use case. It does not mean we tested every feature or that our use case represents every customer.
Bench Tested We ran the product through a versioned SEW test bench using the same tasks, reference data and scoring rules applied to comparable products. It does not guarantee long-term reliability unless the monitoring period is stated.
Hands-On Checked We logged in and completed the specific tasks listed in the article, but did not finish the full comparative bench or a substantial production-use period. It is not a full benchmark. Opening an account or attending a demo is not enough.
Research Verified We checked current facts through official documentation, pricing, demonstrations and other reliable evidence, but did not directly use the product enough to describe the experience. It is not hands-on testing and receives no SEW numerical rating.

If a statement comes only from the company, we identify it as a company claim. If we cannot verify a material claim, we either say so or leave it out.

What may receive an SEW rating

We do not rate a product because a five-star box looks useful beside an affiliate button.

A numerical SEW rating requires direct evidence for all critical tests in its category and at least 70% of the weighted score. If we have not tested a core capability, we do not award a precise overall number that implies we have.

A product may receive a primary “best” recommendation only when it has been Production Tested or Bench Tested for the use case it wins. Research-verified products can appear in a market map or an “also considered” section, but they do not quietly outrank products we actually tested.

Our scores are tied to a category and a test-bench version. A 4.4 for a keyword research tool is not automatically comparable with a 4.4 for a website builder. Scores produced under incompatible major versions of a test bench are not directly comparable until the older product is retested.

How we choose products

No serious publication can test every product in a growing software market. Pretending otherwise would reward volume rather than evidence.

We begin with a longlist built from market usage, reader interest, distinctive capabilities, pricing, availability, new releases and the needs of the audience for that article. We do not require an affiliate programme. A higher commission does not buy a place on the shortlist, and the absence of a commercial relationship does not disqualify a product.

We then apply eligibility criteria appropriate to the category. These may include an active product, clear ownership, a working route to support, adequate documentation, current security information, transparent plan terms and the ability to perform the job being assessed.

The shortlist moves through two stages:

  1. Screening: We verify that the product is real, current, relevant and capable of the claimed core task. This removes abandoned products, misleading offers, duplicate services and tools that do not fit the reader’s use case.
  2. Finalist testing: We run the products most likely to earn a recommendation through the relevant test bench. The number tested is stated in the article.

Large lists may also include research-verified alternatives that serve a niche the finalists do not. They are labelled separately. We would rather publish “six tested picks and nine researched alternatives” than imply that 15 logins received equal scrutiny when they did not.

How we obtain access

We may test products through an existing subscription, a plan purchased by SEW, a free trial or an editorial account supplied by a company. We disclose the source and the plan tested.

A supplied account does not buy favourable treatment. Companies do not choose our test data, supervise our work or approve the verdict. We may ask a company to correct a factual point or explain an unexpected result, but the final judgement remains ours.

The plan matters. A free trial with relaxed limits is not necessarily representative of a paid subscription. An enterprise demonstration is not evidence that the entry plan includes the same features. We record the tier, seats, credits, add-ons and material restrictions so readers can judge the experience we actually had.

The SEW Test Record

Scored reviews and major recommendations carry a test record near the verdict. It includes, where relevant:

  • The evidence label.
  • The named reviewer or test lead.
  • The product plan and version tested.
  • How access was obtained.
  • The testing and price-check dates.
  • The test-bench name and version.
  • The device, browser, country, language and other material conditions.
  • The reference sites, queries or prompt sets, described without exposing confidential data.
  • The tests we did not complete.
  • Any affiliate, coupon, advertising, expedited-testing or other commercial relationship.

This is the software equivalent of showing the hardware, drivers and benchmark settings in a component review. A reader should not have to guess which version produced the result.

What our test bench measures

Every category has its own module, but our software reviews begin with a common set of questions.

Can the product complete its core job?

We define tasks before testing. For an SEO platform, that might mean creating a project, connecting first-party data, crawling a site, finding a ranking change, exporting a report and understanding the subscription cost. For a website builder, it might mean creating, publishing, crawling and measuring the same reference site on each platform.

We record success, failure, error handling, time to a usable result and any intervention required. We do not award points because a feature appears in the navigation.

Is the output dependable enough to guide a decision?

Dependability is not the same as agreement with whichever famous tool happens to be used as a reference.

Where ground truth exists, we use it. A controlled audit site can tell us exactly which technical faults were planted. A same-time capture of a results page can tell us which rank appeared under the stated conditions. A known set of live and removed links can show whether a backlink index found them.

Where ground truth does not exist, we do not invent it. No third-party SEO platform has direct access to every website’s analytics, every search, every link or every answer shown by a stochastic AI system. We therefore examine coverage, repeatability, freshness, outliers, internal consistency and whether the output leads a competent user toward a sound decision.

A competitor traffic estimate is an estimate, even when it contains an exact-looking number. Search volume is modelled and grouped. Keyword difficulty is a proprietary score, not a physical measurement. AI visibility depends on the prompts, models, accounts, locations and moments sampled. Our language and scoring reflect those limits.

What is it like to use for real work?

Controlled tasks expose comparable performance. Real work exposes friction.

We assess setup, navigation, workflow continuity, bulk actions, exports, collaboration, permissions, reporting and the amount of cleaning required before data is useful. We note when an interface makes the scope of a report unclear, when limits interrupt an ordinary task or when a polished dashboard gives a stronger impression of certainty than the data supports.

We distinguish subjective judgement from measurement. “The export took 42 seconds” is a measurement. “The workflow felt unnecessarily fragmented” is the reviewer’s interpretation, supported by the steps recorded.

What does the useful version really cost?

Starting prices are often poor buying guidance.

We build at least one realistic customer scenario and calculate what that user would pay. We include billing frequency, seats, tracked keywords, locations, devices, projects, reports, crawl credits, AI prompts, API usage, add-ons and material overages. We distinguish an introductory discount from the recurring price and state the country, currency, taxes and price-check date.

Value is judged against the job completed, not the length of the feature list.

Can users leave with their work?

We test material export and cancellation routes where the plan and testing period allow it. We check whether data can be exported in useful formats, whether scheduled reports are portable, what happens to stored projects and whether cancellation can be completed without contacting sales.

We do not create a fake support crisis. When support is tested, we submit a real, reproducible question and record first response, substance, escalation and resolution.

Are privacy and security claims supported?

We inspect published security, privacy, retention and subprocessors information relevant to the product. For business software, we may check the availability of multi-factor authentication, SSO, role controls, audit logs, a data processing agreement and recognised assurance reports.

This is a documentation and product-controls review, not a penetration test. We do not claim a product is secure merely because it names a certification, and we do not conduct intrusive testing without explicit authorisation.

How we test common search and marketing products

All-in-one SEO platforms

We use an owned or authorised website so that modelled output can be compared with first-party data without exposing a client. The standard suite covers project setup, Search Console and analytics connections, keyword research, rank tracking, site audit, competitor research, backlinks, reporting, exports and the plan limits encountered.

We also test whether the platform keeps scope clear. Domain, subdomain, subfolder, exact URL, country, language, device and date can change the meaning of a report. A useful platform should not make those changes easy to miss.

An all-in-one tool does not earn a high score simply for having more menu items. We give more weight to the jobs the intended customer will perform repeatedly.

Keyword research tools

We use a fixed seed set across different intents, levels of specificity, markets and languages. We retain the raw export, normalise casing and punctuation, identify close variants and measure how many genuinely distinct topics remain after duplicates and near-duplicates are removed.

We inspect relevance, useful coverage, apparent scraping artefacts, language classification, intent labels, trend data, SERP context, filters and export quality. We test unusual terms and singular/plural pairs because confident-looking outliers often reveal how carefully the numbers need to be handled.

We do not announce a winner based on who returns the largest keyword count. Ten thousand rows can be less useful than two thousand if much of the difference is duplication or noise.

Rank trackers

We define queries, target URLs, locations, languages, devices and the search engine before the test. We run the tracker over multiple days and validate a documented sample against same-time result captures or an appropriate reference service using equivalent settings.

We measure exact agreement, near agreement, missing observations, update delay, handling of local packs and other search features, tagging, competitor tracking and reporting. Because live results vary, disagreement is investigated before it is scored. A single manual search from a logged-in browser is not treated as universal ground truth.

Site audit and crawling tools

Where possible, we use a controlled reference site containing a known set of crawl, indexation, canonical, redirect, metadata, structured-data, hreflang and JavaScript conditions. This gives us something rare in SEO software testing: a known answer.

We measure the proportion of planted issues found, false positives, crawl completion, rendering behaviour, configuration clarity, explanations and prioritisation. Finding more “issues” is not automatically better. A crawler that turns intentional configurations into hundreds of urgent warnings can waste more time than it saves.

Backlink research and monitoring

No accessible index contains every link on the web. We therefore avoid “largest database” claims based solely on a vendor’s published index size.

We use a documented set of known live, redirected, nofollow, sponsored and removed links where available. We examine discovery, freshness, status classification, duplicate handling, link context, export quality and the usefulness of competitor intersection reports. Search Console link samples may provide additional reference data, but we do not treat them as a complete inventory.

Content optimisation and AI writing tools

We give each product the same brief, source pack, target reader and constraints. We assess the research trail, SERP analysis, topic coverage, unsupported claims, duplication, citation quality, editorial control and the amount of human repair required.

We do not publish generated copy merely to prove that a button works. When output quality is scored, reviewers examine it without being told which product produced which draft where practical. Ranking after publication is not attributed to one content score without a controlled basis.

AI detectors and “humanisers” receive especially cautious treatment. We do not accept a detector’s judgement as proof of authorship, and we do not reward a tool for helping users disguise deceptive content.

AI visibility and generative search monitoring

AI answers can vary across repeated runs, accounts, models, locations and dates. A single prompt is not a market share study.

We use a fixed, versioned prompt set, record the exact models and settings available, run prompts more than once and capture brand mentions, citations, sentiment or position according to a predefined rule. We then compare what the monitoring tool recorded with the answers directly observed in the sampled runs.

The result measures performance on the disclosed sample. We do not convert it into a claim about a brand’s total visibility across an entire AI platform.

Analytics and behaviour tools

We install products on the same reference property or equivalent test environments, verify event collection and consent settings, and run documented user journeys. We inspect implementation effort, data delay, filtering, sampling or limits, session reconstruction, exports and the difference between what each product is designed to measure.

Tools with different purposes are not marked inaccurate because their totals do not match. A search click, analytics session and replay are different events produced by different systems.

Website builders, hosting and WordPress SEO products

We build or clone a controlled reference site with the same content, structure and required functions. We test what the platform generates by default and what a competent non-developer can change without unsupported workarounds.

The suite may include crawlability, index controls, canonical handling, redirects, structured data, sitemap and robots controls, image handling, Core Web Vitals under stated conditions, staging behaviour, migrations, backups, uptime, support and the full cost of the required stack.

For plugins, we record the WordPress, PHP, theme and plugin versions and check for conflicts in an isolated environment. We do not test potentially destructive changes on a live publication.

Courses, conferences and professional services

Software methods cannot simply be pasted onto people and services.

We do not say we tested a course unless a reviewer completed enough of it to judge the curriculum. We do not rate a conference we did not attend as though we had. We do not call an agency the best because its own case study reports a large percentage increase.

Depending on the category, evidence may include attendance, purchase, curriculum inspection, instructor credentials, student work, verified client references, contract and pricing review, mystery-shopping enquiries conducted honestly, and independent outcome data. The article states which evidence was used. Services with research-only evidence receive no lab-style score.

How we calculate ratings

Our standard software score begins with seven dimensions. Category test benches may adjust the weights before testing begins, and the article links to the applicable version.

Dimension Standard weight What it covers
Data reliability and decision usefulness 25% Ground-truth performance where available, consistency, freshness, sensible handling of uncertainty and usefulness of output.
Core task performance 25% Whether the product completes the important jobs for its intended user.
Workflow and usability 15% Setup, clarity, speed of routine work, bulk actions and avoidable friction.
Value and plan limits 15% Realistic cost, allowances, add-ons, overages and value for the target use case.
Integrations, exports and data ownership 8% Connections, API, export quality, portability and collaboration.
Reliability and support 7% Errors, stability, documentation and support resolution.
Privacy, security and commercial transparency 5% Relevant controls, documentation, ownership and clarity of terms.

The weighted score is converted to a five-point rating by dividing the underlying score by 20. We show the final rating to one decimal place and retain the underlying 100-point calculation in our test record.

Our ratings mean:

SEW rating Meaning
4.8-5.0 Exceptional under the stated test and use case. Material weaknesses are rare.
4.5-4.7 Category-leading, with limited compromises.
4.0-4.4 Strong and recommendable for the right user, with meaningful limitations disclosed.
3.5-3.9 Good in a defined use case, but not an automatic recommendation.
3.0-3.4 Adequate or narrowly useful. Several buyers should choose elsewhere.
2.0-2.9 Difficult to recommend without an unusually specific reason.
Below 2.0 Failed too many important requirements for the stated use case.

We do not select a verdict and then bend the component scores until the arithmetic agrees. Critical failures can limit a rating even when many secondary features perform well. An unknown critical result does not receive a neutral score; it prevents the overall rating until the missing evidence is resolved.

The numerical rating is accompanied by an evidence-confidence label. High confidence normally requires repeat measurements, a completed critical suite and/or extended production use. Moderate confidence means the critical suite was completed once with stated limitations. Insufficient evidence means no numerical rating.

Editorial judgement still matters

A benchmark does not remove judgement. It disciplines it.

Weights reflect the needs of a defined user. An enterprise platform, a small-business tool and a specialist crawler should not be judged as though they serve the same buyer. We explain whom a recommendation is for and which compromises made it lose points.

Measurements also need interpretation. A tool may find fewer keywords but produce a cleaner usable set. Another may report a technically correct warning so often and so urgently that it becomes operationally unhelpful. Our reviewers are responsible for explaining that difference rather than hiding behind a composite number.

Affiliate links, coupons, access and advertisers

Some SEW articles contain affiliate links or reader coupons. We may earn a commission if a reader buys through one of them. That commercial relationship is disclosed in the article and does not change the rating.

We do not accept payment for a favourable review, a particular score or a ranking in an editorial list. Advertising and sponsorship do not grant control over our test selection or conclusions. Sponsored content is labelled and does not receive an SEW review rating or testing badge.

We may charge a disclosed fee to prioritise an eligible product in our testing queue or to cover unusually demanding testing work.

That fee purchases editorial time, scheduling and testing. It does not guarantee publication, a positive verdict, a rating, a badge, inclusion in a recommendation list or any position within one. If we accept such a fee, we disclose it in the article’s SEW Test Record.

We may negotiate a better public offer for readers after the editorial decision has been made. The existence or size of that offer does not determine the verdict. A useful test is simple: if two products exchanged affiliate payouts tomorrow, our recommendation should remain the same.

How we use AI in reviews

AI may help us transcribe notes, compare structured exports, normalise data, locate anomalies or organise material. It may not invent product use, fill gaps in a test ledger, create a fake support exchange or turn a company claim into an SEW observation.

A human reviewer remains responsible for the test, evidence, wording and verdict. If we cannot show that a material experience occurred, we do not write as though it did.

Updates, retesting and test-bench versions

Software changes after publication. Prices move, limits change, interfaces are rebuilt and AI features can be replaced within weeks.

Every review published under this standard shows a test date and, where relevant, a separate price-check date. A light factual update does not become a new hands-on test. When a product change could alter the verdict, we retest the affected areas and record what changed.

Our test benches use version numbers:

  • A patch update, such as 1.0 to 1.0.1, corrects a procedure without changing the basis of comparison.
  • A minor update, such as 1.0 to 1.1, adds or revises tests while preserving broad comparability where possible.
  • A major update, such as 1.x to 2.0, changes the method or weights enough that older overall scores may no longer be directly comparable.

We publish a change log for material methodology changes. Older reviews display their original bench version. We do not silently apply a new score to old observations.

Corrections and challenges

We welcome specific corrections from readers, practitioners and the companies we cover. Evidence is more useful than indignation, from either side.

If a factual error is confirmed, we correct it and add a note when the change is material. If a company disputes a measured result, we review the test conditions, repeat the test where justified and report the outcome. A company may explain its product. It may not negotiate the conclusion.

What we will not do

  • We will not call a product tested because an editor opened the homepage or watched a sales demonstration.
  • We will not describe every item in a large list as tested when only the finalists received hands-on work.
  • We will not mistake a long feature list for product quality.
  • We will not treat another SEO tool as universal ground truth.
  • We will not hide the plan, region, dates or important limits behind a precise score.
  • We will not sell SEW ratings, winner positions or editorial badges.
  • We will not allow polished language to outrun the evidence.

Search Engine Watch was created to make an opaque industry more understandable. That remains the job.

Readers should not have to trust us simply because we say we are Search Engine Watch. They should be able to inspect how we reached the conclusion. If a publication with our history cannot draw that line clearly, we cannot reasonably ask the rest of the search industry to do so.