Comparing organizations takes more than running the same test.

Two results are only comparable when they come from the same ruler, the same calculation mode, the same engine version, and equivalent coverage. Without those four conditions, the difference between two organizations may be entirely methodological — and a methodological difference presented as a risk difference is worse than no comparison at all.

01The problem

Comparing multiple organizations — units, subsidiaries, acquisition targets, or competitors — with a criterion that survives the question of why one looks worse than another.

What is at stake

  • Comparisons made on different dates attribute to risk what is really a difference in collection.
  • Unequal coverage across organizations produces a ranking that measures collection reach, not exposure.
  • Mixing public and complete readouts creates a ranking that means nothing.
  • Comparing against an abstract threshold does not say whether a result is typical or exceptional for the sector.
02Evidence

GWK's sector ruler is derived from the operational base: 316,911 companies scored in a single run, aggregated by sector with an anonymity floor of K = 30. Every sector in the engine taxonomy survived the floor — none was excluded for having too small a cohort.

GWK's sector ruler is derived from the operational base: 316,911 companies scored in a single run, aggregated by sector with an anonymity floor of K = 30. Every sector in the engine taxonomy survived the floor — none was excluded for having too small a cohort.

Source
GWK operational base, sector aggregates
Date
July 12, 2026
Run
producao_324k_20260712
Population
316,911 companies in the operational base
Coverage
Public layer only: DNS, TLS, PQC readiness, headers, infrastructure, and observable subdomains.

Limitation: The ruler is a static reference base, not a continuous measurement. There is no automatic update between one run and the next, and this site promises no cadence.

03What makes up the comparison

A benchmark is not a report with more rows. It is a ruler built in batch, aggregated and anonymized — and it is that construction that gives meaning to any individual organization's position.

Public collection at scale

In production

The same collection infrastructure applied to every organization in the cohort.

Batch calculation

In production

A single engine version and a single calibration for the entire run.

Anonymized aggregation

In production

Median, dispersion, percentiles, and band distribution, respecting the anonymity floor.

Client benchmark

In production

Positioning a defined set of organizations against the sector ruler.

04What comes out

A comparison across entities you name, against anonymous aggregates. The ruler never exposes another organization's individual result.

  • A ranking of the assessed entities, with ruler and date stated.
  • Each entity's percentile against its own sector cohort.
  • Distribution by risk band within the set.
  • Dispersion and identification of outliers in the set.
  • An executive readout of what the comparison supports.
  • A sanitized dataset, when the scope calls for it.

What it depends on

  • A closed list of organizations and domains to compare.
  • Authorization to assess domains that are not yours.
  • Collection of all entities in the same window, so the comparison does not mix cutoffs.
05The conditions of comparison

The four comparability conditions are not a formality: breaking any one of them turns the ranking into an artifact of method.

Observed

Public collection saw

  • The same public surface, collected with the same modules.
  • The same collection window for the entire set.

Calculated

The IEQ engine derived

  • Score on the same scale, with the same engine version.
  • Percentile against the corresponding sector cohort.
  • Dispersion and relative position within the set.

Inferred

The model estimated

  • Each entity's sector, when not declared.
  • Sector prior for data shelf life.

Out of scope

Other evidence decides

  • That the better-ranked organization is more secure in general cybersecurity.
  • Comparison between results from different modes, public and complete.
  • Comparison with results from another run, calibration, or engine version.
  • Identification of any organization inside the sector aggregates.

What this comparison is measured against

Cutoff date
July 12, 2026
Run
producao_324k_20260712
Population
316,911 companies
Sectors
15 sectors in the engine taxonomy
Anonymity floor
K = 30

Limits of the comparison

  • The ruler is a static reference base, not a continuous measurement: there is no automatic update between one run and the next.
  • Comparison is always against anonymous aggregates, never against another organization's individual result.
  • A public-mode result does not compare to a complete-mode result, because the two readings start from different kinds of evidence.

Boundary

Where measurement ends

Adequacy programNot implemented

Measurement ends at: the technical change in the environment. Measuring exposure does not reduce it: reduction requires changing configuration, replacing certificates, switching negotiation policy, or migrating libraries — work carried out by the organization's own teams and suppliers.

This readout delivers

  • Relative position with ruler, mode, and date stated.
  • Dispersion and outliers within the assessed set.
  • A comparison criterion that withstands audit.

After the change, GWK

  • Re-collects public signals and recalculates IEQ on the same ruler, when contracted to do so.
  • States scope, mode, coverage, and run for both measurements, so the difference is interpretable.
  • Attributes the observed effect only to the scope actually changed and verified.

Not included

  • Executing the change: GWK does not alter the client's configuration, certificates, or infrastructure.
  • Deployment, assisted operation, or change management.
  • An adequacy program: it exists as a GWK engineering project, not as a contractable capability.

Define the set

A comparison starts from a closed list of organizations and from the decision the result will support. The more specific the set, the more interpretable the readout.