1. What the Index measures
The Index ranks published sustainability reports on how well they communicate, not on how sustainable the company is.
The distinction matters. Investors, rating agencies and analysts read corporate sustainability reports with AI before a person does. MSCI applies AI across more than 17,000 issuers; RepRisk screens 150,000 sources daily on behalf of 80% of the world's largest investment managers (sources in section 9). When a report does not carry the work through to those systems, the work is discounted. The Index measures whether it carries.
It does not measure emissions, targets, controversies or performance. Two companies with identical programmes can rank far apart, because one explained the programme in a way that survives a machine read and the other did not.
2. The universe
A ranking is only defensible if the set of companies is defensible. The universe is defined by rules, published with the ranking, and applied without exception.
Rules:
- All stored reports across all countries and years are included, subject to the client, confidentiality and documented source exclusions below. Multiple reports from the same company are included. There is no listing, report-age or market-value cutoff.
- Companies without a recorded exchange appear under sector Private and industry Private. Their original industry classification still determines the alignment benchmark.
- Ranking requires an existing assessment, at least 20,000 extractable characters, all gateway measurements, people readability and an industry benchmark. Reports missing these remain visible without a rank and with a reason.
- Clients are excluded: a company recorded as a client, with a client subscription, or with a report privately uploaded by an outside person is omitted entirely. Reports gathered by our own team remain eligible.
- The 51 documented source exclusions from 10 September 2026 remain excluded by company and report identity. The original evidence is retained; another report from the same company can qualify.
- Missing public identifiers or source links do not hide reports. A preliminary-analysis link requires an assessable report with a unique, nonempty identifier and a public HTTP(S) source. Ambiguous duplicate identifiers are excluded.
Filters: year, sector, industry and company/report search. The initial view includes all years. Filters narrow the display; percentiles and ranks remain calculated within the full eligible sector across all years.
3. What is scored: the four gateways
Every report passes through four gateways in sequence. Each gateway is scored on the text the previous gateway passed through, so a weak result at an early gateway lowers the reliability of everything after it. The gateways are the same ones the platform applies to every client report and every prospect preliminary analysis, unchanged.
Gateway 1: Accessibility. The share of the document that can be extracted as clean text. Password-protected files, scrambled fonts, text rendered as images and broken encodings all lower it. Reported as a percentage. Below 90%, the later gateways are marked as computed on a partial document.
Gateway 2: Comprehensibility. How easily the extracted text is understood, by people and by machines, measured with the Readability Aggregate Index (RAI): a consensus of nine established readability formulas (Flesch Reading Ease, Flesch-Kincaid, Gunning FOG, SMOG, Coleman-Liau, ARI, Linsear Write, Dale-Chall, and a consensus grade), normalised to a 0 to 100 scale. Two figures are reported: - Raw: the formulas' own verdict on the text. - Real: Raw multiplied by prose quality, the share of the document that is structurally a sentence with a subject and a finite verb. Tables, headings and fragments pull Real down. A wide gap between Raw and Real is itself a finding: most of what is being counted as text is not prose.
Each figure is computed twice: on the full extracted text (readability to people) and on an extractive summary of it (readability to AI). The Index ranks on RAI Real, readability to AI, because that is the read the rating agencies' systems perform.
Gateway 3: Emphasis. The share of the document's extracted keywords that are sustainability keywords, matched against a taxonomy of environmental, social and governance terms organised by reporting framework (GRI, CSRD, ESRS, ISSB, TCFD, CDP, SASB). Reported as a percentage. For integrated and annual reports this is expected to be lower and is flagged as such.
Gateway 4: Alignment. How closely the themes the report emphasises match what the market weights in the company's industry. The benchmark is the market-consensus weighting of environmental, social and governance themes for that industry, built from rating-agency methodologies and framework requirements. Alignment is the deviation between the report's theme weights and the benchmark, where over-weighting counts against alignment as much as under-weighting. The overall alignment verdict is set by the single weakest theme, so one neglected material theme cannot hide behind strong coverage elsewhere.
The raw Alignment score is 1 / (1 + the largest absolute proportional deviation) and is shown as a percentage. Exact alignment is 100%; a theme omitted against a non-zero benchmark is 50%; progressively larger deviations score progressively lower. This keeps the weakest-theme rule meaningful without flooring every deviation of 100% or more to the same zero.
4. How a rank is computed
- Each of the four gateway scores is placed in the distribution of the company's own sector: the share of reports in that sector scoring at or below it. That is the company's percentile on that gateway. Percentiles are computed against the universe as defined, not against a stored table, so the same query that ranks a company can be re-run by anyone with the corpus.
- The Index score is the mean of the four percentiles, with one rule: a report whose accessibility is below 90% has its other three percentiles capped at the accessibility percentile. A document that cannot be read cannot rank above the reading of it.
- Companies are ranked on the Index score within their sector, highest first. Ties are broken by the alignment percentile, then by RAI Real to AI, then by company name, so the order is the same every time the query runs.
Small comparison groups. The table states the number of ranked reports in each sector. A sector with one company gives it rank 1 and 100th percentiles; that is not evidence that it outperformed another report. Small groups support limited conclusions. Search, sector and industry filters only narrow the display; they never recalculate the sector ranks on the current page or percentiles.
Why the sector and not the whole universe. A chemicals report and a bank's report are written for different readers under different rules. Ranking them against each other would measure the sector, not the company. The question a company asks is where it sits among its peers, so the peer set is the sector.
Why an average of percentiles and not a weighted composite. Weights are arguments. A percentile average says only that each gateway counts equally and that every company is judged against the same set. That is the claim we can defend.
Why percentiles and not raw scores. Raw scores have different scales and different distributions (accessibility clusters near 98%; RAI spreads across the range). Percentiles put them on one footing and answer the question a company asks: where do I sit among my peers.
5. What is published
The table shows every included report, 50 per page: company and report, sector, year, accessibility, readability to AI and people, emphasis, alignment, Index score and within-sector rank. Every column is sortable in either direction. Reports without enough assessment data show a reason instead of a rank.
The chart places emphasis (%) on the horizontal axis and (People RAI Real + AI RAI Real) / 2 on the vertical axis. Both axes are fixed from 0 to 100. Every visible point has the same radius, decreasing from 5 to 2 pixels as the visible sample grows. Size represents no company metric. Hover, click or use the arrow keys to inspect a report's measurements; repeated clicks cycle nearby overlapping points. Reports missing either chart measurement remain available in the table.
The page reports total reports, ranked reports, companies and report years. It loads fresh SQL results on each visit; no frozen edition or background refresh is required.
The full methodology, this document, is published with the ranking.
What is not published: any per-section diagnostic of a company's report, and any rewrite. Those are what the platform provides to a client. The public Index shows the position; the platform shows how to move it.
6. What the Index does not claim
- It does not rate sustainability performance, and no line of the published ranking says or implies that a higher-ranked company is more sustainable.
- It does not use any information a company has not published.
- It does not apply a pass mark. The standard is the distribution of the universe, and a company at the median is at the median.
- It is not a rating-agency score and does not predict one. It measures the property those scores depend on: whether the document can be read.
7. Corrections
We publish everything the corpus lets us read, and we fix what is wrong when someone tells us. A single indefensible ranking destroys the asset, so the correction route is public, short and answered.
- How to ask. The "tell us if this is wrong" form at the foot of the Index page. Name the company, say what is wrong, and add the report identifier if you have it. You get a reference straight away. Nothing else is required, and you do not need an account.
- What we check. The claim is run against the methodology on this page: the right company, the right document, the right year, and the four gateway scores of that document. We answer within five working days.
- How a correction appears. We correct the underlying report record. The next visit shows the updated ranking, including any changes to other companies’ relative positions. If you provide an email address, we can reply about your request.
8. Versioning and refresh
- The methodology carries a version number. Changes to the scoring or eligibility rules are recorded here.
- Version 3.1, 12 September 2026. The €1 billion market-capitalisation cutoff is removed. Rankings are calculated from current stored reports when the page is requested. Report additions, corrected data and changes in eligibility can change a company’s score or rank. There is no annual publication cycle or dated-edition history.
- Filters narrow the companies shown on the page; they do not recalculate its sector ranks.
- Annual snapshots are not currently provided.
9. Sources
- Friede, Busch and Bassen (2015), meta-analysis of roughly 2,200 studies on sustainability and financial performance.
- NYU Stern Center for Sustainable Business (2021), review of 1,141 papers, 2015 to 2020.
- MSCI (2024), AI in data collection and analysis across 17,000 issuers; 110 basis point cost-of-capital spread across 4,000 issuers.
- RepRisk (2025), daily screening of 150,000 public sources; adoption by 80% of the world's largest investment managers.
- Morningstar Sustainalytics (2024), machine learning and language models across 16,000 companies.
- BIS Innovation Hub, Deutsche Bundesbank and European Central Bank (2024), Project Gaia.
- PwC (2023), 94% of investors believe corporate sustainability reporting contains unsupported claims.
- Schimanski et al. (2024) and Bingler et al. (2024), the link between disclosure language and rating outcomes.
- The Eunoic corpus: more than 30,000 sustainability reports spanning 35 years, 50,000 companies, 150 industries.
(Every source above already appears with its citation in the preliminary-analysis page copy, static/md/prospect_prelim_analysis_copy.md. Full references are carried into the published page from there.)