EnviroBench · Pre-release

The benchmark for AI in environmental work.

EnviroBench is in development: a benchmark that scores AI systems on the documents environmental projects run on, case by case against a fixed answer key. We are finalizing the test sets now; results for leading AI models and Statvis come at release.

In development · Version 0.2.0 · Maintained by Statvis

Coming at release

What is being evaluated

Leading AI models and Statvis, run on the same test sets and scored in six categories. Scores are published once the test sets are final.

Systems under evaluation

AI models, one call per case through the open adapter

  • Claude Sonnet 5.5Anthropic
  • Claude Haiku 4.5Anthropic
  • Qwen3.5 397B-A17BQwen (Alibaba)
  • Qwen3-VL 8B InstructQwen (Alibaba)
  • Gemma 4 31BGoogle
  • Mistral Small 4Mistral AI

Multi-step set-ups

  • Claude Haiku 4.5 · reference agent v2Anthropic · tool-using agent
  • Claude Haiku 4.5 · page map-reduceAnthropic · one call per page, then a merge

Products

  • StatvisDocument import pipeline, run with no human review

Scored in 6 categories, 0 to 100

  • ExtractHow accurately does the system turn documents into data? 6 tasks.
  • VerifyDoes the system know when it might be wrong? 1 task.
  • LocateCan the system place things on site plans? 1 task.
  • ResearchDoes the system find the right evidence, and see where the record conflicts or falls short? 3 tasks.
  • LegalDoes the system read legal instruments without overstating them? 4 tasks.
  • SiteCan the system work from a site's whole document record? 2 tasks.

How the categories are computed

  • 17 tasksfrom a single cell to a site's whole record, each on its own measures
  • Answers withheldnever shown to the system under test
  • Case by caseevery result scored and inspectable on its own
  • Private real-world setsnever published, so never trained on

How it works

From document to dependable data

Each task is one step from document to database, so you can see where a system can be relied on.

Site dataset · Site questions

A site's whole record

Hundreds of pages, with reprinted results, drafts beside finals and corrected values. Report each measurement once, answer questions, and say when the record has no answer.

Measured: records right and cited, duplicates, questions right, invented answers, cost and time per site. Details

Document extraction

Whole reports to records

A whole PDF in, every measurement out, with the right sample, date, depth, unit and non-detect.

Measured: records with the right value, and wrong values reported with no flag. Details

Table mapping

Analytical summary tables

Shading, non-detects like <0.1 or 0.81 U, qualifiers and units in headers or footnotes: every value must reach the right sample, depth, date and matrix.

Measured: the role of every cell, and the records that follow. Details

1 2 3

Expected reading: one record per result (20 in all)

CalloutSampleDepth (m)DateParameterResultUnit
1MW-11.02026-03-03Lead212mg/kg
2BH-12.02026-03-02Arsenic<2.0mg/kg
3MW-11.02026-03-03Toluene0.08 Jmg/kg
MW-11.02026-03-03Arsenic18.6mg/kg
MW-11.02026-03-03Benzene0.21mg/kg
MW-21.52026-03-03PHC F222mg/kg

… and 14 more records, every one with its sample, depth, date and matrix (soil).

  1. 1A result above its guideline, shaded in the report. The reading must still attach it to MW-1 at 1.0 m.
  2. 2A non-detect, reported as less than the reporting limit. It stays <2.0, not 2.0.
  3. 3A qualified value: the J (estimated) stays with 0.08.
  • Sample: location, depth, date, matrix
  • Parameter
  • Unit
  • Result
  • No role (the guideline column is ignored)
Figure 1. Synthetic case mapping-04: outlines show each cell's role in the answer key, and the records on the right follow from them.

Table cell OCR

Reading each cell exactly

One character changes a result: 1.2 or 12, a dropped U.

Measured: exact readings, and how often a wrong one would skip review. Details

Table containers · Table separators

Scanned historical reports

Typewritten, crooked scans whose tables often have no ruled lines. Each table must be found and split into rows and columns.

Measured: table boxes within 1% of the page, pages straightened within 0.3°, row and column lines. Details

1 2 3
4
  • Answer key
  • System output
  1. 1The only box the system kept is too loose: its edges are more than 1% of the page from the table.
  2. 2The borderless table has no box at all.
  3. 3The page is turned 1.4°; the system left it as scanned.
  4. 4On the cut-out table, a weaker rule found all four column gaps but none of the five row lines.
Figure 2. Synthetic cases containers-11 and separators-13. Blue is the answer key; orange is the zero-cost reference adapter.

OCR auto-accept

Deciding what a person must check

Which machine readings are safe to accept without a person checking them?

Measured: accepted wrong values count as errors; caution is reported as cells sent to review. Details

1 2 3 4 5
Callout What the readers saw Checked value Rule A Rule B
1All read 0.81 U; OCR confidence 0.970.81 UAccepted, correctSent to review
212, but one vision reader saw 1.212Sent to reviewSent to review
3All read 8 under a smudge; OCR confidence 0.61Not legibleAccepted an unreadable cellSent to review
4All read a dash; unexplained inkBlankAccepted a mark as a valueSent to review
5All read 1,200; the cell crosses the page edge11,200Accepted a wrong valueSent to review
Figure 3. Five synthetic cells and two accept rules from the reference adapter. Rule B also requires a confident, character-exact OCR match. In cell 5 every reader agreed and every reader was wrong.

Sample location placement

Sample locations on site plans

Wells and boreholes are small labelled symbols; a misplaced one moves its results across the site.

Measured: distance from the checked location as a share of the plan's diagonal. Details

  • Checked location
  • System point
  • Marker not found
Distance and score per location
LocationDistanceScore
MW-10.65%Within 1%
MW-22.71%Near (1–5%)
MW-36.73%Far (over 5%)
MW-40.35%Within 1%
BH-11.73%Near (1–5%)
BH-2—Not found
Figure 4. Synthetic cases location-003 to location-008, a fictional site plan. Orange crosses are the reference adapter's weak baseline.

How scoring works

Scores built to be trusted

  1. 1Inputs onlyThe system gets the documents and the question.
  2. 2System outputIt returns its reading and its cost.
  3. 3Answer key joinedAnswers are added only after it finishes.
  4. 4Every case scoredEach table, page, cell or query gets its own errors.
  5. 5Paired comparisonA new error on any stable case fails the change.
New errors per synthetic table crop, by score profile
Profile Line F1, before → after New errors per crop, crops 1–13 Result
Lines within 5 units 0.26 → 0.57 Fails: 11 crops regressed
Lines split the words the same way 0.58 → 1.00 Passes: no crop regressed
  • Gained an error the previous run did not have
  • No new error
Figure 6. A paired comparison on the 13 synthetic table crops (see the quick start). The new rule doubles the line score, yet 11 crops gained an error, so it fails.

Test sets

Public to reproduce, private to stay honest

Public

Synthetic test sets

Generated documents for every task, published with their answers. Every figure here comes from them.

Private

Real-world test sets

Real project documents, scored by the same rules and never published, so never trained on.

A public development set and a sealed holdout are planned. See the methodology.

Follow the release

EnviroBench is under development and we are finalizing the test sets. Results will be published here at release.