CSOAI · GSPC suite · 12 benchmarks

GSPC-GOV

the governance axis   measured

Classify an AI deployment into its EU AI Act risk tier.

Statutory anchor

EU AI Act (Reg. 2024/1689) Art 5, Annex III, Art 50

How it is graded

Deterministically. A regex extracts the label, scored by macro-F1 — no model judges another model. A response with no readable label is reported as UNMEASURED and is excluded from the denominator; it is never scored as a wrong answer. That distinction separates "the model was wrong" from "the model never answered", and it is enforced in the harness code, not just claimed in prose.

What may be quoted

This axis is measured. n = 24 frozen items — which is below usable_n = 30, so no confidence interval is published on it, including by us. Report the n with any figure you quote.

At a glance

axisgovernance
stateMEASURED
items24
gradingdeterministic
licenceApache-2.0

The dataset Run the harness CSOAI

Run GovBench yourself

The same 24 items sov34 answered, graded by the same deterministic rule. sov34 scored 0.386 macro-F1 on these. No sign-up, nothing leaves your browser.

Items: csoai/gspc-gov · grading is a regex label read plus macro-F1, identical to the published harness · measurement, not certification, and not legal advice.