Practised by
2 links
Graph · Strategy
01 · In focus
The structured facts the source records about Community-defined AI benchmarks and evaluation standards, the count of declared adjacencies in the corpus, and the federation map zoomed on this node and its neighbours.
strategy
↑2 declared connections
02 · Connections
Split by direction. Direct links are the ones Community-defined AI benchmarks and evaluation standards’s source record names; inferred backlinks are records elsewhere in the corpus that point at this entity.
2 links
Other records that name this entity.
2 links
03 · Background
Body prose as it appears in movement-graph’s published markdown for this entity. Links to other corpus entities resolve to their graph page; links to deeper repo paths are kept as text so the page does not invent a route.
Build the evaluation infrastructure itself — benchmarks, datasets, disclosure templates, model documentation formats — from outside the vendor and outside the state, so the tests a system is measured against reflect movement-defined criteria (representational harm, dialect fairness, disability access, Global-South-language performance, culturally-situated notions of accuracy, refusal-rates on high-risk requests) that the vendor's internal evaluation and the incumbent industry benchmark leaderboards do not. The vehicle is community-built artefacts: Data Nutrition Labels, Model Cards for Model Reporting, Datasheets for Datasets, participatory benchmark construction (BOLD, WinoGender, Stereoset for bias; MasakhaneNER for African-language NER; language-specific evaluation sets for under-served languages), harm-taxonomy templates civil society releases as reference standards.
An actor chooses this strategy because the choice of benchmark silently governs the market's account of what a good AI system is — a leaderboard that measures accuracy on English news text will select for models that are good at English news text, and every non-benchmarked failure mode is externalised — and the movement's most durable structural intervention is the one that shifts the benchmark itself. Community-defined evaluation is also a lower-cost strategy than strat-civil-society-inside-technical-standards-bodies: the artefact is released rather than negotiated, and its adoption is by uptake in downstream research and vendor practice rather than by consensus in a standards body. Where the community standard becomes the field's reference (as Model Cards did after 2018, Datasheets did after 2019), the movement's evaluation frame is baked into the vendor's own compliance apparatus with no coercive instrument required.
It trades enforcement for reach. The community standard is not binding on any actor — the vendor can ignore it, the state can decline to incorporate it into regulation, and the industry can build competing corporate-favoured leaderboards that route around it — and adoption depends on the community artefact being demonstrably better than the alternatives for the questions the field is arguing about, which itself takes years of iteration and citation-building. The strategy is also structurally dependent on the community's own capacity to sustain the artefact past its authors' careers: many community benchmarks decay or fail to update as the model class shifts, and the movement's evaluation infrastructure remains chronically under-provided at maintenance.
This strategy differs from strat-civil-society-inside-technical-standards-bodies by venue — inside-standards-bodies works through IEEE, ISO, NIST, and similar formal institutions with member-body voting, community-defined works through publication and uptake — and from strat-empirical-audit-and-expose by object (the evaluation instrument vs. a specific system's result under some evaluation). It is the strategy that shapes what every subsequent measurement means, and its influence is diffuse but structural.
Source: entities/strategies/strat-community-defined-benchmarks-and-standards.md — movement-graph pin 5d136ad.