AI Audits Need a Power Test, Not Just a Fairness Score
Massimo Ragnedda, Maria Laura Ruiu / Aug 20, 2026
Alexa Steinbrück / Better Images of AI / Explainable AI / CC-BY 4.0
In 2019, researchers examined a commercial algorithm used by US health systems to identify patients who should receive additional care. The system predicted future health care costs rather than medical need. Because less money was historically spent on Black patients with comparable illnesses, Black patients assigned the same risk score were considerably sicker than white patients.
Replacing cost with a closer measure of need would have increased the proportion of Black patients selected for additional support from 17.7 percent to 46.5 percent, according to the study published in Science.
The model was not simply inaccurate. It was competently predicting the wrong thing.
That distinction matters as audits become a central instrument of AI regulation. New York City requires bias audits for certain automated employment decision tools. The European Union’s AI Act provides for conformity assessments and fundamental-rights impact assessments for high-risk systems. The US National Institute of Standards and Technology’s (NIST) framework encourages organizations to govern, map, measure, and manage AI risks.
These mechanisms are necessary. But an audit that begins after an institution has chosen the system’s objective may certify an unequal policy with impressive technical documentation. AI accountability therefore needs a power test alongside the fairness test.
A fair result can still serve the wrong objective
Most technical audits ask whether a system performs consistently, whether error rates differ across groups, and whether documented controls are working. These are important questions. They are not the same as asking whether the system’s target is legitimate.
The health care case shows the difference. Predicting expenditure may appear administratively sensible because cost data are available and measurable. Yet expenditure reflected unequal access to care. Once cost was treated as a proxy for need, a social inequality became part of the model’s objective.
Research on fairness in sociotechnical systems has long warned against isolating a model from the institution in which it operates. A statistically balanced system can still support a punitive welfare policy, an exclusionary hiring process, or a surveillance practice that should not have been automated in the first place.
This is where bias laundering can occur. The term does not imply that auditors intentionally conceal discrimination. It describes a narrower institutional risk: political choices about what should be predicted are translated into technical targets, assessed through compliance metrics, and returned to the public with the authority of an audit.
Existing frameworks contain the right pieces, but not a complete test
Some governance frameworks already move beyond model performance. The NIST AI Risk Management Framework asks organizations to document intended purposes, consult relevant external actors, and make a go-or-no-go decision about whether an AI system is appropriate.
UNESCO’s Ethical Impact Assessment similarly asks whether AI adoption is justified, identifies relevant stakeholders, and examines positive and negative consequences.
The EU AI Act’s fundamental-rights impact assessment requires covered deployers to describe the context of use, identify affected groups, assess risks, and establish mitigation and redress arrangements. However, participation by affected groups is encouraged “where appropriate” rather than made a uniform requirement.
Under the current political agreement on the EU’s implementation timetable, the relevant high-risk requirements have also been postponed while standards and guidance are completed. This creates additional time to improve the substance of the assessment process rather than treating the delay only as a concession to compliance concerns.
The building blocks therefore exist, but they remain fragmented. Some are voluntary. Some apply only to particular organizations or categories of systems. Others require documentation without clearly establishing who has authority to challenge the purpose of the system or the contractual arrangements behind it.
The emerging audit regime needs a shared power test.
Four questions every consequential AI audit should answer
1. Who defined the problem?
An audit should identify who selected the objective, which alternatives were considered, why the AI system was chosen as the solution, and what would happen without automation. It should name the proxies used and explain why they represent the underlying social goal. This requirement would have made the difference between health care cost and medical need impossible to treat as a minor modeling choice.
2. Who controls the system?
Audit documentation should disclose who owns or controls the relevant data, model, computing infrastructure, and intellectual property. For public-sector systems, procurement contracts should guarantee access to documentation, change logs, incident reports, and independent testing.
Contracts should also specify whether an institution can switch providers or decommission a system without losing access to essential data or services. Vendor ownership does not make a system inherently harmful, but it determines where and how accountability can be exercised.
3. Who receives the benefits and bears the errors?
Group-level error rates remain important, but audits should also examine how benefits, burdens, and savings are distributed. A system may reduce administrative costs while transferring investigation, delay, documentation, or appeal costs to applicants, workers, patients, or welfare recipients. These effects should be treated as part of system performance rather than as externalities.
4. Who can contest the decision?
Notice and explanation are insufficient without practical routes to correction. Audits should verify that affected people can obtain relevant information, challenge data and classifications, reach a responsible human decision-maker, and receive timely remedies.
Affected communities should also have standing to influence system objectives before deployment, not merely report harms afterward.
Make the power test part of existing governance
A power test does not require a new regulatory agency or an additional audit industry. It can be incorporated into existing impact assessments, procurement reviews, conformity procedures, and public transparency records.
The United Kingdom’s Algorithmic Transparency Recording Standard already asks public bodies to publish information about system ownership, rationale, deployment context, data, risks, and accountability.
The EU’s fundamental-rights assessment template could make affected-group participation mandatory for consequential public uses and add questions about vendor dependence, benefit distribution, and the no-AI alternative. Regulators could require a confidential annex where genuine security or trade-secret concerns prevent full publication, while still publishing a meaningful public summary.
The requirements should be proportionate. A power test is most justified for systems that materially shape employment, education, health, welfare, credit, migration, policing, or access to public services. Standardized templates could reduce compliance costs, particularly for smaller organizations.
Auditor independence also matters. Providers should not be able to define the benchmark, select the evidence, and determine what counts as a successful result without external scrutiny. Previous Tech Policy Press analysis has identified unresolved questions concerning audit objects, credentialing, and declining audit quality.
Regulators can establish minimum methodological standards, approve qualified auditors, require disclosure of financial relationships, and trigger reassessment after material system changes or serious incidents.
Auditing power is not the same as politicizing technical review
A predictable objection is that auditors should assess systems, not decide public policy. That is correct. Auditors should not replace legislators, regulators, courts, workers, or affected communities.
But the objective of an AI system is already a policy choice. Treating it as a fixed technical specification does not make the choice neutral. It merely shields it from scrutiny.
The purpose of the power test is to make these decisions visible and assign responsibility for them.
Technical auditing remains indispensable. Regulators need evidence about accuracy, robustness, discrimination, security, and compliance. Yet these measures answer whether a system works according to its specification. Democratic accountability also requires asking who wrote the specification, whose interests it serves, and who can refuse its consequences.
An AI system can be accurate, documented, and compliant while still allocating care, work, welfare, or scrutiny on an unjust premise. As audit regimes harden into governance infrastructure, regulators should ensure that the audit object does not stop at the model. The institution, objective, contractual dependencies, and distribution of power must also be examined.
A fairness score asks whether a system treats groups consistently. A power test asks who has the authority to define consistency, and who has the power to challenge it.
Without that second test, compliance can certify inequality. With it, an audit can become a democratic checkpoint rather than a technical seal of approval.
Authors


