Two regulators now ask questions about interview scoring, and they are not the same question. New York City asks whether your tool's outcomes are skewed across protected categories, and whether you published the answer. The European Union asks whether the system was built, documented, and supervised well enough to be trusted in a high-risk setting. A compliance program that answers only one of them is half a program.
What NYC Local Law 144 actually requires
Local Law 144 covers automated employment decision tools used to substantially assist hiring or promotion decisions for roles in New York City. It took effect January 1, 2023, with enforcement beginning July 5, 2023. It requires an independent bias audit conducted within the year preceding use, a public summary of the results posted on the employer's site, and notice to candidates at least 10 business days before the tool is used.
The audit math is prescribed, which is useful. For a scoring tool you compute the rate at which candidates in each category score above the sample median, then divide each category's rate by the rate of the highest-scoring category. That quotient is the impact ratio. It must be reported by sex category, by race and ethnicity category, and by the intersections of the two, along with the number of people in each cell.
What the EU AI Act asks instead
The EU AI Act, in force since August 1, 2024, classifies AI systems used for recruitment, application filtering, and candidate evaluation as high risk under Annex III. High risk is not a prohibition. It is a documentation and governance regime, and the obligations sit on both the provider that builds the system and the deployer that puts it in front of candidates.
- A risk management system maintained across the lifecycle, not a launch-day document
- Data governance covering training and evaluation sets, including examination for bias
- Technical documentation and automatically generated logs, retained by deployers for at least six months
- Human oversight assigned to named people with the authority and the competence to override an output
- Accuracy, robustness, and cybersecurity commitments that can actually be tested
The scheduled application date for the Annex III high-risk obligations is August 2, 2026, and the Commission has proposed adjustments to that timeline. Treat the exact date as a question for counsel. Treat the engineering as work to start now, because none of it retrofits quickly.
An impact ratio tells you the outcome was uneven. It does not tell you why, and the why is where the fix lives.
Audit the score, not just the decision
Most audit programs measure the end of the funnel, where the model score has already been blended with recruiter judgment, scheduling luck, and offer negotiation. That is the legally interesting number and it is close to useless for debugging. Score the model in isolation as well, on a held-out sample, with the human step removed, so you can tell a model problem from a process problem.
The EEOC's four-fifths guidance, an 80 percent threshold on selection rate ratios, is a screening heuristic rather than a safe harbor. Ratios above 0.8 can still reflect real disparity at scale, and ratios below 0.8 can be noise in a cell of 40 people. Report the ratio, the cell size, and a confidence interval together, or the number will be misread by whoever reads it next.
What we audit on Nova Recruiter
- Impact ratios by category and by intersection, recomputed quarterly rather than annually
- Per-question score distributions, since bias usually enters through one badly written question rather than the whole rubric
- Transcription accuracy by accent and by first language, because a scoring model cannot outperform its input
- Drift between the calibration sample and the live population, checked whenever a role family shifts
- Override rates, meaning how often a human recruiter reverses the model and in which direction
The per-question view is the one that pays for itself. In a 2025 review of a logistics client's supervisor req, a single situational question about weekend availability drove almost the entire subgroup gap in the composite score. The question was job-relevant and lawful. It was also doing work that belonged in a scheduling conversation, not in a competency score.
What an audit cannot tell you
A bias audit measures outcomes on the population that reached the tool. It says nothing about who never applied, which sourcing channel built the pool, or whether the job description filtered people out before the model saw anyone. A clean impact ratio on a pool that Beacon Sourcing assembled from three job boards is a clean ratio on those three job boards.
It also cannot certify the rubric. If the underlying competency model rewards a communication style that correlates with background rather than with performance, the audit will cheerfully report evenly distributed scores on a bad construct. Validation and bias audit are separate exercises, and validation is the harder of the two.
Publish the audit, keep the logs, name the humans, and re-run the numbers more often than the law demands. The regulatory floor is a floor. Candidates and clients both notice the difference between a company that measures quarterly and a company that measures once a year because it has to.
Naomi FeldsteinHead of Responsible AI at Novexhire