How to Evaluate an AI Hiring Tool for Bias: A Checklist for Employers
Every vendor says their tool is fair. That promise is worth nothing. Judge the tool by what it refuses to do and what it can show you: the questions to ask, what a good answer sounds like, and the red flags that should stop you.
By Umair Ali · ·
You are about to let an AI hiring tool have a say in who gets hired, and the sales pitch will lean hard on words like objective and unbiased. The vendors getting sued used the same words. "Unbiased" is a marketing term, not a fact you can check.
So judge it on something you can verify instead: the practices a vendor will rule out in writing, and the reasoning they can put on screen for any candidate you pick. Fairness claims are free. Refusals and evidence are not. This checklist is the set of questions to put to any AI hiring vendor before you buy, with what a good answer sounds like and the red flag that should stop you. Run it before signing anything. None of this is legal advice, so confirm the rules where you operate.
This is the buying half of a bigger question. For the evidence on whether AI hiring is actually biased, and for the rules of using AI in hiring without adding bias once a tool is in place, see the companion guides.
How to use this
Ask every question below of the vendor directly, in writing if you can. You are listening for two things: clear answers, and willingness to show proof. A vendor who gets vague, or who answers a "can you show me" with a "trust us," has told you something. On the questions marked walk-away, a wrong answer is reason enough to stop, on its own.
1. What does the tool score?
Ask: Does it analyze a candidate's face, expressions, voice, tone, or accent in any way?
Good answer: It evaluates the content of what a candidate says or does, in text, and ignores how they look or sound.
Walk-away: It scores video appearance, facial expression, or speech patterns. This is the most reliably discriminatory thing in the category. It penalizes non-native speakers, disabled candidates, and neurodivergent ones for traits with no bearing on the job, and it has drawn the heaviest criticism and legal attention of any AI hiring feature. There is no version of appearance scoring that measures competence. If the answer is yes, you are done here.
2. Who makes the decision?
Ask: Does the tool reject candidates automatically, or does a human make the final call?
Good answer: It supports a human decision. It surfaces, ranks, or recommends, and a person accepts or rejects. Nobody is cut without a human in the loop.
Walk-away: It auto-rejects behind the scenes. Automated rejection with no human review is where black-box bias does its quiet damage, filtering the same people out at scale with nobody accountable. Courts are moving toward holding employers responsible for exactly this, so "the tool decided" will not protect you.
3. Can it explain itself?
Ask: For any given candidate, can it show me why they ranked where they did, in terms I can read?
Good answer: Yes, every ranking comes with a specific, readable reason tied to the candidate's own information, for example a claim they made and how it was tested.
Walk-away: It returns a score or a "recommend / do not recommend" with no explanation. A black box you cannot interrogate is a black box you cannot defend, to a rejected candidate, to a regulator, or to yourself. If the vendor cannot show the reasoning, assume there is a reason they cannot.
4. What was it trained on?
Ask: Is the model trained on our past hires, or on historical hiring outcomes?
Good answer: It evaluates each candidate against the requirements of the role and what they can demonstrate, rather than learning a pattern from who was hired before.
Red flag: It learns from your past hires or past "good hire" labels. This is the single most common source of AI hiring bias, because it teaches the tool to reproduce whoever you happened to hire before, discrimination included. Amazon scrapped a tool for exactly this. Training on history means inheriting history.
5. Has it been audited for bias?
Ask: Has an independent bias audit been run on this tool, and can I see the results?
Good answer: Yes, and here are the results. In New York City this is legally required for covered tools, with public reporting, so a vendor operating there should have it ready.
Red flag: They cannot produce one, or will not show it. Even where an audit is not legally required, a vendor confident in their tool's fairness should be able to point to some testing of outcomes across groups. Vagueness here is a signal. Note that a young or small tool may honestly not have a formal third-party audit yet, so weigh a candid "not yet, but here is how we test" very differently from an evasive non-answer. The refusal to engage is the red flag, more than the absence of a certificate.
6. Does it accommodate disability?
Ask: How does the tool handle candidates who need accommodations, and does anything in it penalize non-standard responses?
Good answer: It judges the substance of what a candidate says, offers or allows accommodations, and does not score speed, timing, or manner in ways that disadvantage disabled or neurodivergent people.
Red flag: Timed tests, game-based assessments, or scoring that rewards a particular response style. These are a documented way disabled candidates get filtered out, and they carry direct legal exposure.
7. Who is accountable, and does it fit the law?
Ask: Where does liability sit, how is candidate data and consent handled, and does the tool meet the rules where we hire?
Good answer: The vendor can speak clearly to data protection and consent, and to the relevant regimes, New York City's audit law, California's and Colorado's rules, the EU treating hiring as high-risk, and accepts that you as the employer remain responsible for outcomes.
Red flag: They imply the tool takes the legal risk off your hands. It does not. The Workday case points toward employer and vendor sharing liability, so a vendor promising to absorb yours is either naive or misleading.
The red flags, in one place
If you remember nothing else, these are the walk-away signals:
- It scores face, voice, or accent.
- It rejects candidates automatically, with no human deciding.
- It cannot explain, in readable terms, why a candidate ranked where they did.
- It learns from your past hires and calls that objectivity.
- The vendor will not discuss auditing or outcomes, and gets vague when you press.
- The vendor claims the tool is unbiased, or promises to carry your legal risk.
Any one of those is a reason to keep looking.
Running our own tool through this
It would be strange to hand you this checklist and dodge it, so here is where BestHire lands on it, honestly.
On scoring, it reads only what a candidate says in the interview, never face, voice, tone, or accent. On decisions, it never auto-rejects; it hands the founder a ranked shortlist and a person makes every call. On explainability, every ranking comes with the evidence attached, the exact resume line and the interview answer behind it, so there is no black-box score. On training, it tests what each candidate can back up about their own work rather than matching them to past hires. On data, it runs on consent under GDPR.
On the audit question, the honest answer is the one this checklist tells you to respect: judge us by whether we engage, not by a certificate we wave. Ask us how we test for biased outcomes and we will talk specifics rather than hide behind the word "unbiased", which we do not use, because no tool has earned it. That candor is the point of the whole list. A vendor's honesty about its limits tells you more than its confidence about its strengths.
The short version
- Ignore the vendor's fairness marketing. Judge the tool on what it refuses to do and what it can show you.
- Walk away if it scores face, voice, or accent, or if it auto-rejects with no human deciding.
- Demand a readable reason for every ranking. Reject black-box scores.
- Ask what it was trained on. Training on past hires bakes in past bias.
- Ask to see a bias audit, and weigh candour about limits over confident claims of none.
- You stay legally responsible for what the tool does. No vendor takes that off you.