Research

Bias in AI Hiring: What the Evidence Shows

Last reviewed

Summary

Bias in hiring predates AI, and AI can reduce some sources of inconsistency while introducing others. Research has documented disparities in technologies used in assessment, such as speech recognition and facial analysis, so no AI hiring system should be described as bias-free.

What the evidence says

  • Manual resume screening has been shown in a field experiment to favor resumes with names associated with white applicants over otherwise comparable resumes with names associated with Black applicants. [1]
  • Automated speech recognition systems from several major providers were found to make more errors for Black speakers than for white speakers. [2]
  • Commercial facial-analysis systems have shown higher error rates for some demographic groups, and NIST found demographic differentials in many face recognition algorithms. [3][4]
  • US federal guidelines use the "four-fifths rule" as a practical rule of thumb for identifying adverse impact in selection procedures. [5]
  • Many vendors of algorithmic hiring tools disclose little about how they test for and mitigate bias. [6]

AuraSync's interpretation

Because the components of multimodal assessment, such as transcription and face verification, have documented disparities, we do not describe AuraSync as bias-free. We aim for consistency and reviewability: one structured method for every candidate, evidence attached to every result, adjustments on request and a person making the decision.

What AuraSync claims

  • AuraSync does not describe itself as bias-free.
  • Results can be affected by audio and video quality, accent, connection, device, disability or assistive technology, and candidates can request adjustments and human review.
  • Customers remain responsible for monitoring their own hiring outcomes.

Limitations

The studies above measured specific systems at specific times; error patterns change as technology changes. None evaluated AuraSync. Consistent methods reduce some kinds of bias but do not guarantee fair outcomes.

Sources

  1. [1]Bertrand, M., & Mullainathan, S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4), 991–1013. https://doi.org/10.1257/0002828042002561
  2. [2]Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., & Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689. https://doi.org/10.1073/pnas.1915768117
  3. [3]Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 77–91. https://proceedings.mlr.press/v81/buolamwini18a.html
  4. [4]Grother, P., Ngan, M., & Hanaoka, K. (2019). Face Recognition Vendor Test (FRVT) Part 3: Demographic Effects (NISTIR 8280). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.IR.8280
  5. [5]Uniform Guidelines on Employee Selection Procedures (1978), 29 C.F.R. Part 1607. https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XIV/part-1607
  6. [6]Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020). Mitigating bias in algorithmic hiring: Evaluating claims and practices. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 469–481. https://doi.org/10.1145/3351095.3372828

This page summarizes third-party research for orientation. It is not an evaluation of AuraSync, and any summary of law is not legal advice.