Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • The Pymetrics game is rigged by design:

    Only 40% self report gender/race

    no resume data, no education information, degrees, schools, GPA, major, work experience, skills/certifications

    Zero job qualifications

  • Well, they're only looking at whether the pymetrics gameplay algorithm ML thing recommends the candidate, not any of that other stuff. The outcome they're looking at here isn't whether the people actually got hired, or got passed by other screening layers or anything.
  • Some job application websites I've seen actually have a yes or no option to consent to AI review that they claim is to simply assist HR and not actually screen you. I always select no. There is no way that selecting yes would ever be in my interest. I'm sorry, I'm going to force a real human to look at my stuff if I still can.
  • My fear is that pressing "no" on stuff like that is going to become an auto-rejection in the vast majority of cases
  • > We find that people who submit multiple applications to positions screened by the same algorithmic hiring vendor are more likely to be rejected from every position to which they apply than would be true if the companies made decisions statistically independently from one another. Ten percent of applicants who submit four applications are rejected from all the places to which they apply.

    > Our research also found that this pattern does not appear to be the case in other circumstances. We analyzed data from the largest prior study of hiring decisions, which sent 83,000 applications to 108 Fortune 500 firms during the same time period as our study and did not focus on whether AI was used to make decisions. We found that the rate at which applicants were rejected from every firm they applied to in this data was no higher than what you’d expect if each company decided independently of the others.

    It sounds like this study was using real-world applicants, and the other study they're comparing against was using synthetic applicants.

    Consider the chance of being accepted as being composed of signal+bias+noise. Noise is random. Signal is a per-applicant value, and what's meant to be measured. Bias is a per-group value, and an artifact of the measuring process.

    If acceptance/rejection is independent between positions applied for (as in the synthetic applicant study), that suggests that it's random or composed entirely of noise; ie there is no signal; ie the applicants are all equally qualified.

    If acceptance/rejection is correlated, that means there is some nonzero amount of (signal+bias). But real-world applicants are not all identical, so there should be some amount of signal. So you can't just assume zero signal in order to infer that there must be bias.

  • I think I am confused.

    A inferior candidate (by skill) is going to be consistently rejected, no?

  • Ayres, I., Banaji, M. and Jolls, C. (2015), Race effects on eBay. The RAND Journal of Economics, 46: 891-917. https://doi.org/10.1111/1756-2171.12115

    "Cards held by African-American sellers sold for approximately 20% ($0.90) less than cards held by Caucasian sellers, and the race effect was more pronounced in sales of minority player cards."

  • The paper is here: https://arxiv.org/pdf/2605.27371

    They find "disparate impact" of pymetrics across racial groups, but it doesn't seem like they controlled for anything.

  • They also say that if they do the analysis globally the effect goes away. Curious, does that not imply that if one domain is biased against some group there would be another where the bias was in its favor?
  • > To measure adverse impact, we apply the EEOC’s “four-fifths rule,” which flags a position when one group is recommended at less than 80% of the rate of the most-recommended group

    That seems like a nonsensical way to measure racial discrimination. What could justify it?

  • The desire to subsidize employment for Democratic constituencies by threatening legal action if they aren't given enough jobs.
  • It's a starting point to flag.

    Here's some analysis of what it is and why it's useful as a canary in the coal mine: https://www.prevuehr.com/resources/insights/adverse-impact-a...

  • ‘Every one is the same’, even when one group or another doesn’t like doing some kind of work for some reason.

    Because surely no one would have legitimate preferences based on their gender, cultural norms, etc. or real differences in aptitude due to childhood exposure, education, or said norms and preferences.

  • I guess it measures if there's more than one std deviation gap between highest and lowest? Assuming that's twenty percent here

    it sounds like how you'd get that kind of metric at least

  • >What could justify it?

    The assumption that applicants from all races are on average equally qualified for every position. Whole subfields of modern academia are based on that assumption.

  • This is an application of the disparate impact doctrine. Even facially neutral policies are considered suspect if they produce results that correlate against protected groups, irrespective of intent.

    This doctrine is the basis for much of employment law. It is a significant reason why employers don't administer IQ tests (or equivalents) to screen candidates since ~the 90s.

    A common objection to the doctrine is that it leads to unfalsifiable discrimination claims, which is why it seems nonsensical to you.

  • Have you googled this? The EEOC is a federal agency, and they've published on this topic quite extensively. The four fifths rule is used to define if there is a "substantially different selection rate". It does not measure racial discrimination. It measures selection rate.

    It indicates there may be adverse impact to one group. It specifically is not used to resolve racial discrimination.

    It's purely a signal for "we should consider asking more questions, because this appears unusual". That's what your quote says too, it "flags" a low recommendation -- it's indicating further study and investigation is likely warranted.

  • Anyone who’s done hiring wouldn’t be shocked by this:

    We find applicants are more likely to be rejected from every position they apply to than would be predicted by the baseline of each position making statistically independent decisions.

    Obviously a rejected resume is more likely to be rejected by every other employer and an accepted resume is more likely to be accepted by every other employer. Like online dating, most employers are looking for some baseline indicators that you are going to be successful and stable.

  • This does not sound right to me. It would be correct if all companies and their hiring managers had the same requirements/looking for candidates with the same qualifications - but they're not.

    Obviously as a hiring manager you're looking for a hard working individual with a number of successfully completed projects and glowing referrals from multiple places of employment, but you're also looking for a person with expertise in particular technologies/industries/whatever other areas of expertise. To a large extent requirements for each role are unique, however many do have some overlap.

    So being rejected from one position might simply mean there's a misalignment between what the company is looking for and what the individual has. Which might not be the case with other companies.

    So if we're seeing increasing number of candidates being consistently rejected at multiple places the question "why" is a valid one.

  • > a rejected resume is more likely to be rejected by every other employer

    This makes sense to me, albeit intuitively and in a way I can't articulate.

    > an accepted resume is more likely to be accepted by every other employer

    but this doesn't necessarily follow from the prior for me. Plenty of people get really good jobs and are really successful in them only after dozens or hundreds of rejections with a nearly-identical resume.

    by pc86
  • Yes I don’t understand why this is surprising or problematic at all?

    Actually the fact that they found this result didn’t hold in a different dataset is especially weird.

  • > Obviously a rejected resume is more likely to be rejected by every other employer and an accepted resume is more likely to be accepted by every other employer.

    But that wasn't the case for non-algorithmic screening. From the paper:

    "By contrast, we find that when first round screening is not mediated by a single screening procedure, systemic rejections are close to the baseline. To support the empirical validity of our baseline, we study homogeneous outcomes in the largest study of first-round screening at U.S. employers to date. Kline et al. [38] generated 83000 synthetic resumes and submitted these resumes to vacant positions at 108 US companies between October 2019 and April 2021, a similar time period to our data. The companies, which are a subset of the Fortune 500,15 collectively employ 15 million workers. We analyze the homogeneity observed in the resulting callback outcomes in their data. We find that the baseline is an effective estimator of the systemic rejection rate for this dataset. As shown in Figure 3, the observed systemic rejection rate is accurately predicted by the baseline and a chi-squared goodness-of-fit test cannot reject equality of the two distributions (2 = 20.05, = 0.69). In other words, while the largest previous study observes systemic rejection rates consistent with employers making statistically independent decisions, the algorithmic hiring data shows significantly correlated outcomes that lead to higher-than-baseline systemic rejection rates."

  • The European Union passed The Artificial Intelligence Act, which classifies:

    High-risk – AI applications that are expected to pose significant threats to health, safety, or the fundamental rights of persons. Notably, AI systems used in health, education, recruitment, critical infrastructure management, law enforcement or justice. They are subject to quality, transparency, human oversight and safety obligations

    That's a pretty common sense legislation to me.

  • > That's a pretty common sense legislation to me.

    There's no reason to single out AI vs any other approach to the same topics.

  • This is one of those things where the first sentence sounds completely fine and reasonable, maybe even objectively good.

    Of all the things listed "recruitment" doesn't belong to me. Is the argument that it is someone's fundamental human right to get someone else to pay them to do a job? Or is it strictly about human oversight?

    by pc86
  • The AI “safety” industry is lobbying for federal preemption so that states won’t have the power to enact these types of sensible regulations.
  • Misleading title the paper [0] does not mention any CV screening that might suggest racial or gender bias. It is purely about assessment tool. No AI or LLMs.

    I'm not saying AI is not biased, but this study does not prove that.

    [0] https://arxiv.org/pdf/2605.27371

    From the paper:

    > Fig. 1. The pymetrics process. > Stage 1: Applicants apply to positions. > Stage 2: Applicants are directed to the pymetrics platform to play assessment games. > Stage 3: pymetrics algorithms use applicant gameplay features to recommend 58.2% of applicants per position on average. > Stage 4: Employers decide which applicants to interview or hire, typically rejecting applicants that were not recommended by pymetrics.

    by Oras
  • Did I miss the part of the article where they break down how they determined race? Is the algorithm blind to race? It looks like they specifically looked at 83k people applying to ~100 companies which notably were Fortune 500 companies. Could there simply be candidate discrepancies here? Hard for me to follow the full methodology but it doesn't necessarily seem either malicious or that well structured. Don't you need to have a control group of applicants who are similar on paper? To allege DISCRIMINATION is quite bold.

    Definitely open to opposing or critical views

  • id expect any algorithm to learn race by other properties in the data?

    its going to be in the rest of the data because race has a meaningful correlation, and pleanty of causation with being disadvantaged in real ways, that can also affect the ability to then do certain jobs.

    like, the environmental pollution and building interstates and freeways through black communities, on purpose to do bad things to those communities, then results in a bunch of noise and particulate pollution, that is bad for developing brains.

    you wont be able to do some meritocratic non-racist hiring without fixing the environmental racism. otherwise youre just mirroring racism other people built for you

  • Yes. You missed it. They are using a test dataset of 83k resumes generated in 2022 for this paper and comparing it as a baseline against their observational data: https://www.nber.org/papers/w29053

    The dataset is constructed, deliberately, to hold candidate performance constant and vary the names of candidates to appear to be associated with a specific race.