AI Reality Case Studies

Almost Magic Tech Lab seal: Clarity over cynicism. Less noise, more nerve.

The shortlist nobody could explain

An AI Reality case study - People and Hiring

Case ARC-HR-01 | Fictional composite | Free reader edition, 1 October 2026

This is a fictional composite. North Quay Utilities, its people, figures, internal extracts and results are invented for discussion; none are AMTL client findings. The public incidents and regulator guidance in the source note are separate and explicitly labelled.

Tuesday, 9:10 am: the backlog

North Quay maintained power equipment across a broad regional service area. Twelve field technician jobs had stayed open after an intense summer. The vacancies did not mean twelve identical gaps: four were for crews qualified to work independently, and eight were for teams where supervision was available. The operations director, Helen, had already moved two maintenance jobs back a fortnight. Each shift she covered with the contractor deferred a choice about which maintenance task could wait. A union delegate had asked whether the overtime roster was still voluntary; Helen could not answer for the next month. Staff were covering overtime, and managers had begun to argue over whether a slower repair was a hiring problem or a scheduling problem.

The recruitment mailbox held 486 applications. Laila, the only recruiter assigned to this campaign, had spent two weeks sorting them among other work. A ranking vendor offered a pilot: extract skills and experience from the files, compare them with North Quay's last three years of successful hires, and show a scored list with one-line rationales. Priya, the people director, permitted queue ordering on the understanding that Laila would choose whom to interview. No one had put an automatic rejection rule in the pilot approval.

On Friday, the vendor returned a sorted list. Laila read the top 60 applications against the job description. She booked none yet. She also opened six applications below the cut line when their licences caught her eye in a separate search. Two came from applicants returning after caring breaks; one had moved among contractors whose names the parser did not recognise. Their rationales said “limited recent continuity”. The attached work samples were not reflected in the short explanations. These six were leads, not a random audit and not proof that six qualified people had been wrongly excluded.

Helen wanted the interview invites out by Friday. A local contractor had offered to keep covering shifts for another month, but the price would rise. Laila could take on the extra screening only by dropping another campaign. The vendor could change the continuity weight inside its system immediately, yet could not supply Priya with a reproducible score calculation or an export of every factor used. The training set was drawn from North Quay’s own past hires: that made the tool feel tailored, but also made Priya ask whose career paths had been represented in those past hires. She had no analysis that could answer it. If Priya simply accepted the top 60, more than eight out of ten applicants would have no human review before the interview list closed.

Tuesday, 10:40 am: the letter

An applicant, Samira, wrote: “I hold the required licence and did this work for eight years. Was my application read?” Her score placed her at 214. The licence was in the PDF, and a separate attachment described a recent year of contract work. The vendor rationale called it “limited recent continuity.” Laila had not opened Samira's file before the email. A drafted reply said every application had received fair consideration. Priya stopped it: the words made a claim about the 426 people below the reviewed group that the team could not evidence.

Counsel asked what North Quay had told applicants about the ranking, which files had gone to the vendor, whether it retained them for training, and who could erase them. Procurement had a sales deck and a pilot order, but not an agreed retention schedule or a usable decision log. These were open questions, not proof that the vendor had misused the data. Priya could ask for the documents; she could not tell a candidate they existed.

At noon Helen offered a compromise: “Interview the obvious 60 now. We can audit the rest later.” Priya understood the cost of waiting. She also saw the trap. Once interviews began, a “temporary” pool could become the final pool because managers' diaries filled up, candidates accepted other jobs, and the team stopped looking below the line. A later audit would explain a decision already made. At the same time, a long pause would not distribute its costs evenly: an applicant with another offer would leave, while the current field crews would continue to cover. Neither speed nor caution was neutral.

Exhibit 1. What the pilot actually did

Stage Count Human action What remains unknown
Applications received 486 Basic completeness check only Whether every attachment was extracted correctly
Ranked by vendor 486 None before ranking Score weights, error rate and data retention terms
Above proposed interview cut line 60 Laila opened each application Whether the top group is the best group for each of the two role types
Below proposed cut line 426 Six opened after a separate licence search How many qualified applicants sit below the line
Interviews issued at decision time 0 None Which review rule Priya will authorise

The proportion below the cut line is 426 / 486, about 87.7%. That is the exposure to a process risk, not an estimate of the percentage wrongly excluded. “Human review” is accurate for the 60 opened files; it is not yet accurate as a description of the whole pool.

Exhibit 2. The Tuesday check

Priya asked an analyst to select ten files from each of four score bands: ranks 1-60, 61-160, 161-320 and 321-486. Forty original applications were compared with the extracted fields and the vendor rationales. In that deliberately balanced, non-random sample, five licences or licence dates were inconsistently extracted. Ten “continuity” rationales omitted context that was present in attachments. Some files had both problems. The six that first caught Laila's eye were not used as the forty-file sample.

The check does not establish a population error rate. Taking ten from each band over-represents the top 60, and the sample was designed to find failure modes, not to estimate their prevalence. It also says nothing on its own about disparate outcomes by age, disability, gender or other groups: North Quay had not designed or lawfully obtained an appropriate dataset for that test. Some low scores could reflect genuinely missing essential skills. Still, an extraction inconsistency in a required licence field has a different decision weight from a cosmetic wording error. A responsible next test would stratify by both role type and rank, use the source files, blind the reviewer to score, record disagreements and name a rule for reopening an application.

Exhibit 3. Capacity under the Friday deadline

Route Working assumption Capacity consequence What it cannot settle
Top 60 only Laila has reviewed these already Interviews could be invited within one day No check of the 426 excluded by the cut line
Full manual re-screen Seven minutes per file, including attachments and a recorded reason 486 x 7 minutes = 56.7 reviewer-hours, before reconciliation or appeals A hurried manual screen can make different errors
Expanded targeted review Laila can free 18 hours; a trained second reviewer can provide 24 hours before Friday About 360 additional files at seven minutes each beyond the top 60 already opened, if no rework; less if disagreements require a second pass Even with 60 files already opened, this leaves at least 66 of the other 426 for later review, before rework; no group-level fairness finding
Additional cover Helen can fund 24 reviewer-hours at an estimated loaded cost of $1,920 this week Buys time, but not an automatic decision rule Cover may not be trained or available next month

These are North Quay's invented planning assumptions. The 42 hours comprise Laila's 18 and the additional 24. The targeted route must say which files are selected, why, who resolves disagreements, and what happens to the at least 66 remaining. The top 60 were opened previously, not all checked with this new score-blind method; redoing them would reduce the 360 additional-file capacity. The $1,920 estimate is the fictional organisation's cost, not a market rate. No option makes all errors disappear by Friday. The operations case is not just recruitment speed: the contractor cover protects service this week but may crowd out maintenance later. The candidate case is not just a score: delayed interviews also affect people with less ability to wait.

Exhibit 4. Two messages on Priya's desk

Candidate message, Tuesday 10:40 am: “I hold the required licence and did this work for eight years. Was my application read?”

Vendor reply, Tuesday 2:15 pm: “The score is a recommendation only; final selection remains with your team. We can alter the continuity weighting today. A detailed score export and retention specification require an engineering follow-up.”

The vendor may be right that it does not make a final selection. That does not answer what the cut line does in practice. Nor does a weight adjustment explain the files already processed. What could Priya write to Samira now without implying a review that did not happen?

The case in numbers

Most of the numbers in this case are stated in the story or follow from it by arithmetic. One is an invented teaching assumption. The last column says which is which, so you can replace the invented ones with your own.

Metrics for this case
MeasureFigureWhy it mattersBasis
Field technician vacancies12 (4 independent crews, 8 supervised)The gap is two different problems, not twelve identical onesStated
Applications and recruiters486 applications, 1 recruiterOne person cannot read all of them in a weekStated
Applications a human read before the decision60 (12.3%)This is all the human review the ranking has hadDerived: 60 / 486
Applications nobody read426 (87.7%)This is exposure to a process risk, not an error rateStated
Tuesday check: files with a licence or date extracted inconsistently5 of 40 (12.5%)A sample built to find failures, so it is not a population rateStated count; percentage derived
Tuesday check: rationales that left out context in the attachments10 of 40 (25%)Same caveat: a balanced sample, not a random oneStated count; percentage derived
Full manual re-screen56.7 reviewer-hoursAgainst 42 hours of capacity, a shortfall of 14.7 hoursStated; shortfall derived
Extra cover Helen can fund24 hours for $1,920 ($80 an hour)Buys time, not a decision ruleStated; hourly rate derived
Contractor shift cover while roles stay openAbout $4,800 a week, rising next monthThe cost of waiting is real, and it falls on operationsInvented

Read the bottom row against the 87.7%. Waiting has a price that shows up on an invoice. The unreviewed 426 have a price that does not show up anywhere yet.

Critical evaluation

Question Working answer What remains open
What happened? All 486 applications were ranked; a person opened 60, plus six leads outside the group. The rule that turns an ordering into an exclusion.
What does the evidence support? Faster queue ordering and real extraction/rationale problems in a designed sample. Population false-negative rate, role-specific fit and group-level effects.
What else could explain the scores? Missing essential skills could account for some low ranks; attachment parsing could account for others. Source-file comparison by a reviewer who cannot see the score.
What test changes the decision? A resourced cross-rank, cross-role review with recorded reasons and a correction route. Vendor logs and data terms; whether the review can finish before invites.

DECISION MAP
Application + attachments -> field extraction -> score and rationale -> proposed cut line -> human review -> interview. The invisible decision is between score and review. Draw an override before the cut line and a correction path from candidate back to source file. If neither exists, the final human interviewer cannot see what the screen erased.

Figure 1. The interview cut line is a map of decision rights, not a claim of measured bias.

Application, extraction, scoring, cut line and human review, with an override and candidate correction path
Figure 1. The cut line marks a decision right, not a measured bias rate.

The decision room: Tuesday, 4:30 pm

Priya has until 5 pm to tell Helen and Laila which invitations can go out. She must also approve a response to Samira. Three paths are defensible only if their costs and limits are named.

Choice Immediate gain Immediate cost Second-order risk
A. Invite from the top 60 now Keeps Friday interview slots and relieves operations pressure 426 people remain unseen by a human The cut line becomes an unapproved rejection rule; the later audit cannot restore lost opportunities
B. Freeze invitations for a full manual and vendor review Makes no claim the organisation cannot support The 56.7-hour estimate exceeds this week's reviewer capacity; service pressure persists A safety pause without a dated decision owner becomes a hiring freeze
C. Use a temporary, documented manual selection path while auditing the ranking Some interviews can begin with reasons tied to published role criteria At least $1,920 of cover and 42 reviewer-hours; decisions may be slower and incomplete The temporary path can become another undocumented filter unless its sampling rule, escalation and end date are written

There is another boundary to test: is vendor data retention a precondition for using the tool at all, even merely to order the queue? If the answer is yes, Priya may have to stop processing while counsel checks the terms. If no, explain why keeping the service on is proportionate.

A fourth variant is possible: interview a small number of applicants whose required licences and work evidence are independently verified, while keeping every other application open for review. Is that C with a sharper rule, or A disguised? The answer depends on what happens to the rest of the pool. Do not read the follow-up until you have chosen a path, recorded what evidence would reverse it, and drafted Samira's reply.

OPERATING PRINCIPLE A human at the end cannot review the people who were removed before the human entered the process.

Separate teaching epilogue: open only after choosing a path

Priya chose a bounded version of C. She rejected A because the review record supported only a queue order, not the claim that the 60 contained all qualified candidates. She rejected an indefinite B because 56.7 hours exceeded the week's capacity while vacant roles carried a real cost for staff and customers. C bought a reversible, time-limited way forward, but she did not pretend it validated the vendor.

She authorised a first wave only after a reviewer checked the source licence and job evidence against the published criteria, without seeing the score. The team selected files from both above and below the proposed cut line and across the two role types; no score alone could close an application. The criteria, selection method, reviewer, reason and exceptions went in a register. Laila and the second reviewer each had an assigned workload; disagreements went to Priya before an invitation or rejection. Priya set Friday for a capacity review and two weeks for a go/no-go on continued use of ranking. If the review could not cover the remaining applications fairly, the team would extend the timetable rather than issue unreviewed rejections.

Samira's original licence had been extracted correctly, but her separate contract-history attachment was absent from the one-line rationale. Priya's reply did not say “every application was reviewed.” It said her application was being reassessed against the published criteria, named a person to contact about an error, and gave a date for a further response. It did not promise an interview.

At thirty days North Quay had filled eight roles. That invented outcome measures hiring progress, not the tool's fairness. The review register now identified who had examined each decision and on what grounds; unresolved vendor questions still barred score-only rejection. At ninety days the ranking service remained a way to organise the queue, not a gate. Priya had commissioned a larger, role-stratified validation with an explicit error and disagreement log. She would not treat an absent vendor retention schedule or a short rationale as a substitute for accountable records. Another organisation with less review capacity might reasonably choose B and reset its hiring dates. The case is not teaching C as a universal answer; it is teaching the need to defend the chosen trade-off.

Real incidents and guidance: parallels, not proof of this story

In 2023, the US EEOC announced a $365,000 settlement of allegations that iTutorGroup's software automatically rejected more than 200 older US applicants. That case concerns an explicit age rule, not North Quay's continuity score; its relevance is that software can create an exclusion before any manager meets an applicant. https://www.eeoc.gov/newsroom/itutorgroup-pay-365000-settle-eeoc-discriminatory-hiring-suit

In November 2024, the UK's ICO said audits of recruitment AI providers led to almost 300 recommendations. It described risks including filtering by protected characteristics, inferring characteristics from names and retaining more candidate data than needed. These are regulator observations about other providers, not findings about the fictional vendor here. https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2024/11/ico-intervention-into-ai-recruitment-tools-leads-to-better-data-protection-for-job-seekers/

The ICO's questions for recruiters include testing bias, explaining AI use to candidates, defining provider responsibilities and limiting unnecessary processing. Australia's OAIC guidance separately recommends checking intended-use testing, human oversight, privacy, security and access to personal data. Different jurisdictions have different legal tests; this case does not substitute for legal advice. https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2024/11/thinking-of-using-ai-to-assist-recruitment-our-key-data-protection-considerations/ ; https://www.oaic.gov.au/privacy/privacy-guidance-for-organisations-and-government-agencies/guidance-on-privacy-and-the-use-of-commercially-available-ai-products

For the teaching method: Harvard Business School's case-method note on quantitative material asks learners to test assumptions and connect calculations to practical decisions. We use that as a design reference, not as an affiliation or endorsement. https://www.hbs.edu/teaching/case-method/leading-case-discussion/teaching-quantitative-material

Editorial source: original AI Reality teaching-case format as shown in ARC-01-02, newly written scenario and exhibits. The story and exhibits are fictional, with public sources cited as parallels. Almost Magic Tech Lab Pty Ltd.

Facilitator guide

A private 90-minute session kit: run sheet, stakeholder questions, model analysis, transfer worksheet and quality check.

Request the facilitator guide