Game-Based Assessments in Hiring: What They Measure and How We Keep Them Fair
A CV tells you where someone has been. An interview tells you how they talk about it. Neither tells you how someone actually holds four items in working memory under time pressure, resists a tempting but wrong response, or keeps calibrating risk after a loss. Those are measurable behaviors, and cognitive science has been measuring them for decades with short, standardized tasks. StormInterview now brings 14 of those tasks into the hiring flow as game-based assessments: small, friendly, one-to-three-minute games a candidate can play on any phone or laptop as part of an async interview.
This article explains what the games measure, where the science comes from, how the scoring stays explainable, and the guardrails that keep the whole thing fair: for candidates, for recruiters, and under the EU AI Act's rules for hiring technology.
What the games measure
Each game targets a specific, well-studied construct. A few examples from the library:
- Working memory: repeating digit sequences backwards (Backward Digit Span, in the tradition of Wechsler's scales) and tapping spatial sequences (Corsi block-tapping task, Corsi, 1972).
- Attention and impulse control: responding to go-signals while withholding on no-go signals (Go/No-Go, Donders' tradition; Verbruggen & Logan, 2008), ignoring distracting flankers (Eriksen flanker task, Eriksen & Eriksen, 1974), and cancelling an already-started response (stop-signal task, Logan & Cowan, 1984).
- Planning: rebuilding a target tower in as few moves as possible (Tower of London, Shallice, 1982).
- Risk and decision-making: pumping a balloon for points and cashing out before it bursts (Balloon Analogue Risk Task, Lejuez et al., 2002) and learning which card decks pay off over time (Iowa Gambling Task, Bechara et al., 1994).
- Effort-based choice: choosing between a small easy reward and a large effortful one (EEfRT, Treadway et al., 2009).
- Social judgment: sending and returning coins in the classic trust game (Berg, Dickhaut & McCabe, 1995) and reading emotions from faces (in the Ekman tradition).
The citation is not marketing decoration. It is shown inside the product, next to the game in the template builder and next to every score in the review screen, so everyone can check what a task is and where it comes from.
Playful for candidates, precise underneath
Candidates see warm, character-driven mini-games: a balloon with a face that gets nervous as the risk grows, stepping stones in a pond, a drum to tap as fast as you can. That craft is deliberate, because a relaxed candidate gives a more representative sample of behavior than a stressed one. Underneath, the measurement layer is strict: stimuli appear instantly, nothing animates between a stimulus and the response it measures, and all the celebration happens after an answer is recorded. The playfulness lives in the presentation layer only, and the recorded data is identical to what the classic lab versions of these tasks would capture.
Explainable scoring, versioned like software
Every game submission is scored server-side by a documented, deterministic formula. The review screen shows named traits on a 0-to-1 scale (for example risk tolerance or impulse control), the supporting numbers behind them, the percentile when enough normative data exists, and the scoring engine version. When there is not enough data for a fair percentile yet, the product says exactly that instead of inventing one. There is no neural network guessing at a personality; there is a formula you can read, tied to an engine version that never silently changes old scores.
The fairness guardrails
- Advisory only, human decides. Game outcomes never auto-reject anyone. The platform structurally blocks automations from acting on game results, and scores appear as one signal next to interview answers. This matches how the EU AI Act treats hiring as a high-risk domain: transparency, human oversight and no automated rejection.
- Consent first. Games run under an explicit, separate consent step. A candidate who declines is not scored and not penalized for it.
- An accessible alternative, always. A game step cannot be published without a configured alternative question. Skipping a game routes the candidate to that question. Keyboard play, screen-reader announcements and reduced-motion modes are built in.
- Integrity checks. Server-side validation catches implausible speed, impossible sequences, manipulated payloads and missing trial types. Suspicious attempts are flagged for a human or invalidated, never quietly scored.
- Norms grow honestly. Percentiles come from real, consented usage and appear only once the sample behind them is large enough to be meaningful.
Where this fits in your funnel
Games work best as an enrichment layer, not a gate. A common setup: an async video interview carries the role-specific questions, and two or three short games add signal on the traits the role genuinely needs, for example working memory and attention for analytical roles, or risk calibration for commercial ones. Recruiters review everything in one place, compare candidates on the same structured evidence, and decide.
Game-based assessments are available per workspace as a pilot feature. If you want your team to try them, they can be enabled for your environment in minutes.
Start a free trial of StormInterview and see what a candidate's working memory, focus and judgment look like next to their interview answers, with a human making every call.