Assessment · Hiring

The Work-Sample Test: Hiring's Most Honest Hour

By Tom Christian July 23, 2026 ~9 min read

Every interview process contains a quiet substitution that almost nobody notices. You need to know whether a candidate can do the work, so you spend three hours asking them to talk about the work — and then you grade the talking as if it were the doing. But talking about work and doing work are different skills, held by overlapping-but-different people. The confident narrator who interviews beautifully and delivers nothing is a hiring cliché for a reason. So is the quiet candidate who fumbles the story of their own competence and then quietly outperforms everyone.

Decades of selection research keep pointing at the same unglamorous conclusion: if you want to predict how someone will perform at a job, one of the strongest signals you can collect is a sample of them performing the job. Not describing it. Doing a small, representative, fairly-scored piece of it. The work-sample test is the most honest hour in hiring — and most small companies either skip it entirely or run one so badly designed it measures stamina and free time instead of skill.

Here's how to build one that measures what you think it measures — and what has to change now that every candidate has AI in the room with them.

Design it small, representative, and paid-shaped

Three design rules separate a fair work sample from an exploitative one, and they're the three rules candidates silently judge you by.

  1. Small. Sixty to ninety minutes of work, hard-capped, and say the cap out loud: "We're expecting about an hour. Please stop there — more polish will not score higher." An uncapped take-home is a stamina contest that filters for people with free evenings, which is a filter for life circumstances, not ability. If the task can't be meaningfully attempted in ninety minutes, the task is too big — shrink the task, not the fairness.
  2. Representative. The sample must be a slice of the actual job — a support candidate answers two realistic tickets; an ops candidate untangles a scheduling conflict; a trainer outlines a fifteen-minute lesson from raw material. Resist the clever abstract puzzle. You're not measuring cleverness; you're measuring the Tuesday-morning work, so sample the Tuesday-morning work.
  3. Respectful of the exchange. Use invented scenarios, never your real backlog — free labor dressed as assessment is how you end up on the wrong kind of hiring forum. For larger samples or later stages, paying a modest honorarium is both fair and a signal of seriousness that costs less than one bad-hire week.

Score it blind, against a rubric written first

The sample is only as honest as its scoring. Two disciplines carry all the weight here. First, the rubric exists before the first submission arrives — three to five criteria that define what weak, solid, and strong look like for this exact task. Writing it afterward means writing it around your favorite candidate; that's not scoring, that's rationalizing. Second, score blind where you can: strip names before review, score all submissions on one criterion at a time rather than one candidate at a time, and write one line of evidence next to every mark — the specific thing in the submission that earned the level. A mark with a receipt survives a challenge, a hiring debrief, and your own second thoughts a week later. A mark without one is a vibe with a number attached.

The AI-era update: assess the judgment, not the typing

Now the part that has changed under everyone's feet. Any take-home you send today will be attempted with AI assistance by a meaningful share of candidates, and you will not reliably detect which ones. You have two options. You can run an arms race — forbid the tools, deploy the detectors, escalate the surveillance — and lose, slowly and expensively, while filtering for rule-followers rather than performers. Or you can redesign the assessment around the question that actually matters now: can this person produce good work with the tools they'll actually use — and can they defend it?

The redesign is simpler than it sounds. Allow the tools explicitly. Then move the assessment weight from the artifact to the reasoning behind it: after the submission, a twenty-minute conversation where the candidate walks you through their work. Why this approach and not the obvious alternative? What's the weakest part of what you submitted? What did the AI get wrong that you had to catch — and how did you know? A candidate who genuinely owns their work answers these easily, and often eagerly. A candidate who piped the prompt through a model and skimmed the output cannot — the seams show within three questions. The defense of the work, not the work alone, is the assessment. Which is exactly the principle serious learning assessment has been forced to adopt for the same reason: when the artifact is cheap to fake, evaluate the thinking that produced it.

What this filters for — and what it filters out

Run samples this way and the profile of who clears your process shifts in directions you want. The confident narrator without the skill stops clearing, because the sample doesn't care about narration. The skilled-but-modest candidate starts clearing, because the work speaks where they wouldn't. The AI-fluent pragmatist clears — and you want them; that fluency is now part of the job. The prompt-piper doesn't, because the defense round exists. And the candidate with a full life and no free weekends stays in your funnel, because you capped the hour and meant it. Every one of those shifts is the assessment getting more honest, and honesty in selection compounds: the team you build this way trusts your next process more, and says yes to it faster.

The bottom line

Interviews will always have a seat — culture, motivation, and the shape of a person don't show up in a work sample. But the core question of hiring is "can they do the job?", and the most honest answer available is an hour of the job itself: small enough to be fair, representative enough to mean something, scored blind against a rubric that existed first, with every mark carrying its evidence — and, in the AI era, finished with the only test that survives cheap artifacts: walk me through why. The candidates who can, are the ones you were always trying to find.

The defense round, made rigorous

Crucible runs structured oral defenses and scores every rubric criterion with a verbatim quote from the candidate's own words, a confidence level, and an honest "insufficient evidence" flag — a human confirms every mark. The same evidence discipline your work-sample deserves.

See how Crucible works