How to evaluate a take-home assignment fairly
On this page
To evaluate a take-home assignment fairly, decide what the task must prove before you write it, score every submission against the same weighted rubric with written anchors, have two reviewers score independently without the candidate's name or resume in front of them, and finish with a short conversation where the candidate explains their choices. The artifact alone is weak evidence; the artifact plus the explanation is strong.
Below: how to design an assignment worth a candidate's evening, a rubric and scoring sheet you can copy, the review process, what not to score, how to handle pay and AI use, and the fairness questions a work sample raises.
Decide what the assignment has to prove
A take-home is expensive for the candidate and slow for you, so it should answer a question the interview cannot. Before writing a brief, finish this sentence: "We are giving this assignment because we cannot tell from conversation whether the candidate can ___."
- Good reasons: whether they can structure a messy problem; whether their writing is clear to a reader who was not in the room; whether they can find the actual error in a real dataset; whether their design handles the edge cases nobody mentions.
- Poor reasons: "everyone does one"; to break a tie you could break with evidence you already have; to test enthusiasm; to see who will spend the most time.
- Pick one or two competencies, the same ones that appear on the scorecard. If the assignment tests four things, the score becomes an impression again.
Place it late in the process, after the hiring manager interview, when you already believe the candidate could do the job. Asking fifteen applicants to spend three hours each so you can shortlist four is a cost you are moving onto people who may never hear back from you.
Design rules that make a submission scoreable
- Time-box it and mean it. State the limit in the brief ("please spend no more than three hours"), and say what you will do about extra effort: "we score the rubric below, not polish or volume."
- Make the scope smaller than feels right. Reviewers consistently underestimate how long a brief takes someone without your context, your data and your shortcuts.
- Use a realistic problem, not a live one. A sanitized version of something you solved last year is ideal: realistic, already answered, and impossible to mistake for free labour.
- Give everyone the same brief, data and deadline, usually a week to return three hours of work, and be flexible on the deadline when someone asks. Flexibility on when costs you nothing; flexibility on what destroys comparability.
- Say how it will be assessed. Publishing the criteria, and even the weights, does not make the task easier; it removes the guessing that favors candidates who have seen your company's process before.
- Offer accommodations in the same sentence for everyone. The EEOC's page on job applicants and the ADA explains that reasonable accommodation covers changes to the job application process so a qualified applicant can be considered; extra time on an assignment is a common one.
- Name a contact for questions and share every answer with all candidates, so one person's clarification does not become an advantage.
The brief template
TAKE-HOME: [role] - [assignment name]
Time limit: [2-4] hours. Please do not exceed it; we score the criteria
below, not extra work.
Return by: [date]. Tell us if you need longer; that is not held against you.
Accommodations: if you need an adjustment to take part, reply to this
email and we will arrange it.
THE SITUATION
[3-6 sentences of context, sanitized, realistic]
WHAT TO PRODUCE
1. [artifact, with a length or size limit]
2. [a short note: assumptions, what you would do with more time]
DATA / MATERIALS
[links or attachments, everything they need]
HOW WE ASSESS IT
[criterion 1] - [weight]
[criterion 2] - [weight]
[criterion 3] - [weight]
[criterion 4] - [weight]
TOOLS AND AI
[your policy, stated plainly]
WHAT HAPPENS NEXT
Two reviewers score it against the criteria without your name attached.
You will hear by [date], and if you reach the next stage we will spend
20 minutes walking through your choices.
Pay, scope and the free-work question
Candidates object to take-homes for two reasons: the hours, and the suspicion that the work is useful to you. Both are solvable in the design.
- Never use a submission. If the output could ship, the assignment is a job, not an assessment. Say in the brief that submissions are deleted after the decision and used for nothing else, then actually delete them.
- Pay when it is long. Set your own threshold and apply it to everyone. Automattic publishes that its trial is paid at a standard rate of $25 USD an hour and typically takes 25 to 40 hours for non-Happiness roles (as of September 2026). That is the far end of the scale, and it shows the principle: at that length, pay is not a gesture.
- Offer an alternative. Some candidates cannot take unpaid hours, and some have existing work that answers your question. A portfolio walkthrough, a paired working session in your office, or a live exercise on a call can replace the take-home. Score it against the same rubric.
- Check the total ask. Add the hours across your whole process. Four interviews plus a four-hour assignment plus a presentation is a lot to ask of someone who is also doing a full-time job.
The rubric
Write the rubric before the first submission arrives, and write it from the brief, so a candidate who did exactly what was asked scores well. Four criteria, weights that add to 100, and anchors for each point on a 1-4 scale. The interview rating scale covers how to write anchors that two people read the same way.
RUBRIC: [assignment] Scale: 4 strong / 3 meets the bar / 2 below / 1 missing
Criterion Weight 4 3 2
--------------------------------------------------------------------------------
[Correctness of the 35% [what a 4 does] [what a 3 does] [what a 2 does]
core answer]
[Reasoning and 30%
assumptions stated]
[Communication to a 20%
non-expert reader]
[Judgment about 15%
what to leave out]
--------------------------------------------------------------------------------
Weighted score = sum(weight x score). Gate: any criterion at 1 = no.
A filled example
Invented role: Operations Analyst at a regional wholesaler. Assignment: "Here is a month of delivery data with known problems. In two hours, tell us where the late deliveries come from and what you would change. One page plus any workings."
| Criterion | Weight | A 4 looks like | A 2 looks like |
|---|---|---|---|
| Finds the real driver | 35% | Identifies that two depots account for most late deliveries and checks it against volume, not just counts | Lists every late delivery reason without ranking them |
| States assumptions and data problems | 30% | Names the duplicated rows and the missing timestamps, and says how they handled each | Uses the data as given and does not mention the gaps |
| Writes for the reader | 20% | One page, the recommendation first, numbers that support it | Six pages, the conclusion on the last one |
| Chooses what to leave out | 15% | Says what they did not do in the time and what they would do next | Attempts everything, finishes nothing |
Invented result: Candidate A scores 4, 4, 3, 3 = 0.35(4) + 0.30(4) + 0.20(3) + 0.15(3) = 3.65. Candidate B scores 4, 2, 4, 2 = 1.40 + 0.60 + 0.80 + 0.30 = 3.10, with the gap entirely in stating assumptions. That is a specific thing to probe in the walkthrough, not a tie to break by feel.
How to run the review
- Strip the identity. One person who is not scoring removes names and any obvious identifiers, and gives each submission a code. Reviewers do not open the resume before scoring.
- Two reviewers, independently. Same rubric, no discussion, scores submitted before they compare. The stage-by-stage bias controls apply here exactly as they do to interviews.
- Calibrate on the first two. Both reviewers score the first two submissions, compare line by line, and adjust the wording of the anchors before the rest are scored. Then rescore those two against the final wording.
- Time-box the review. 20 to 30 minutes per submission. A reviewer who spends two hours on one and ten minutes on another has produced two incomparable scores.
- Write evidence under each score: the line, the number, the paragraph that earned it. "Good work" is not a rating.
- Resolve gaps of two points or more by reading each other's evidence, not by averaging. Averaging a 4 and a 2 into a 3 describes nobody's view.
SCORING SHEET
Submission code: [ ] Reviewer: [ ] Time spent reviewing: [ ] min
Criterion 1 [ ]/4 - evidence: [quote or line reference]
Criterion 2 [ ]/4 - evidence:
Criterion 3 [ ]/4 - evidence:
Criterion 4 [ ]/4 - evidence:
Weighted score: [ ]
Questions for the walkthrough: [2-3]
Anything I could not assess: [ ]
What not to score
| Do not score | Why | Unless |
|---|---|---|
| Polish and design of the document | Rewards time and tools, not the skill you are testing | Presentation is in the rubric because the job requires it |
| Hours visibly spent | Punishes the candidate who respected your limit | Never |
| Tool or language choice | Familiarity, not capability | The brief named the tool |
| Writing style or fluency | Tied to first language and education; a proxy for national origin | Clear written communication is a defined criterion, judged on whether the reader can follow the argument |
| Going beyond the brief | Ignoring a stated limit is not a strength | You asked for a stretch section explicitly |
| Whether it matches how you would have done it | Different does not mean wrong | The rubric defines what correct means, and it is wrong |
AI tools and take-homes
Any assignment a candidate takes home can be done with an assistant, so an unstated policy just means you do not know what you are scoring. Write one into the brief and assess accordingly.
- Allowed, disclosed: "Use whatever tools you use at work. Add a line saying what you used and what you changed." This is closest to the job for most roles, and it makes the walkthrough the real test.
- Not allowed: only if you can say why the unaided skill matters for the job, and expect that you cannot verify it.
- Design around it. Ask for the reasoning, the assumptions, the trade-offs and what they would do with more time. Those are the parts a candidate has to own regardless of what produced the first draft.
- Never accuse on a hunch. Detection tools are unreliable, and an accusation is a serious thing to put on a record. If the submission and the walkthrough do not match, score what the candidate could explain and say that in the feedback.
The walkthrough conversation
Fifteen to twenty minutes, same questions for everyone, scored against the same rubric. It is the cheapest quality check in the process, and it turns the assignment from an artifact into evidence.
- "Walk me through how you approached it in the first twenty minutes."
- "Which assumption were you least comfortable with?"
- "What did you leave out because of the time limit, and what would you do next?"
- "If the data showed the opposite, what would change in your recommendation?"
- "Change one thing live with me: [small variation on the brief]. Talk me through it."
- "What did you use to produce this, and what did you change yourself?"
Score the walkthrough on the same criteria and record it on the scorecard alongside everything else, so the assignment feeds the same decision matrix as the interviews. When you compare finalists, the assignment is one row among several: see how to compare candidates after interviews.
Fairness, records and feedback
Not legal advice. This is a general summary of US federal material as of September 2026, and rules differ by location. Confirm your own process with HR or employment counsel.
A take-home used to decide who moves forward is a selection procedure. The Uniform Guidelines on Employee Selection Procedures define a selection procedure as "any measure, combination of measures, or procedure used as a basis for any employment decision" (29 CFR 1607.16(Q)), which brings work samples inside the same framework as tests, including the attention paid to whether a procedure screens out one group at a substantially different rate. In practice that means three habits:
- Keep the rubric, the scores and the evidence, with the rest of the hiring record for the role.
- Watch who drops out. If candidates regularly decline the assignment or fail to return it, the length or the timing is doing the screening, not the skill.
- Check the assignment against the job. Every criterion should map to something the person will do. A puzzle that correlates with nothing in the role is hard to defend and harder to justify to a candidate.
Give every candidate who completed it a real answer, quickly, and two lines of specific feedback: it is the least you owe someone who spent three hours on your problem. The rejection email templates include a version for after a work sample.
Checklist
- The assignment answers one question the interviews cannot.
- Time limit stated, scope tested by someone on your team against that limit.
- The problem is sanitized and already solved internally; submissions are never used.
- Same brief, data, deadline and rubric for every candidate at the stage.
- Criteria and weights published in the brief; accommodations offered to everyone.
- Pay decided by a rule you apply to all, and an alternative offered for those who cannot take unpaid hours.
- Rubric with anchors written before the first submission; calibrated on the first two.
- Names stripped; two reviewers scoring independently; evidence under every score.
- AI policy stated in the brief, and a walkthrough that tests understanding either way.
- Scores, evidence and rubric kept with the hiring record; feedback sent within days.
Questions people ask
How long should a take-home assignment be?
Long enough to show the skill and short enough that a candidate with a job and a family can do it: two to four hours is a workable ceiling for most roles. Write the time limit in the brief, and say plainly that you will not reward extra hours.
Should candidates be paid for a take-home assignment?
Pay if the task is long or if the output is useful to you. Automattic, for example, publishes that its trial is paid at $25 an hour and typically takes 25 to 40 hours for non-Happiness roles. If you would feel uncomfortable paying for it, that is a sign the assignment is too big.
Who should review a take-home assignment?
Two reviewers, scoring independently against the same rubric, with the candidate's name and resume out of view. One reviewer with the resume open is the setup most likely to produce a score that matches the pedigree.
What if a candidate used an AI tool on the assignment?
Decide the policy before you send the brief and state it there. Whatever you allow, a 15-minute walkthrough where the candidate explains their choices and changes something live tells you more about their understanding than the artifact does.
Do we have to give every finalist the same assignment?
Yes, if you plan to compare the results. Same brief, same time limit, same rubric, same reviewers where possible. An assignment given to one candidate only is evidence about that candidate and nothing else.