Looking for work with a US company? Apply to the Rolemote talent roster — free →

The scoring scale

Scores aren't vibes — they're anchored to a fixed distribution so a number means the same thing every time. Most real applicants land 55–75; a score of 85+ has to be genuinely in the top tenth of everything we'd see for that role.

ScoreBandWhat it means
90–100ExceptionalTop ~2%. Reserved for answers that teach the screener something. Almost never awarded.
80–89Top ~10%Specific and evidenced — we'd present them to a paying client without hesitation.
70–79StrongReal experience, minor gaps. Advances to the second-round interview.
55–69Competent but genericThe modal band — where the typical decent applicant lands. Competent, but no specifics.
40–54ThinVague claims, no numbers or tools, boilerplate phrasing.
0–39UnusableOff-role, copy-paste, AI-slop, or incoherent English.
The production score bands — a candidate's number maps to exactly one of these.

What actually counts, in order

Not every answer weighs the same. The graded work sample outweighs any interview answer — because talk is cheap and the work isn't.

WeightWhat we scoreWhy
35%Graded work sampleThe single heaviest input. We ask for a real deliverable — a set of replies, a reconciliation, a JD, working code — and grade the actual work, not a description of it.
25%Scenario judgmentA hard, role-specific situation. We're reading for how they'd actually handle it — the decision, the tradeoff — not a textbook answer.
20%Experience specificityReal names, numbers, and tools — not adjectives. 'I cut first-response time from 6 hours to under 2' beats 'I have great communication skills' every time.
15%Written & spoken EnglishThey'll work directly with a US company. Written answers carry most of this; when a candidate's recorded intro is in, its transcript is read under the same dimension — and the clip itself goes to the client unedited.
5%Salary-band fitA small nudge — whether their ask lands in the band you set.
Round-one weighting. The work sample is deliberately the heaviest input.

What a strong answer looks like — and a weak one

The line is specificity. Two candidates can answer the same customer-support scenario and score 40 points apart:

Scores well — specific, evidenced

“I'd acknowledge the second breakage first and own it, then make the refund right before it's asked for twice. I ran a Shopify inbox where I cut repeat-contact rate by rewriting the refund macro and adding a proactive replacement offer.”

Scores low — adjectives, no evidence

“I'm very hardworking and a fast learner with excellent communication skills. I always go the extra mile to make customers happy and I'm great at handling difficult situations.”

The rubric penalizes hard for the second kind: adjectives without evidence, AI-sounding boilerplate, answers that restate the question, unverifiable claims with no specifics.

Then a second round, written from their own answers

Anyone who clears round one gets a short second interview — three questions the AI writes from that candidate's own round-one answers and your brief, not a canned list. That's what catches a pasted or AI-written first round: it's hard to keep a fabricated story straight when the follow-ups are specific to what you claimed. The bar in round two is higher, and it's graded on whether they've actually done the thing — specific systems, numbers, decisions — not whether they can describe it.

Don't take our word — run it

The rubric above is live on this site. Answer a real scenario question and it scores you on the spot, with the same calibration and the same penalize-vague rules our pipeline uses.

Try the screener on yourself → See a scored shortlist