The rubric, in the open
How we screen — no black box
Every agency in this category says 'rigorous vetting' and then shows you nothing — 'filters out 99.7%', 'top 1%', numbers you can't check. Here is the exact rubric we score every candidate against. Read it, then run it on yourself and watch it behave the same way.
The scoring scale
Scores aren't vibes — they're anchored to a fixed distribution so a number means the same thing every time. Most real applicants land 55–75; a score of 85+ has to be genuinely in the top tenth of everything we'd see for that role.
| Score | Band | What it means |
|---|---|---|
| 90–100 | Exceptional | Top ~2%. Reserved for answers that teach the screener something. Almost never awarded. |
| 80–89 | Top ~10% | Specific and evidenced — we'd present them to a paying client without hesitation. |
| 70–79 | Strong | Real experience, minor gaps. Advances to the second-round interview. |
| 55–69 | Competent but generic | The modal band — where the typical decent applicant lands. Competent, but no specifics. |
| 40–54 | Thin | Vague claims, no numbers or tools, boilerplate phrasing. |
| 0–39 | Unusable | Off-role, copy-paste, AI-slop, or incoherent English. |
What actually counts, in order
Not every answer weighs the same. The graded work sample outweighs any interview answer — because talk is cheap and the work isn't.
| Weight | What we score | Why |
|---|---|---|
| 35% | Graded work sample | The single heaviest input. We ask for a real deliverable — a set of replies, a reconciliation, a JD, working code — and grade the actual work, not a description of it. |
| 25% | Scenario judgment | A hard, role-specific situation. We're reading for how they'd actually handle it — the decision, the tradeoff — not a textbook answer. |
| 20% | Experience specificity | Real names, numbers, and tools — not adjectives. 'I cut first-response time from 6 hours to under 2' beats 'I have great communication skills' every time. |
| 15% | Written & spoken English | They'll work directly with a US company. Written answers carry most of this; when a candidate's recorded intro is in, its transcript is read under the same dimension — and the clip itself goes to the client unedited. |
| 5% | Salary-band fit | A small nudge — whether their ask lands in the band you set. |
What a strong answer looks like — and a weak one
The line is specificity. Two candidates can answer the same customer-support scenario and score 40 points apart:
Scores well — specific, evidenced
“I'd acknowledge the second breakage first and own it, then make the refund right before it's asked for twice. I ran a Shopify inbox where I cut repeat-contact rate by rewriting the refund macro and adding a proactive replacement offer.”
Scores low — adjectives, no evidence
“I'm very hardworking and a fast learner with excellent communication skills. I always go the extra mile to make customers happy and I'm great at handling difficult situations.”
The rubric penalizes hard for the second kind: adjectives without evidence, AI-sounding boilerplate, answers that restate the question, unverifiable claims with no specifics.
Then a second round, written from their own answers
Anyone who clears round one gets a short second interview — three questions the AI writes from that candidate's own round-one answers and your brief, not a canned list. That's what catches a pasted or AI-written first round: it's hard to keep a fabricated story straight when the follow-ups are specific to what you claimed. The bar in round two is higher, and it's graded on whether they've actually done the thing — specific systems, numbers, decisions — not whether they can describe it.
Don't take our word — run it
The rubric above is live on this site. Answer a real scenario question and it scores you on the spot, with the same calibration and the same penalize-vague rules our pipeline uses.