Your best candidates say no
The first people to decline a long unpaid exercise are the ones holding competing offers. The exercise does not filter your shortlist — it thins it, from the top.
A work sample for data analysts
A take-home case study eats your candidate's weekend and your analyst's afternoon, and still cannot tell you whether they understood the data or got lucky. This is fifty minutes inside a real company — and it shows you the working.
One real question against a real company, with the real console. No sign-up, and nothing is recorded.
Six hours of theirs. An afternoon of yours. And you still cannot tell whether they understood it.
Checked the data before committing on seven of nine. Right on seven. Never once asserted something they had not tested.
Fifty-five minutes of theirs. Two of yours. And you see every step they took.
The first people to decline a long unpaid exercise are the ones holding competing offers. The exercise does not filter your shortlist — it thins it, from the top.
An eight-hour task quietly removes anyone with caring responsibilities, a second job, or a current role with real demands. None of those correlate with being good at the work.
A tidy final answer arrives and someone spends an afternoon on it. Did they understand the data, or get lucky? You never see what they tried first, what they checked, or what they chose to ignore.
So: a shorter exercise that is the actual job, and a report that shows the working rather than only the answer.
People cost is 21% over plan. Divide the overspend by the planned rate and it looks like nine extra heads. Freeze hiring.
That arithmetic cannot work — a pay rise shows up in it as an extra person. Count the payroll instead and most of the gap is paying over the odds to replace leavers. A hiring freeze fixes the smaller half.
The clock does not start until they finish reading. Reading speed is not the thing being measured, so it is not measured.
Twelve colleagues bring them real disagreements — is the cheap channel actually working, is it the leads or the reps? A console sits inside the assessment: six to twelve tables depending on the seat, tens of thousands of rows, the company's own CRM and billing and marketing spend. Nobody has done the analysis for them.
A third of the deals have no source recorded. The column called
stage_changed_on holds the date a rep updated the record, not the date
anything happened. Some activity rows are duplicated. Take the tables at face value
and you reach a confident wrong answer — the same one the company's own dashboard
already reached.
Nine decisions, each with a stated confidence and the result they are based on. Optionally they write 40–150 words explaining any conclusion they assert firmly.
Every candidate draws a different company, from a pool of forty that have each been measured for how hard they are — so two people are compared on a common scale rather than on identical questions. The first person you invite cannot hand the second one an answer key, because the answers are in rows only their own company has.
You pick the seat when you set up the role. The candidate meets the questions that seat actually gets asked — a marketing analyst is not asked to size a sales team.
Pipeline, forecasting, quota attainment, sales productivity.
Where the pipeline is stuck, whether it is the leads or the reps, what discounting costs, and whether the forecast can be trusted at all.
One thing that catches people: half the team has not been here long enough to be at full productivity. Compare them on conversion alone and you manage out the wrong person.
Channel mix, acquisition cost, attribution, campaign performance.
Which channel earns the next pound, what a customer really costs once attribution is accounted for, and how much of the reporting rests on records that never captured a source.
One thing that catches people: a third of the deals never recorded a source. The channel report is built from the ones that survived.
Churn, cohorts, onboarding, net revenue retention, expansion.
Why accounts leave, whether the ones that landed badly leave faster, and what a segment is worth once you count how long it stays.
One thing that catches people: by count, the answer is always the smallest accounts, because there are more of them. Only a rate says anything.
Activation, usage, retention curves, experiments.
Whether the accounts you signed ever got going, how much revenue is sitting in ones nobody uses, whether the usage numbers can be trusted, and what an experiment actually showed.
One thing that catches people: bigger accounts generate more of every event and churn less. Count rows and you have measured account size.
Budget against actual, variance, headcount cost, forecasting.
Whether the company is spending what it said it would, whether being over is more people or costlier people, which team is furthest from its own plan, and whether the monthly numbers could be shown to a board.
One thing that catches people: every cost carries two dates — the month it belongs to and the day finance recorded it. Total on the wrong one and money moves between months that never moved.
Pick one or two per role. Nothing to install; it all runs in the candidate's browser.
Five decisions, none of which takes longer than a minute. Everything after that is generated — the company, the questions, the data.
These are the whole of what is scored, in the words the report uses. Each one is counted rather than judged — there is no language model reading anything, anywhere in the scoring path.
Did they run the analysis before committing, and did the result they cited actually bear on the question? Near-deterministic, and the most predictive thing here: being right is partly the draw, how they got there is what you are hiring.
Every call carries a stated confidence. This counts the times they committed hard to something they had not checked and were wrong — and credits the times they hedged exactly where the evidence was thin. It is the rarest thing on the report.
Were they correct. Deliberately not the largest measure, and never read on its own: over a dozen decisions this figure carries enough noise that a good analyst and a mediocre one can change places by luck.
Every number in their written reasoning is matched against analyses they actually ran. The engine knows exactly what they queried and what it returned, so it knows whether a quoted figure is real. The one thing here a machine can verify perfectly, and the one thing a polished answer can hide.
Twelve colleagues act on this person’s advice only as far as they believe it, and revise that view each month on the record so far. It is how the work landed on the people around them, which is not the same question as whether it was right.
More is brought to an analyst than they can do. What they got to, and what they chose to leave, is a judgement in itself — and under a clock it is often the most revealing one.
Two candidates. Both got 3 of 9 right. Here is everything a right-or-wrong screen would have filed as identical.
Answered fast and firmly, checked almost nothing, and quoted two figures that appear in none of their own results.
Did the analysis on seven of nine, hedged where the data was thin, and every number they wrote down traces back to a query they ran.
Thirty-seven points apart — on identical accuracy. One of these people is worth an interview and the other is a risk, and nothing about which boxes they ticked tells you which is which.
How well they write, or how they think. Nothing here reads their prose and forms a view of it — that would need a model in the scoring path, which is the one thing this product will not do. What you get instead is their reasoning in full, with every figure in it marked green where it matches their own work and amber where it appears nowhere. Reading the argument takes you thirty seconds, and the judgement stays yours.
See a finished report A real one, scored by the same engine, and downloadable as a PDF. No sign-up, and there is no real person in it. It comes with a shorter version too — the six measures and what would move each of them, carrying no answers and no verdict — which is the one you forward to a manager, or send back to the candidate.
The demo above has this switched off, because you should be able to check Slack in the middle of it. A real sitting does not.
The clock never stops. The deadline is wall-clock and enforced on the server, so time spent away is time spent. The two minutes are not extra thinking time, which is what removes the incentive to use them.
There is no grace period, and there used to be. Sixty seconds, so a screen lock would not cost anybody a strike. It protected the wrong person: a minute is long enough to paste the question into a language model, read the answer and come back, with the product saying nothing throughout.
No camera, no microphone, no screen recording. Those are denied by a header the browser enforces, not by a promise in a policy. The only thing observed is whether this page was in front. See Security.
There is no language model anywhere in the scoring path. We check whether the figures a candidate quotes match the analyses they actually ran, and mark each one for you. Judging the argument stays with a person — yours.
That is a commercial position as much as a technical one. A model deciding part of the score makes a tool an automated employment decision system: high-risk under the EU AI Act, with mandatory risk assessment, bias testing, technical documentation, human oversight and continuous monitoring, and an automated employment decision tool under New York City Local Law 144. Those obligations land on the employer as well as the vendor.
Every number here is arithmetic you can inspect, and every judgement is made by a person. There is nothing to audit because there is no model to audit.
Not for New York City roles yet — Local Law 144 requires an independent bias audit, and we do not have one.
The same fifty minutes works as a course capstone. Your students find out where they actually stand before they start interviewing, which is the one thing a certificate cannot tell them.
I have sat on both sides of this. I have been sent take-home case studies that ate a weekend and told the company almost nothing, and I have been the one reading them — trying to work out from a tidy final answer whether somebody understood the data or got lucky, with no way to see what they tried first.
The people who decline a long unpaid exercise first are the ones with competing offers, so the exercise quietly makes the shortlist worse. And an eight-hour task filters out anyone with caring responsibilities or a demanding current job, neither of which has anything to do with being good at the work.
So: fifty minutes, the real job, and every query they ran on the way to their answer. Judged by you, not by a model.
Sanyam Gulati, founder
Tell us what you are hiring for and we will set you up, usually the same day. Bring your own analysts too — running your team through it first is free, and it is what makes every candidate result afterwards mean something.
Already set up? Open the hiring console.