A work sample for data analysts

See them work before you hire them.

A take-home case study eats your candidate's weekend and your analyst's afternoon, and still cannot tell you whether they understood the data or got lucky. This is fifty minutes inside a real company — and it shows you the working.

One real question against a real company, with the real console. No sign-up, and nothing is recorded.

The take-home
“The reps are the problem.”

Six hours of theirs. An afternoon of yours. And you still cannot tell whether they understood it.

Trial Day
Every decision, and what they did first

Checked the data before committing on seven of nine. Right on seven. Never once asserted something they had not tested.

checked, right checked, wrong asserted

Fifty-five minutes of theirs. Two of yours. And you see every step they took.

The take-home case study is failing you three ways

Your best candidates say no

The first people to decline a long unpaid exercise are the ones holding competing offers. The exercise does not filter your shortlist — it thins it, from the top.

It filters for free time, not for skill

An eight-hour task quietly removes anyone with caring responsibilities, a second job, or a current role with real demands. None of those correlate with being good at the work.

You still cannot tell what happened

A tidy final answer arrives and someone spends an afternoon on it. Did they understand the data, or get lucky? You never see what they tried first, what they checked, or what they chose to ignore.

So: a shorter exercise that is the actual job, and a report that shows the working rather than only the answer.

What the candidate actually does

What the total says
  • Planned £4.1m
  • Actual £5.0m

People cost is 21% over plan. Divide the overspend by the planned rate and it looks like nine extra heads. Freeze hiring.

What the payroll says
  • More people 42%
  • Paying above plan 58%

That arithmetic cannot work — a pay rise shows up in it as an extra person. Count the payroll instead and most of the gap is paying over the odds to replace leavers. A hiring freeze fixes the smaller half.

  1. They read the brief, untimed

    The clock does not start until they finish reading. Reading speed is not the thing being measured, so it is not measured.

  2. They query the company's database

    Twelve colleagues bring them real disagreements — is the cheap channel actually working, is it the leads or the reps? A console sits inside the assessment: six to twelve tables depending on the seat, tens of thousands of rows, the company's own CRM and billing and marketing spend. Nobody has done the analysis for them.

  3. They find out what the tables are hiding

    A third of the deals have no source recorded. The column called stage_changed_on holds the date a rep updated the record, not the date anything happened. Some activity rows are duplicated. Take the tables at face value and you reach a confident wrong answer — the same one the company's own dashboard already reached.

  4. They commit, and say how sure they are

    Nine decisions, each with a stated confidence and the result they are based on. Optionally they write 40–150 words explaining any conclusion they assert firmly.

Every candidate draws a different company, from a pool of forty that have each been measured for how hard they are — so two people are compared on a common scale rather than on identical questions. The first person you invite cannot hand the second one an answer key, because the answers are in rows only their own company has.

Five kinds of analyst, and the questions each one gets

You pick the seat when you set up the role. The candidate meets the questions that seat actually gets asked — a marketing analyst is not asked to size a sales team.

Revenue & Sales Operations

Pipeline, forecasting, quota attainment, sales productivity.

Where the pipeline is stuck, whether it is the leads or the reps, what discounting costs, and whether the forecast can be trusted at all.

One thing that catches people: half the team has not been here long enough to be at full productivity. Compare them on conversion alone and you manage out the wrong person.

Marketing & Growth

Channel mix, acquisition cost, attribution, campaign performance.

Which channel earns the next pound, what a customer really costs once attribution is accounted for, and how much of the reporting rests on records that never captured a source.

One thing that catches people: a third of the deals never recorded a source. The channel report is built from the ones that survived.

Customer & Retention

Churn, cohorts, onboarding, net revenue retention, expansion.

Why accounts leave, whether the ones that landed badly leave faster, and what a segment is worth once you count how long it stays.

One thing that catches people: by count, the answer is always the smallest accounts, because there are more of them. Only a rate says anything.

Product & Growth

Activation, usage, retention curves, experiments.

Whether the accounts you signed ever got going, how much revenue is sitting in ones nobody uses, whether the usage numbers can be trusted, and what an experiment actually showed.

One thing that catches people: bigger accounts generate more of every event and churn less. Count rows and you have measured account size.

Finance & FP&A

Budget against actual, variance, headcount cost, forecasting.

Whether the company is spending what it said it would, whether being over is more people or costlier people, which team is furthest from its own plan, and whether the monthly numbers could be shown to a board.

One thing that catches people: every cost carries two dates — the month it belongs to and the day finance recorded it. Total on the wrong one and money moves between months that never moved.

Assessed in the tools the job actually uses

  • SQL against the company's own database — six to twelve tables, tens of thousands of rows.
  • Python with pandas, for the seats where retention curves and experiments are the work.
  • Excel — a pivot table over the same data, with several measures at once, filters and calculated columns by name. Budget against actual, side by side, which is how that work is done almost everywhere it is done at all.

Pick one or two per role. Nothing to install; it all runs in the candidate's browser.

You set it up for the role, not for analysts in general

Five decisions, none of which takes longer than a minute. Everything after that is generated — the company, the questions, the data.

The seat
Revenue operations · Marketing · Customer retention · Product · Finance The candidate meets the questions that seat is actually asked. A marketing analyst is not asked to size a sales team.
The tool
SQL · Python · a pivot table Pick one or two. Nothing to install — it all runs in the candidate's browser, against a real database of tens of thousands of rows.
How much help
Guided · Standard · Unaided Guided says what to look at and what to watch out for. Unaided hands them the situation and the company and nothing else. Same questions, same right answers, same clock — only the scaffolding moves.
Written reasoning
On or off When it is on, a candidate asserting something firmly is asked to write 40–150 words explaining why, and every figure they quote is checked against analysis they actually ran. Switching it off does not cost them marks.
The clock
50 minutes by default Adjustable per candidate, which matters more than it looks: a fixed timer with no way to grant extra time is a legal problem in hiring, not merely an unkind one. Reading the briefing is untimed either way.

Six things about a candidate that a right-or-wrong screen cannot see

These are the whole of what is scored, in the words the report uses. Each one is counted rather than judged — there is no language model reading anything, anywhere in the scoring path.

Shows their working

Did they run the analysis before committing, and did the result they cited actually bear on the question? Near-deterministic, and the most predictive thing here: being right is partly the draw, how they got there is what you are hiring.

Knows when they’re sure

Every call carries a stated confidence. This counts the times they committed hard to something they had not checked and were wrong — and credits the times they hedged exactly where the evidence was thin. It is the rarest thing on the report.

Reached the right answer

Were they correct. Deliberately not the largest measure, and never read on its own: over a dozen decisions this figure carries enough noise that a good analyst and a mediocre one can change places by luck.

Quotes figures they can back

Every number in their written reasoning is matched against analyses they actually ran. The engine knows exactly what they queried and what it returned, so it knows whether a quoted figure is real. The one thing here a machine can verify perfectly, and the one thing a polished answer can hide.

Kept the room’s trust

Twelve colleagues act on this person’s advice only as far as they believe it, and revise that view each month on the record so far. It is how the work landed on the people around them, which is not the same question as whether it was right.

Serves the business

More is brought to an analyst than they can do. What they got to, and what they chose to leave, is a judgement in itself — and under a clock it is often the most revealing one.

Two candidates. Both got 3 of 9 right. Here is everything a right-or-wrong screen would have filed as identical.

Candidate A
  • Shows their working 1 of 9
  • Knows when they’re sure 5 overclaimed
  • Quotes figures they can back 2 unbacked

Answered fast and firmly, checked almost nothing, and quoted two figures that appear in none of their own results.

Candidate B
  • Shows their working 7 of 9
  • Knows when they’re sure 0 overclaimed
  • Quotes figures they can back all backed

Did the analysis on seven of nine, hedged where the data was thin, and every number they wrote down traces back to a query they ran.

Thirty-seven points apart — on identical accuracy. One of these people is worth an interview and the other is a risk, and nothing about which boxes they ticked tells you which is which.

And one thing we do not measure

How well they write, or how they think. Nothing here reads their prose and forms a view of it — that would need a model in the scoring path, which is the one thing this product will not do. What you get instead is their reasoning in full, with every figure in it marked green where it matches their own work and amber where it appears nowhere. Reading the argument takes you thirty seconds, and the judgement stays yours.

See a finished report A real one, scored by the same engine, and downloadable as a PDF. No sign-up, and there is no real person in it. It comes with a shorter version too — the six measures and what would move each of them, carrying no answers and no verdict — which is the one you forward to a manager, or send back to the candidate.

What stops them opening a second tab

The demo above has this switched off, because you should be able to check Slack in the middle of it. A real sitting does not.

  1. Under 2 seconds Nothing happens. A dropdown, a window manager settling, an alt-tab that bounced. Nobody reaches another application and back in two seconds, so it is not recorded as leaving.
  2. First departure A countdown appears on the page and starts running down from 2 minutes. They can see it from the other application. Coming back and pressing the button stops it.
  3. Second departure The same two minutes, and the page says the next one ends it. No ambiguity left about what happens next.
  4. Third departure The assessment ends. No countdown — it has been said twice already. The report says so on its front page rather than burying it.

The clock never stops. The deadline is wall-clock and enforced on the server, so time spent away is time spent. The two minutes are not extra thinking time, which is what removes the incentive to use them.

There is no grace period, and there used to be. Sixty seconds, so a screen lock would not cost anybody a strike. It protected the wrong person: a minute is long enough to paste the question into a language model, read the answer and come back, with the product saying nothing throughout.

No camera, no microphone, no screen recording. Those are denied by a header the browser enforces, not by a promise in a policy. The only thing observed is whether this page was in front. See Security.

Your assessment vendor's compliance problem is not yours

There is no language model anywhere in the scoring path. We check whether the figures a candidate quotes match the analyses they actually ran, and mark each one for you. Judging the argument stays with a person — yours.

That is a commercial position as much as a technical one. A model deciding part of the score makes a tool an automated employment decision system: high-risk under the EU AI Act, with mandatory risk assessment, bias testing, technical documentation, human oversight and continuous monitoring, and an automated employment decision tool under New York City Local Law 144. Those obligations land on the employer as well as the vendor.

Every number here is arithmetic you can inspect, and every judgement is made by a person. There is nothing to audit because there is no model to audit.

Not for New York City roles yet — Local Law 144 requires an independent bias audit, and we do not have one.

Teaching analysts rather than hiring them?

The same fifty minutes works as a course capstone. Your students find out where they actually stand before they start interviewing, which is the one thing a certificate cannot tell them.

Trial Day for courses

Why this exists

I have sat on both sides of this. I have been sent take-home case studies that ate a weekend and told the company almost nothing, and I have been the one reading them — trying to work out from a tidy final answer whether somebody understood the data or got lucky, with no way to see what they tried first.

The people who decline a long unpaid exercise first are the ones with competing offers, so the exercise quietly makes the shortlist worse. And an eight-hour task filters out anyone with caring responsibilities or a demanding current job, neither of which has anything to do with being good at the work.

So: fifty minutes, the real job, and every query they ran on the way to their answer. Judged by you, not by a model.

Sanyam Gulati, founder

Try it on a role you are hiring for

Tell us what you are hiring for and we will set you up, usually the same day. Bring your own analysts too — running your team through it first is free, and it is what makes every candidate result afterwards mean something.

No card, and nothing automated — a person reads these.

Already set up? Open the hiring console.