Updated October 2026
A startup data scientist interview works best in four stages: a problem-framing screen, a technical round on the flavour you are hiring for, a work sample presented to the founder, and a final round with references. They test, in turn, judgement, technical depth, communication of uncertainty and stage fit.
Choose the flavour first: analytics, decision science or machine learning. Then cut the questions that don't apply. A loop that tests all three rewards generalists who interview well and punishes the specialist you need.
The running example is an illustrative Seed-stage language-learning app hiring a decision science profile to work out why users drop off in week two and whether a new paywall helped.
The data scientist interview plan
| Stage | Who | Length | What it tests |
|---|---|---|---|
| 1. Framing screen | Founder | 30 min | Flavour, decisions their work changed |
| 2. Flavour deep dive | Senior data person or advisor | 60 min | Depth in your chosen flavour |
| 3. Work sample readout | Founder plus product lead | 60 min | Analysis and explaining uncertainty |
| 4. Final and references | Founder | 45 min | Stage fit; a reference from a stakeholder, not a peer |
Framing the problem (every flavour)
- Tell me about work of yours that changed a decision. What was the decision, and who made it?
- A founder asks why revenue dipped last month. What do you ask before opening a notebook?
- When did you last tell a stakeholder the data could not answer their question?
- What's the simplest approach you'd try before reaching for a model?
Strong answer: Names a decision and a decision-maker, asks about timing, segments and tracking changes first, and starts with a rule or a chart.
Red flags: Lists models built with no decision attached, or never says no to a question.
Analytics
- Weekly active users rose but revenue fell. Give me three explanations and how you would check each.
- How would you build a retention cohort table from our raw events?
- Which metric in our pitch deck would you challenge, and why?
Strong answer: Several competing explanations, each with a quick check, and a willingness to question your own numbers.
Red flags: One explanation defended to the end, or cohorts built without checking how events are logged.
Experimentation
- Walk me through testing a new paywall, from hypothesis to readout.
- How do you fix the run time before the test starts?
- Halfway through, the product lead wants to stop because it looks like a win. What do you say?
- We lack the traffic for a clean A/B test. What would you do instead?
- A test shows no difference. What can and can't we conclude?
Strong answer: Sets a primary metric and sample size in advance, resists peeking, offers alternatives such as before-and-after with a holdout, and knows no difference is not proof of no effect.
Red flags: "Run it until it is significant", or no plan for low traffic.
"Yes, it was statistically significant, so we should roll it out to everyone."
"Probably, for monthly plans. For annual plans the range still includes no effect. Roll out monthly now and run annual for two more weeks."
Getting models into production
- Tell me about a model that never reached production. Why not?
- What baseline would you compare a first recommendation model against?
- How do you know a live model has quietly got worse?
- Who owned deployment in your last team, and what did you hand over?
- When would you recommend not building a model at all?
Strong answer: An honest story of a shelved model and what they learned, a simple baseline such as most popular items, and monitoring of inputs as well as accuracy.
Red flags: Every model shipped, no baseline, or "engineering handled that" with no idea what happened next.
Communicating uncertainty
- Explain a confidence interval to a founder who must decide by Friday.
- Tell me about a recommendation of yours that turned out wrong. What did you say afterwards?
- The founder wants one number for the board. Your analysis gives a range. What goes on the slide?
Strong answer: Plain language, a recommendation despite the range, and ownership of past mistakes.
Red flags: Jargon, false precision, or refusing to recommend anything until the data is perfect.
Work sample: a paywall readout
Share synthetic results from a paywall test with a borderline overall result, one segment that moved and one tracking gap. Ask for a five-slide readout for the founder with a recommendation. Cap it at 3 hours. For the ML flavour, ask instead for a baseline and a monitoring plan, not a tuned model.
Look for three things: they question the data before the result, they state a range, and they still make a call. Present it in stage three to the founder, who is the real audience.
Scoring rubric
| Outcome | 1 | 2 | 3 | 4 |
|---|---|---|---|---|
| Decisions changed | None named | Vague influence | Clear decision | Several, with decision-makers named |
| Experiment quality | Peeks until significant | Basic design | Sized in advance | Handles low traffic well |
| Communicating uncertainty | Hides or ignores range | Jargon | Clear range | Range plus a call |
| Data foundations | Trusts all data | Spots issues if told | Checks tracking first | Found the planted gap |
| Models in production (ML only) | Never shipped | Shipped with help | Shipped and monitored | Knows when not to build |
Where Funded.club fits
We run data and engineering searches on a fixed fee agreed upfront, with one dedicated recruiter and first screened candidates within 7 days. See pricing.
Frequently asked questions
How do I interview a data scientist if I'm not technical?
Borrow an advisor for stage two and run stage three yourself. Explaining uncertainty to a non-specialist is the skill you are best placed to judge.
Should I hire for ML or for analytics?
Analytics, unless a model is part of your product. Our guide to hiring ML and AI engineers covers the ML route.
What should the job description say?
Name the flavour and the decisions the hire will change. Start from our data scientist job description template.
Worth a brief chat about your next hire? Book a free call.
Hiring after a funding round?
First candidates in 7 days. Fixed-fee, no commission.