Behavioral
16 questions
Tell me about a challenging ML project and what you learned.▶
Structure: STAR (Situation, Task, Action, Result)
What interviewers look for:
- Clear problem definition - did you understand the real problem before building?
- Data thinking - what were the data challenges? How did you handle distribution issues?
- Iteration mindset - what didn't work and why? How did you debug?
- Quantifiable impact - accuracy improved X%, latency reduced Y%
- Honest self-assessment - what would you do differently?
interview angle: lead with the CV production work. Pick one project - object detection or a model serving challenge - and tell it with specific technical details (what backbone, what loss, what data augmentation strategy, what evaluation metrics).
Avoid: vague descriptions, no numbers, blaming external factors, no reflection.
How do you stay current with ML research?frontier▶
Frontier labs want specifics, not "I follow arXiv."
Good answers include:
- arXiv daily: cs.LG, cs.CV, cs.CL - subscribe to feeds, use Semantic Scholar
- Papers one've read recently: be specific (name the paper, why it interested one, key insight)
- Twitter/X follows: @karpathy, @ylecun, @sama, @hardmaru, @ilyasut
- Blogs: Lilian Weng (lilianweng.github.io), Sebastian Raschka, The Gradient
- Podcasts: Dwarkesh Patel, Lex Fridman, Machine Learning Street Talk
- Code: re-implement interesting papers in PyTorch (this teaches more than just reading)
Tip: before the interview, read 2-3 recent papers relevant to the company's work. Have a concrete opinion on each. "I read the Flash Attention 3 paper last week and found the warp specialization insight fascinating because..." is a much stronger answer than "I follow arXiv."
A model worked great in testing but failed in production. Walk me through the debugging.▶
This is a classic ML system debugging question. Interviewers want systematic thinking.
Common causes + checks:
- Data leakage: did features in test set contain future information? Check feature timestamps vs label timestamps.
- Distribution mismatch: is production data distribution similar to test set? Compare feature distributions. Did anything change after data collection?
- Preprocessing bugs: is the production preprocessing pipeline identical to training? Missing normalization? Different tokenizer version?
- Batch norm inference mode: is model in eval mode? (Dropout + BatchNorm behave differently at inference)
- Feature drift: are the features being computed correctly? Log some production inputs and compare to training examples.
- Label delay: if labels come with delay, the validation might have been evaluated with delayed labels that aren't available in production.
Process: 1) Reproduce the failure offline with a production input. 2) Shadow mode test (log but don't serve production predictions). 3) Compare data distributions. 4) Check every step of the pipeline.
How do you approach a project with very limited labeled data?▶
- Transfer learning: fine-tune a pretrained model (CLIP, DINOv2 for CV; BERT/GPT for NLP). Often works with 100-1000 labeled examples.
- Self-supervised pretraining: if you have unlabeled data in-domain, pretrain on it (MAE, contrastive learning) before fine-tuning on the small labeled set.
- Data augmentation: for CV: random crop, flip, color jitter, mixup. For NLP: back-translation, synonym replacement, EDA.
- Active learning: query a model to label the most informative unlabeled examples first. Uncertainty sampling, core-set selection, query-by-committee.
- Semi-supervised learning: use unlabeled data with pseudo-labels (FixMatch, MixMatch). Train on labeled + predicted-label unlabeled examples.
- Weak supervision (Snorkel): write labeling functions (heuristics, regex, knowledge bases) to programmatically label unlabeled data. Combine with denoising model.
- Synthetic data: generate training examples with LLMs (text), or with GANs/diffusion (images).
How do you debug a model with a flat loss curve?▶
Systematic debugging process for training issues:
- Verify data pipeline first: print a batch. Is it the same every iteration (forgot to shuffle)? Are labels correct? Are inputs in expected range?
- Check learning rate: too low → flat loss. Too high → NaN or exploding loss. Try 10× higher, then 10× lower.
- Verify loss computation: is reduction correct (mean vs sum)? Is loss actually being backpropagated (no.detach accidentally on the right variable)?
- Check gradient flow: print gradient norms per layer. If all zeros, something is wrong in the graph. If all NaN, exploding gradients.
- Overfit a single batch: train on 1 batch for 100 steps. If loss doesn't reach near-zero, there's a fundamental bug. Loss should be -log(1/n_classes) ≈ 2.3 initially for 10-class.
- Check initialization: initial loss should be close to -log(1/n_classes). If very different, initialization or loss formula is wrong.
How do you explain AI limitations to a non-technical PM who wants 99% accuracy?startup▶
This is a communication + technical maturity question.
Framework:
- Ground with the baseline: "The current rule-based system achieves X%. No model achieves 99% on this kind of data."
- Explain irreducible error: "Some inputs are inherently ambiguous - even humans disagree. Our target can't exceed human-level performance on ambiguous cases."
- Reframe the metric: "99% overall accuracy may not be the right goal. If Type II errors (missed fraud) cost 10× more than Type I, we should optimize that, not accuracy."
- Set expectations with data: "SOTA for this task is 93%. We can realistically target 88-90% at launch."
interview angle: the experience teaching AI to seniors makes one well-suited for this. Draw on how you explain model limitations in InterviewReady.
How do you decide whether a problem needs ML at all?▶
Strong candidates know ML is not always the answer.
Don't use ML when:
- Rule-based logic handles it fully (route emails by subject line)
- Insufficient data (<1000 labeled examples for a complex task)
- Interpretability is legally required (regulated credit or medical decisions)
- Latency constraint makes inference infeasible (<1ms requirement)
- Training cost exceeds value (one-time task, simple pattern)
Use ML when: patterns are too complex for hand-crafted rules, sufficient labeled data exists, prediction errors are acceptable, the task generalizes across many inputs.
Rule: always baseline with a simple heuristic first. If the rule gets 90%, the ML bar is 95%+ with acceptable cost and complexity. The delta must justify the complexity.
Tell me about a time you improved model performance significantly.▶
Structure: what was the baseline, what did you try, what worked, what did you measure, what was the outcome?
Strong signals interviewers look for:
- You measured first (profiled, identified bottleneck) rather than guessing
- You tried multiple hypotheses systematically, not randomly
- You can explain WHY something worked, not just that it worked
- You have specific numbers (mAP went from 0.76 to 0.91, not "significantly improved")
the story: the 54% latency reduction at Beltech via NMS-free YOLOv10 + TensorRT is a strong story. Or 0.9+ F1 across 10+ detection modules. Pick one, tell it precisely with numbers.
How do you handle disagreements about model design with teammates?▶
Separate disagreements about opinions from disagreements about facts. Most ML disagreements are empirical - test both approaches.
Process:
- Understand their position first: "Help me understand why you prefer X" - collect reasoning, don't argue
- Agree on success criteria before debating: what metric determines who's right?
- Run both experiments if feasible: a 2-hour ablation study resolves more than a 2-hour debate
- Opinion only: defer to whoever has more relevant domain experience, or whoever will implement it
- Document the decision and reasoning for future revisit if results diverge
Avoid: ego-driven disagreements, HiPPO effect (Highest Paid Person's Opinion wins), deciding without data when data is available.
Describe the experience teaching AI. What's the hardest concept to explain?▶
Use concrete examples from InterviewReady (200+ engineers) and TensorTonic (30k+ sign-ups).
Hard concepts (pick one you actually struggled to teach):
- Backpropagation: the chain rule is easy; why we can compute all gradients in one backward pass is hard. Key insight: they are computing a product of Jacobians with memoization - O(forward pass).
- KV cache: most engineers understand caching. Specific confusion: why cache K and V but not Q? Q is the "question" for the current new token; K and V are history we reuse.
- RLHF vs DPO: engineers understand supervised learning, not RL. Insight: DPO replaces the RL loop with a direct supervised objective derived from the RLHF optimal policy - alignment without RL training.
interview angle: mention specific misconceptions students had. Shows pedagogical depth beyond just knowing the content.
What interests you specifically about this company's work? (Prepare per company.)frontier▶
Generic answers fail here. Prepare something specific enough it couldn't apply to any other company.
Per company:
- OpenAI: o3/o4 reasoning models, Whisper, Operator (agents). "I find the o3 approach - scaling verification rather than generation - more interesting than pure scale."
- Anthropic: safety + interpretability research. "The Superposition paper's feature visualization work made me curious about what circuits Claude learns for reasoning."
- Google DeepMind: AlphaFold, Gemini, robotics. "Gemini 1.5's million-token context and how they achieved it via ring attention."
- Meta AI: LLaMA open ecosystem, SAM, DINO. "Meta open-sourcing frontier models changes the entire research ecosystem - democratizes alignment research."
- Sarvam AI: "The challenge of high-quality ASR for code-switched speech across 22 Indian languages with limited labeled data."
- Krutrim: "Building India's sovereign AI stack - multilingual tokenization and culturally-grounded pretraining for Indian languages."
Where do you see the AI field heading in the next 3 years?frontier▶
Tests intellectual engagement with the field. Give a genuine opinion, then defend it.
Strong view areas (pick 1-2 and go deep):
- Reasoning and verification: the o1/o3 paradigm - scaling compute at inference via "thinking" separates capability tiers. Models that verify their own outputs become more reliable than those that don't.
- Agentic systems: LLMs will run entire software development workflows, not just assist. Bottleneck moves from model capability to reliable long-horizon tool use.
- On-device intelligence: 7B models on laptops and phones, enabled by quantization + mobile hardware. Privacy-preserving AI becomes practical at scale.
- Multimodal natively: models trained jointly on text, image, audio, video, code from scratch - not modality bolted on.
Be honest: "I don't know" is acceptable for specific timelines. A genuinely held view beats a diplomatic non-answer. Karpathy's insight: progress is fast but reliability and safety lag capability by years.
Tell me about a time you made a decision with incomplete data.▶
Tests comfort with uncertainty and decision-making under ambiguity - core skill at startups and research labs.
STAR for ML context:
- Situation: goal, timeline, why data was incomplete
- Task: what decision needed to be made? Stakes?
- Action: how did you quantify uncertainty? What assumptions did you make explicit? Did you prioritize getting more data vs. acting with what you had?
- Result: what happened? Was the decision vindicated? What would you do differently?
Key insight to demonstrate: you make uncertainty explicit (confidence intervals, sensitivity analysis) rather than pretending more confidence than the data supports. Show you can act despite uncertainty, not just analyze it.
How do you measure ROI of an AI feature?startup▶
Measures business acumen + ML engineering maturity. Pure ML engineers skip this; strong candidates think end-to-end.
Framework:
- Baseline cost: what does the non-AI solution cost? (human labor, rule engine, or no solution at all)
- AI solution cost: compute (inference cost), human review, maintenance overhead, annotation cost
- Value: what metric does the AI feature move?
- Fraud detection: TPR improvement × avg fraud value × transactions = annual savings
- Recommendation: A/B test CTR lift × monetization per click × traffic = revenue
- Automation: tasks/hour × human cost/hour = labor savings
- Payback period: when does cumulative value exceed development + operating cost?
Gotcha: include model maintenance and retraining costs. Most ROI calculations omit this and underestimate total cost.
A PM wants 90% accuracy but the model achieves 82%. What do you do?startup▶
This is about requirements negotiation, not just model improvement.
Step 1: understand where 90% came from. Business requirement (need to replace current process)? Or round number someone invented? Very different problems.
Step 2: challenge the metric. Is accuracy right? With class imbalance, 82% accuracy might be 0.3 F1. 82% at high precision might beat 90% at lower precision depending on the cost structure.
Step 3: price the gap. "Going from 82% to 90% requires 10× more labeled data, 3 months, possibly a different architecture, and may not be achievable given label noise. Here's what 85% costs for 2 weeks."
Step 4: propose alternatives. Confidence thresholding: achieve 92% accuracy on 70% of cases (high-confidence), route remaining 30% to humans. System-level accuracy meets the 90% bar.
The right answer is not "I'll improve the model." It's a business/engineering conversation first.
Tell me about a time you pushed back on a technical direction and were wrong. What did you learn?▶
This question tests intellectual honesty and how you update beliefs under evidence - arguably more important in ML than in most engineering domains because ML involves constant uncertainty.
What interviewers want to hear:
- You had a clear technical opinion and stated it (not passive)
- You were genuinely wrong, not just deferring to authority
- You updated the view based on evidence (data, results, a colleague's reasoning)
- You extracted a durable lesson, not just "I was wrong once"
Structure: situation → the position → counter-evidence you encountered → how you changed course → what you now look for before pushing back on similar decisions.
Trap to avoid: picking a story where you were "wrong" but actually right, or where the stakes were trivially low. The question is a calibration check - interviewers (especially at Anthropic, DeepMind) will probe whether they are genuinely intellectually honest or performing humility.