Everything they can
reasonably ask, derived.
Not a flashcard dump. Each idea is built from the assumption underneath it, because the follow-up question is always "why does that work?" — and the answer to that is where offers are decided.
Weighted for your profile: classical ML is developed in full from first principles, deep learning and LLM sections run past standard interview depth since your paper invites it.
The night-before plan
You cannot learn ML tonight. You can make what you already know retrievable under stress, and close the two or three gaps most likely to be probed.
| Block | Do this | Why |
|---|---|---|
| Hours 1–2 | §02, §03, §04 — bias-variance, regularization, metrics, logistic regression derivation | These appear in almost every ML screen and they're the part your coursework skipped |
| Hour 3 | §05 ensembles + §06 PCA/k-means/EM | Classic "explain gradient boosting" and "derive PCA" territory |
| Hour 4 | §07–§08 backprop by hand, optimizers, normalization | Whiteboard favourites; you know these — this is retrieval practice |
| Hour 5 | §10–§11 attention, KV cache, decoding, RLHF/DPO, calibration | Your resume promises this. They will push until you're at the edge |
| Hour 6 | §15 — say your paper's contribution and its weakness out loud, three times | The single highest-leverage 30 minutes on this page |
| Morning | §14 rapid-fire + §16 checklist only | Warm the cache; don't open new material |
Answer in three beats: (1) the one-line definition, (2) the mechanism or derivation, (3) when it breaks. Beat 3 is what separates a candidate who read a blog post from one who has trained models. Almost nobody volunteers it. You should, every single time.
Never bluff — interviewers detect it instantly and it retroactively discounts everything you said before. Say: "I haven't worked with that directly. My guess from first principles is X, because Y — is that the right intuition?" This scores nearly as well as knowing, because it demonstrates the thing they're actually testing.
Math you'll actually be asked
Interviewers don't test linear algebra for its own sake. They test the four or five results that keep reappearing inside models — and they test whether you can connect the symbol to the thing.
Linear algebra: the five facts
1. A matrix is a function. y = Ax maps ℝⁿ → ℝᵐ. Rank = dimension of the output space actually reachable. If rank < n, the map destroys information and is not invertible — this is exactly why collinear features break ordinary least squares.
2. Eigen-decomposition. For symmetric A, A = QΛQᵀ with Q orthonormal. Eigenvectors are directions the matrix only stretches; eigenvalues are the stretch. Covariance matrices are symmetric PSD, so they always decompose this way — that is the whole of PCA.
3. SVD generalises this to any matrix: A = UΣVᵀ. Always exists. Truncating to the top-k singular values gives the best rank-k approximation in Frobenius norm (Eckart–Young). This single fact underlies PCA, LSA, recommender factorization, and LoRA.
4. Norms. L2 = √Σxᵢ², smooth, penalises large values quadratically. L1 = Σ|xᵢ|, non-differentiable at 0, and that kink is precisely why it produces exact zeros. L∞ = max|xᵢ|.
5. Gradients of the things you'll write on a whiteboard:
Probability: the parts that bite
- Expectation / variance: 𝔼[aX+b] = a𝔼[X]+b; Var(aX+b) = a²Var(X). Var(X) = 𝔼[X²] − 𝔼[X]².
- Independence vs uncorrelated: independent ⟹ uncorrelated; the converse fails (X ~ N(0,1), Y = X² has zero correlation, total dependence). Interviewers love this one.
- CLT: the mean of n i.i.d. samples → Normal with variance σ²/n, whatever the source distribution. This is why standard errors shrink as 1/√n, and why you need 4× the data to halve a confidence interval.
- Covariance matrix: Σ = 𝔼[(x−μ)(x−μ)ᵀ]. Diagonal = variances, off-diagonal = covariances. Always symmetric PSD.
Information theory (you need this for losses)
Why is squared error the "right" loss for regression and cross-entropy the "right" loss for classification?
Both fall out of maximum likelihood under different noise assumptions.
Assume y = f(x) + ε with ε ~ N(0, σ²). The log-likelihood of the data is −Σ(yᵢ − f(xᵢ))²/2σ² + const. Maximising it is identical to minimising squared error. MSE assumes Gaussian noise.
For classification, assume y ~ Bernoulli(p) with p = σ(f(x)). Log-likelihood is Σ[y log p + (1−y) log(1−p)]. Maximising it is minimising binary cross-entropy. Cross-entropy assumes a categorical/Bernoulli likelihood.
Beat 3 — when it breaks: if your noise is heavy-tailed, the Gaussian assumption makes MSE wildly outlier-sensitive; switch to Huber or MAE (which is MLE under a Laplace noise model). And if you use MSE on a sigmoid output, the gradient contains a σ′(z) factor that vanishes when the model is confidently wrong — cross-entropy cancels it, which is why it trains far faster.
Derive why an L2 penalty is equivalent to a Gaussian prior on the weights.
MAP estimation maximises log P(D|w) + log P(w). Put a prior w ~ N(0, τ²I). Then
Maximising the sum is minimising −log P(D|w) + ‖w‖²/(2τ²), i.e. the data loss plus λ‖w‖² with λ = 1/(2τ²). A tighter prior (small τ) means stronger regularization. Substituting a Laplace prior P(w) ∝ exp(−|w|/b) gives the L1 penalty identically.
Learning theory & generalization
The core question of ML: why should performance on data you've seen say anything about data you haven't? Every regularizer, every validation scheme, every architectural prior is an answer to it.
The setup, stated properly
There's an unknown joint distribution P(x, y). You want the hypothesis f minimising risk, R(f) = 𝔼₍ₓ,ᵧ₎~P [L(f(x), y)]. You can't compute it, so you minimise empirical risk over your n samples, R̂(f) = (1/n)Σ L(f(xᵢ), yᵢ). Everything interesting is in the gap R(f) − R̂(f).
That gap grows with how expressive your hypothesis class is (it can chase noise) and shrinks with n. Formally VC dimension / Rademacher complexity bound it; practically, you control it with regularization, data, and inductive bias.
Bias–variance, derived
For squared error, with y = f(x) + ε, Var(ε) = σ², and f̂ trained on a random dataset:
The U-curve is the classical picture. Empirically, as you keep growing capacity past the interpolation threshold (where the model fits training data exactly), test error rises, peaks, then falls again. Over-parameterised networks generalise well because SGD has an implicit bias toward low-norm / flat-minimum solutions among the infinitely many that interpolate. Mentioning this signals you've read past the textbook — but state the classical decomposition first.
Regularization: five distinct mechanisms
| Method | Mechanism | Effect on solution |
|---|---|---|
| L2 / weight decay | Penalises ‖w‖². Gradient adds 2λw, shrinking weights each step | Shrinks all weights smoothly toward 0, never exactly 0. Handles collinearity by splitting weight between correlated features |
| L1 / Lasso | Penalises Σ|w|. Constant-magnitude gradient regardless of size | Drives weights to exactly zero → feature selection. Among correlated features it arbitrarily picks one |
| Elastic net | αL1 + (1−α)L2 | Sparsity from L1 plus grouped selection from L2 |
| Early stopping | Halt before the model fits noise | Provably approximates L2 for linear models — the effective norm grows with training steps |
| Dropout / augmentation / noise | Injects stochasticity so no feature can be relied on | Approximate ensembling; forces redundant, distributed representations |
Why does L1 give exact zeros but L2 doesn't? Give the geometric and the calculus argument.
Geometric: constrained optimisation. The L2 constraint region is a ball (round, no corners); the L1 region is a diamond with vertices on the axes. The elliptical contours of the squared-error loss expand until they touch the constraint region. A round ball is almost surely touched at a non-axis point; a diamond is most likely touched at a vertex — and a vertex has coordinates that are exactly zero.
Calculus: look at the gradient of the penalty near zero. For L2 it's 2λw → 0 as w → 0 — the shrinkage force vanishes exactly where you'd need it to finish the job, so weights asymptote to zero without arriving. For L1 the subgradient is λ·sign(w), a constant push that does not weaken as w shrinks. So as long as the data's gradient is smaller than λ, the weight is pinned at exactly zero. That's the soft-thresholding operator: w ← sign(w)·max(|w| − λ, 0).
Your model gets 99% train accuracy and 71% test accuracy. Walk me through what you do.
Diagnose before treating. In order:
- Confirm it's variance, not a broken split. Check the test set has the same class balance and feature distributions as train. A distribution shift looks identical to overfitting from the metrics alone but has the opposite fix.
- Check for leakage in reverse: 99% train is suspiciously high. Is a feature a proxy for the label? Did I fit the scaler/imputer before splitting?
- Plot a learning curve — test error vs training set size. If the curves are still converging, more data helps and I should get it before touching the model. If they've plateaued with a wide gap, it's capacity/regularization.
- Then, cheapest interventions first: stronger regularization (increase λ, dropout, weight decay), early stopping on a validation split, data augmentation, then reduce capacity, then simpler model class.
- Re-examine features — high-cardinality categoricals with rare levels are a classic overfit source; target-encoding them without out-of-fold encoding will do exactly this.
The point they're testing is whether you jump straight to "add dropout" or actually diagnose.
Validation done right
- k-fold CV: split into k folds, train on k−1, validate on 1, rotate, average. Lower-variance estimate than one holdout. k=5 or 10 by convention; k=n is leave-one-out (nearly unbiased, high variance, expensive).
- Stratified k-fold preserves class ratios per fold — mandatory for imbalanced data.
- Time-series: never shuffle. Use forward-chaining (train on [0,t], validate on [t,t+Δ], roll forward). Random CV on temporal data leaks the future and is the single most common résumé-project mistake.
- Grouped CV: if you have multiple rows per patient/user/session, all rows for a group must be in the same fold, or the model memorises the group.
- Nested CV: when you tune hyperparameters and want an unbiased performance estimate, you need an inner loop for tuning and an outer loop for evaluation. Reporting the best inner-loop score as your final number is optimistically biased — a subtle point that impresses.
Evaluation & metrics
Choosing a metric is choosing what kind of mistake you're willing to make. Interviewers use metric questions to check whether you think about the product, not just the model.
ROC-AUC vs PR-AUC — the distinction they're fishing for
ROC plots TPR against FPR as you sweep the threshold. AUC = probability that a randomly chosen positive is ranked above a randomly chosen negative. It's threshold-free and invariant to class balance.
That invariance is both its strength and its flaw. FPR has TN in the denominator, and with 99.9% negatives, TN is enormous — so hundreds of false positives barely move FPR. ROC-AUC can look excellent while your fraud alerts are 95% noise.
PR curve plots precision against recall. Precision has no TN term, so it feels every false positive. On heavily imbalanced problems where you care about the positive class, PR-AUC is the honest metric. Note the PR baseline is the positive rate (0.001), not 0.5.
Calibration — your home turf
A model is calibrated if, among all predictions with confidence 0.8, exactly 80% are correct. Accuracy and calibration are independent: a model can rank perfectly (AUC 1.0) and be catastrophically overconfident.
- Fixes: temperature scaling (single parameter T fitted on a held-out set, divides logits — preserves ranking so accuracy is unchanged), Platt scaling (logistic fit), isotonic regression (non-parametric, more flexible, needs more data, can overfit).
- Why it matters: anything that thresholds on confidence — abstention, routing, early-exit, self-consistency stopping, cascades — is only as good as the calibration underneath.
- Critical: the temperature must be fitted on data disjoint from both training and test. This is the leak-free requirement your own work is about; expect to be asked why it matters.
Fraud detection, 0.2% positive rate. Which metric, which threshold, and why?
Metric: PR-AUC for model selection (ROC-AUC will be near 0.99 for everything and won't discriminate), plus recall at a fixed precision or precision at a fixed alert budget as the operating metric — because the review team can only handle k alerts a day, and that's the real constraint.
Threshold: derived from cost, not from 0.5. If a missed fraud costs C_FN and a false alarm costs C_FP (analyst time plus customer friction), the optimal threshold on a calibrated probability is t* = C_FP / (C_FP + C_FN). With C_FN = 500 and C_FP = 5, t* ≈ 0.01 — flag anything above 1%.
Then: I'd also monitor precision over time, because fraud is adversarial and non-stationary — the distribution changes in response to your model, so I'd retrain on a schedule and alert on drift. And I'd never evaluate with random CV; fraud has temporal structure and the split must be forward in time.
Is a model with 90% accuracy and terrible calibration worse than one with 85% accuracy and good calibration?
Depends entirely on how the output is consumed, and saying so is the answer.
If the model's argmax is the final decision and every error costs the same, take the 90%. Calibration is irrelevant to argmax.
If anything downstream reads the probability — expected-value calculations, ranked triage, abstain-and-escalate, cascading to a bigger model, combining with another signal — the 85% calibrated model is strictly more useful, because the 90% model's confidences are lies and every threshold you set on them is arbitrary.
Best answer ends here: you usually don't have to choose. Temperature scaling fixes calibration without changing accuracy at all, since dividing all logits by a positive scalar is monotone and preserves the argmax. So take the 90% model and calibrate it on held-out data.
Classical supervised models
Each of these is a set of assumptions plus an optimisation. Know the assumptions and you can answer any follow-up; memorise the algorithm and you can't.
Linear regression
Model: ŷ = Xw. Loss: ‖Xw − y‖². Set the gradient 2Xᵀ(Xw − y) to zero:
Assumptions: linearity in parameters, independent errors, homoscedasticity (constant error variance), no perfect collinearity, and for inference, normally distributed errors. Complexity: O(np² + p³) closed-form — so use gradient descent when p is large.
Logistic regression — derive this tonight
It's a linear model of the log-odds:
Likelihood over the dataset, taking y ∈ {0,1}: ∏ pᵢ^yᵢ (1−pᵢ)^(1−yᵢ). Negative log-likelihood:
Derivation detail worth knowing: σ′(z) = σ(z)(1 − σ(z)). In the chain rule, ∂L/∂p · ∂p/∂z, the 1/(p(1−p)) from the log-loss cancels the p(1−p) from the sigmoid, leaving (p − y). That cancellation is why cross-entropy avoids the vanishing-gradient problem that MSE-plus-sigmoid has.
No closed form — the NLL is convex but transcendental, so you use gradient descent or Newton's method (IRLS). Convexity means one global optimum. On perfectly separable data the weights diverge to infinity to push σ toward exactly 0 or 1; regularization is what keeps it finite, and that's a great detail to volunteer.
p_k = e^(z_k) / Σⱼ e^(z_j). Gradient of cross-entropy w.r.t. logits is again p − y (one-hot). Softmax is shift-invariant — subtracting max(z) changes nothing mathematically but prevents overflow, which is the standard numerical-stability question. Note softmax has one redundant degree of freedom; with K classes you only need K−1 independent logits.
Support vector machines
Idea: among all separating hyperplanes, pick the one with the largest margin, because a wide margin is a more robust decision boundary.
The dual replaces w with Σ αᵢ yᵢ xᵢ and the objective depends on data only through inner products xᵢᵀxⱼ. That's the opening for the kernel trick.
Support vectors are the points with αᵢ > 0 — only those on or inside the margin. The solution depends on nothing else, which makes SVMs sparse in examples (unlike Lasso, sparse in features) and robust to distant outliers. Scaling is mandatory — RBF's distance metric is meaningless with mismatched feature scales.
Cost: training is roughly O(n²)–O(n³), so SVMs don't scale past ~100k rows. That's the honest reason they lost to boosted trees and neural nets in practice.
Decision trees
Greedily choose the (feature, threshold) split that most reduces impurity, recurse, stop on a depth/leaf-size/gain criterion.
- Strengths: no scaling needed, handles mixed types, captures interactions automatically, interpretable, invariant to monotone feature transforms.
- Weaknesses: extremely high variance — change a few rows and the top split flips, restructuring the whole tree. Greedy and myopic (can't see a split that only pays off two levels down). Axis-aligned, so diagonal boundaries need staircases. Cannot extrapolate beyond the training range — a regression tree predicts a constant outside it.
- Bias in split selection: impurity-based importance favours high-cardinality and continuous features, because they offer more thresholds to try. Use permutation importance or SHAP instead when it matters.
k-NN and Naive Bayes
Zero training, all cost at inference — the definition of a non-parametric lazy learner. Small k = low bias, high variance, jagged boundary. Large k = smoother, higher bias. Must scale features. Suffers badly from the curse of dimensionality: in high dimensions, the ratio of nearest to farthest distance approaches 1, so "nearest" stops meaning anything. Inference is O(nd) per query unless you use KD-trees (which themselves degrade past ~20 dimensions) or approximate methods like HNSW/IVF-PQ — which is exactly what your vector database does.
Applies Bayes with the "naive" assumption that features are conditionally independent given the class: P(x|y) = ∏ P(xⱼ|y). The assumption is nearly always false, yet it works well for text, because you only need the argmax to be right, not the probabilities. Needs Laplace/add-α smoothing so an unseen word doesn't zero out the entire product. Famously badly calibrated — the independence assumption multiplies correlated evidence repeatedly, pushing probabilities to 0 or 1.
Why can't you use linear regression for classification?
Four reasons, in increasing sophistication:
- Outputs aren't bounded to [0,1], so they aren't probabilities.
- Squared error penalises confidently-correct predictions. A point with label 1 predicted at 3.0 incurs loss 4, and the model drags the boundary to fix a "mistake" that isn't one. So a single far-away point can shift the decision boundary and cause misclassifications — the classic figure.
- The Gaussian noise assumption behind MSE is wrong for a binary target; the residuals are heteroscedastic by construction (variance p(1−p) depends on the mean).
- With sigmoid + MSE, the gradient includes σ′(z), which is ≈0 when the model is confidently wrong — so the worst-off examples produce the smallest updates and learning stalls. Cross-entropy's log cancels that term exactly.
Generative vs discriminative — define it and say when each wins.
Discriminative models learn P(y|x) directly (logistic regression, SVM, neural nets). Generative models learn P(x|y) and P(y), then use Bayes to get P(y|x) (Naive Bayes, GDA/LDA, GMMs, and in the modern sense, LLMs modelling P(text)).
The Ng & Jordan result: the generative model (Naive Bayes) has higher asymptotic error but converges to its ceiling in O(log p) samples, versus O(p) for logistic regression. So generative wins in the small-data regime and discriminative wins once you have enough data. That's the crisp answer.
Generative models also let you sample new data, handle missing features naturally by marginalising, and detect out-of-distribution inputs via low P(x) — none of which a discriminative model gives you.
Ensembles & boosting
The most reliable performance gain in classical ML, and the most common source of "explain the difference" questions. Anchor everything to the bias–variance decomposition and you can't get lost.
Bagging attacks variance
Train M models on bootstrap resamples (sample n rows with replacement), average their predictions. If the models were independent with variance σ², the average has variance σ²/M. They aren't independent — they share data — so the real result is:
Random forest = bagging + a second decorrelation trick: at each split, consider only a random subset of m features (typically √p for classification, p/3 for regression). This stops every tree from choosing the same dominant feature at the root. Bias goes up slightly; ρ drops a lot; variance drops more.
Out-of-bag estimate: each bootstrap sample omits about (1 − 1/n)ⁿ → 1/e ≈ 36.8% of rows. Predict each row using only the trees that didn't see it — a free, nearly unbiased validation estimate with no separate holdout. Deriving the 36.8% is a favourite question.
Boosting attacks bias
Train models sequentially, each one correcting the previous ensemble's errors. Base learners are deliberately weak and high-bias (shallow trees, "stumps" or depth 3–6).
| Bagging / RF | Boosting / GBM | |
|---|---|---|
| Trained | In parallel, independently | Sequentially, each depends on the last |
| Base learner | Deep, low-bias, high-variance trees | Shallow, high-bias, low-variance stumps |
| Reduces | Variance | Bias (and variance, via shrinkage) |
| Overfits with more trees? | No — plateaus, adding trees is safe | Yes — needs early stopping on a validation set |
| Outliers / noisy labels | Robust — averaging dilutes them | Sensitive — repeatedly focuses on the hardest (possibly mislabelled) points |
| Tuning | Forgiving; defaults usually fine | Fussy; learning rate × n_estimators must be traded off |
XGBoost / LightGBM / CatBoost — what each actually adds
- XGBoost: second-order approximation of the loss (uses gradient and Hessian, so the leaf weight is
−G/(H+λ)), explicit regularization on leaf count and leaf weights in the objective, sparsity-aware split finding with a learned default direction for missing values, approximate quantile sketch for candidate thresholds, and column subsampling. - LightGBM: histogram binning of features (O(#bins) instead of O(#rows) per split), leaf-wise growth instead of level-wise (grows the leaf with max loss reduction — faster convergence, more overfitting risk, control with num_leaves), GOSS sampling, and EFB feature bundling. Usually the fastest on large data.
- CatBoost: ordered target statistics for categorical features that avoid target leakage, and ordered boosting to correct the prediction-shift bias. Best out-of-the-box on categorical-heavy tabular data.
"On tabular data, gradient-boosted trees still generally beat neural networks — the 2022 Grinsztajn et al. result. Trees are invariant to monotone feature transforms, robust to uninformative features, and handle irregular non-smooth target functions that MLPs' smoothness bias fights against. I'd reach for a neural net on tabular data mainly for multi-modal inputs or when I need embeddings for transfer."
Explain AdaBoost's weighting scheme.
AdaBoost is gradient boosting with exponential loss L = e^(−y·F(x)), but the original formulation is in terms of reweighting:
Read off the behaviour: if εₘ = 0.5 the classifier is useless and αₘ = 0, it gets no vote. If εₘ < 0.5, αₘ > 0. If εₘ > 0.5, αₘ is negative — the ensemble uses its inverted predictions. Misclassified points get their weight multiplied by e^(αₘ) so the next learner concentrates on them.
Weakness worth volunteering: exponential loss grows without bound for confidently-wrong points, so a mislabelled example gets exponentially increasing weight and can dominate the ensemble. That's why AdaBoost is notably noise-sensitive and why log-loss GBMs are preferred in practice.
Bagging reduces variance. Prove the 36.8% out-of-bag figure.
Drawing n items with replacement from n items: on any single draw, the probability a specific row is not chosen is (1 − 1/n). Across n independent draws, the probability it's never chosen is (1 − 1/n)ⁿ. Take the limit:
So each bootstrap sample contains about 63.2% unique rows, and 36.8% are out-of-bag. Averaging OOB predictions across all trees where a row was excluded gives a validation estimate for free — practically equivalent to leave-one-out CV, at no extra training cost.
What's stacking, and why is it easy to get wrong?
Stacking trains a meta-learner on the predictions of several base models, so it can learn which model to trust in which region of the input space — unlike simple averaging, which weights them uniformly.
The failure mode: if you train the meta-learner on predictions the base models made on their own training data, those predictions are optimistically good (the base models have partly memorised), so the meta-learner learns to trust an overfit signal and the ensemble generalises worse than its parts. The fix is out-of-fold predictions: for each fold, train base models on the other folds and predict on the held-out one; assemble those into the meta-features. Keep the meta-learner simple (regularized logistic/linear) — it has very few effective samples.
Unsupervised & dimensionality reduction
Less commonly asked than supervised learning, which is exactly why a confident PCA derivation or a clean EM explanation stands out.
k-means
Minimise within-cluster sum of squares: J = Σₖ Σ_{x∈Cₖ} ‖x − μₖ‖². Alternate two steps:
- Assign each point to its nearest centroid (holding μ fixed, this is the optimal assignment).
- Update each centroid to the mean of its points (holding assignments fixed, the mean provably minimises squared distance).
Each step is non-increasing in J and there are finitely many assignments, so it converges — but to a local minimum, depending on initialisation. Hence k-means++: pick the first centroid at random, then pick each subsequent one with probability proportional to D(x)² (squared distance to the nearest existing centroid), which spreads the seeds out and gives an O(log k) approximation guarantee in expectation.
- Assumes spherical, similar-sized, similar-density clusters — because Euclidean distance to a centroid defines a sphere. It will slice a crescent moon or two elongated clusters straight down the middle.
- Choosing k: elbow on the inertia curve (subjective), silhouette score (mean of
(b−a)/max(a,b), ranges −1 to 1, higher is better), gap statistic, or a downstream task metric. Say the elbow is a heuristic, not a criterion. - Complexity O(nkdi). Scale features first — otherwise the largest-range feature dominates the distance.
GMM and the EM algorithm
Model the data as a mixture: p(x) = Σₖ πₖ N(x | μₖ, Σₖ). Unlike k-means, assignments are soft (responsibilities), and clusters can be elliptical and differently sized because each has its own covariance.
PCA — derive it two ways
Setup: centre the data (subtract the mean — non-negotiable, or the first component just points at the mean). Compute Σ = (1/n)XᵀX.
Derivation 1 — maximum variance. Find the unit vector w maximising the variance of the projection, wᵀΣw subject to wᵀw = 1. Lagrangian: wᵀΣw − λ(wᵀw − 1). Differentiate, set to zero:
Derivation 2 — minimum reconstruction error. Find the k-dimensional subspace minimising Σ‖xᵢ − x̂ᵢ‖². Because total variance is fixed, minimising the residual is the same as maximising the retained variance. The two objectives are equivalent — say this, it shows real understanding.
In practice use the SVD of the centred X directly (X = UΣVᵀ, principal directions are columns of V, and forming XᵀX squares the condition number and loses precision).
- Components are orthogonal and ordered by variance. Standardise features first if units differ, otherwise PCA just finds the largest-variance unit.
- Limits: linear only; variance ≠ discriminative information (the class-separating direction can be a low-variance one — that's what LDA optimises instead); components are dense combinations, so interpretability suffers.
Preserves local neighbourhoods by matching pairwise similarity distributions and minimising KL. Uses a heavy-tailed t-distribution in the low-dim space to solve the crowding problem. Read distances and cluster sizes as meaningless — only local grouping is trustworthy. Non-parametric, so it can't embed new points. Perplexity ≈ effective neighbour count.
Similar goal via a fuzzy topological construction; faster, preserves more global structure, and supports transforming new points. Still: don't over-read the geometry. Both are visualisation tools, not preprocessing — don't feed t-SNE output into a classifier.
Your clusters look terrible with k-means. Give me three reasons and three alternatives.
Reasons: (1) the true clusters aren't spherical or equally sized, violating the model; (2) features weren't standardised, so one high-variance feature dominates the Euclidean distance; (3) high dimensionality — distances concentrate and every point is roughly equidistant from every other, so the objective becomes flat and uninformative.
Alternatives: GMM if clusters are elliptical or overlapping; DBSCAN if they're arbitrarily shaped and there's noise — it defines clusters by density (eps, min_samples), finds k itself, and labels outliers, but struggles when densities vary; HDBSCAN for varying density; hierarchical/agglomerative if you want a dendrogram and no fixed k, with Ward linkage minimising within-cluster variance. Also consider clustering in a learned embedding space rather than raw features.
What is the curse of dimensionality, concretely?
Several distinct effects, and naming more than one is the point:
- Sample sparsity: to keep a fixed density you need samples exponential in d. Covering [0,1]^10 at 10 points per axis takes 10¹⁰ samples.
- Distance concentration: as d grows,
(max_dist − min_dist)/min_dist → 0. Nearest-neighbour queries lose meaning, which breaks kNN, k-means and RBF kernels. - Volume moves to the shell: almost all the volume of a high-dimensional ball is in a thin shell near its surface, so "typical" points are all far from the centre and from each other.
- Everything is nearly orthogonal: random vectors in high d have inner product ≈0 — which is, incidentally, why superposition works in LLM residual streams and why you can pack far more features than dimensions.
Escape hatch: real data lies on a much lower-dimensional manifold. That manifold hypothesis is why deep learning works at all despite nominally enormous input dimensions.
Neural nets from scratch
If they hand you a marker, this is what they hand it to you for. Be able to do the two-layer backprop with shapes, out loud, without hesitating.
Forward pass
Backprop, term by term
Why the softmax + cross-entropy gradient is just (ŷ − y): the Jacobian of softmax is ∂pᵢ/∂zⱼ = pᵢ(δᵢⱼ − pⱼ). Multiply by ∂L/∂pᵢ = −yᵢ/pᵢ and sum over i; the pᵢ cancel and, using Σy = 1, you're left with p − y. Libraries fuse these two ops for exactly this reason (numerical stability plus the clean gradient).
A network is a computation graph. Forward-mode differentiation costs one pass per input dimension; reverse-mode costs one pass per output dimension. Loss functions have one scalar output and millions of inputs, so reverse mode gives you all parameter gradients in a single backward pass at roughly the cost of the forward pass. That asymmetry is the whole reason deep learning is computationally feasible.
Activations — and why each was invented
| Function | Form / derivative | The problem it has or solves |
|---|---|---|
| Sigmoid | 1/(1+e⁻ᶻ); σ(1−σ), max 0.25 | Saturates both ends → vanishing gradients. Not zero-centred, so all gradients into a weight share a sign, causing zig-zag updates. Keep it for output layers only |
| tanh | zero-centred, derivative 1−tanh², max 1 | Fixes the centring; still saturates. Standard inside LSTMs |
| ReLU | max(0,z); derivative 1 or 0 | No saturation for z>0, trivially cheap, induces sparsity. Dying ReLU: a unit pushed permanently negative has zero gradient forever and never recovers |
| LeakyReLU / PReLU | αz for z<0 (α≈0.01 or learned) | Keeps a gradient path alive on the negative side |
| GELU | z·Φ(z), a smooth gate by the input's percentile | Smooth and non-monotone near zero; the default in transformers (BERT, GPT) |
| SwiGLU | (xW₁ ⊙ σ(xW₂))W₃ — a gated variant | Consistently better perplexity per FLOP; used in LLaMA, PaLM. Note it needs 3 matrices, so hidden dim is scaled to ⅔ to match parameter count |
Initialisation — the failure they'll ask about
Initialise all weights to zero and every neuron in a layer computes the same thing, receives the same gradient, and stays identical forever — symmetry is never broken and the layer has an effective width of 1. Initialise too large and activations explode; too small and the signal dies out over depth.
Derive the vanishing gradient problem.
The gradient at layer l is a product of Jacobians along the path back from the loss:
This is a product of L−l terms. If each term has magnitude consistently below 1, the product decays exponentially in depth. With sigmoid, g′ ≤ 0.25, so even with perfectly scaled weights you lose at least a factor of 4 per layer — after 10 layers that's ~10⁻⁶. Early layers stop learning while late layers train fine. The mirror case, terms >1, gives exploding gradients, which shows up as NaN losses and is handled by gradient clipping.
The fixes, and which mechanism each uses: non-saturating activations (ReLU family, g′ = 1 on the positive side); careful initialisation (keeps the W terms near unit scale); normalization layers (rescale activations so the Jacobian doesn't drift); and most importantly residual connections — with x + F(x), the Jacobian is I + ∂F/∂x, so the identity term guarantees a gradient path of magnitude 1 straight back to any earlier layer regardless of depth. That's why 100+ layer networks became trainable.
What does the universal approximation theorem actually say — and not say?
Says: a feed-forward network with a single hidden layer and a non-polynomial activation can approximate any continuous function on a compact set to arbitrary accuracy, given enough hidden units.
Does not say: how many units (the width may be exponential in the input dimension); that you can find those weights (it's an existence result, silent on optimisation); or anything about generalisation to unseen data. It's also why the theorem doesn't justify shallow networks — depth buys you exponentially more efficient representations of compositional functions, which is the actual reason deep beats wide in practice.
Training deep nets
Optimisers, normalization, and the debugging story. The debugging question — "your loss isn't going down, what do you do" — is asked in nearly every ML interview and rewards a systematic answer.
Optimisers, built up in order
AdamW — the one to name for transformers. In Adam, an L2 penalty added to the loss gets divided by √v̂ along with everything else, so parameters with large historical gradients get less decay — the opposite of what regularization should do. AdamW decouples it, applying w ← w − ηλw directly, separate from the adaptive step. This measurably improves generalisation and is the default for every modern LLM.
Why SGD+momentum still wins sometimes: on vision tasks it often generalises slightly better than Adam. The common explanation is that Adam's per-parameter scaling steers toward sharper minima; SGD's isotropic noise biases toward flatter ones, which generalise better.
Learning rate — the single most important hyperparameter
- Warmup: start near zero and ramp up over the first few thousand steps. Essential for transformers — at initialisation, Adam's variance estimates are unreliable and large early steps destabilise LayerNorm statistics and attention patterns permanently.
- Cosine decay from the peak to ~10% of it is the standard LLM schedule. Also: step decay, linear decay, one-cycle.
- LR range test: exponentially increase the LR over a few hundred steps and plot loss; pick roughly an order of magnitude below where it diverges.
- Batch size interaction: larger batches give lower-variance gradients, permitting a larger LR. Linear scaling rule: multiply batch by k, multiply LR by k (with warmup). The √k rule is the alternative and often more stable for Adam.
Normalization
Why transformers use LayerNorm, not BatchNorm: sequences have variable length and the batch statistics of a padded token position are meaningless; BatchNorm makes each example's output depend on the others in its batch, which breaks autoregressive inference (batch size 1 at serving time); and NLP batch statistics are far noisier than vision's. LayerNorm needs no running statistics and behaves identically at train and test time.
Pre-LN vs Post-LN: the original transformer put LayerNorm after the residual add (Post-LN) and needed careful warmup to train at all. Pre-LN — x + Attn(LN(x)) — keeps a clean identity path through the whole network and trains stably without warmup. Every modern LLM uses Pre-LN. RMSNorm drops the mean subtraction (empirically the re-centring wasn't doing the work; the re-scaling was) for a ~10-15% speedup; used in LLaMA and most current models.
BatchNorm's train/test asymmetry is a classic bug source: at train time it uses batch statistics, at test time the running averages. Forget model.eval() and your inference silently changes. Also: BN's regularizing effect comes from the noise in batch statistics, which is why very small batches hurt it badly (use GroupNorm below batch size ~8).
Your training loss isn't decreasing. Debug it out loud.
Bisect the problem — is it the data, the model, or the optimisation?
- Overfit a single batch first. Take 8 examples and train until loss ≈ 0. If you can't, the bug is in the model or the training loop, not the data or hyperparameters. This is the highest-information single test and I always do it first.
- Check the initial loss. For C-class classification it should be ≈ ln(C) — 2.30 for 10 classes. If it isn't, the labels, the loss, or the output layer is wrong. A wildly high initial loss usually means the output layer bias is badly initialised.
- Learning rate. Too high → loss oscillates or NaNs; too low → flat line. Sweep by orders of magnitude, not by 10%.
- Check gradients actually flow. Print gradient norms per layer. All zeros → detached tensor, wrong
requires_grad, dead ReLUs, or a missingbackward(). Exploding → clip to a max norm of ~1.0. - Data pipeline. Are labels aligned with inputs after shuffling? Is normalization applied? Visualise a batch as the model sees it — mismatched channel order or a double-normalized image is common.
- Common code bugs: forgetting
optimizer.zero_grad()(gradients accumulate), applying softmax before a loss that already includes it, wrong reduction axis, a mask broadcasting incorrectly. - Only then: capacity, architecture, normalization placement, initialisation.
The last line to say: "and I'd log everything — loss curves, grad norms, LR, activation statistics — because debugging without instrumentation is guessing."
Dropout: mechanism, and what happens at inference?
During training, zero each unit independently with probability p. This prevents co-adaptation — no unit can rely on a specific other unit being present — and forces redundant, distributed representations. It approximates training an exponential ensemble of subnetworks that share weights.
At inference, dropout is off and all units are active, so the expected input to the next layer is larger by a factor of 1/(1−p). Two ways to correct it: scale activations by (1−p) at test time (original), or — universally now — inverted dropout: divide by (1−p) during training, so inference code needs no change at all.
When not to use it: with BatchNorm (the variance shift between train and test interacts badly — this is a documented conflict), on convolutional layers (spatially correlated activations make it weak; use DropBlock or rely on BN), and in modern LLM pretraining, where dropout is typically 0 because with a corpus larger than the model can memorise there's nothing to regularize against and it just wastes compute.
CNNs & sequence models
Less central than they were, but the reasoning — inductive bias, parameter sharing, receptive field — is what transformer questions build on.
Convolutions
Two inductive biases: locality (nearby pixels are related, so a small kernel suffices) and translation equivariance (shift the input, the feature map shifts identically). Pooling adds approximate translation invariance. These priors are why CNNs beat MLPs on images with far less data — and why ViTs need much more data or heavy augmentation to match them, since they must learn the priors instead of assuming them.
- Receptive field grows linearly with depth for stacked 3×3 convs (layer L sees 2L+1 pixels), and exponentially with dilation. Two stacked 3×3 convs have the same receptive field as one 5×5 but use 18C² instead of 25C² parameters and insert an extra nonlinearity — the VGG argument.
- 1×1 convolutions don't mix spatially; they mix channels and change channel depth cheaply. That's the bottleneck block in ResNet and the core of Inception.
- Depthwise separable (MobileNet): one spatial conv per channel, then a 1×1 to mix — cuts cost by roughly 1/k² + 1/C_out.
- ResNet:
y = x + F(x). Solves degradation (deeper plain nets did worse even on training error, so it wasn't overfitting) by making the identity easy to represent — a block only has to learn the residual. The gradient consequence is in §07.
RNNs, LSTMs, GRUs
RNN: h_t = tanh(W_hh h_{t−1} + W_xh x_t). Backprop through time unrolls it, and the gradient across k steps carries W_hhᵏ — so if the largest eigenvalue is <1 gradients vanish, >1 they explode. In practice an RNN can't learn dependencies beyond ~10 steps.
Why transformers replaced them: the recurrence is inherently sequential — you cannot compute h_t before h_{t−1}, so training can't be parallelised over the time axis. Attention computes all positions simultaneously, giving O(1) sequential operations instead of O(n), and a constant path length between any two tokens instead of O(n). You trade O(n²) compute for parallelism and short gradient paths, and on modern hardware that trade is overwhelmingly worth it.
Transformers, in full
Given your résumé, assume this is where the interview spends its time. Know the shapes, the complexity, and the reasons behind every design choice.
Scaled dot-product attention
Why divide by √d_k — the question that separates memorisers from understanders. If q and k have i.i.d. components with mean 0 and variance 1, their dot product Σᵢ qᵢkᵢ is a sum of d_k such terms, so it has variance d_k and typical magnitude √d_k. Feed logits of that scale into softmax and it saturates — the output approaches a one-hot vector, and softmax's gradient p(1−p) goes to zero. Dividing by √d_k returns the logits to unit variance, keeping the softmax in its responsive region and gradients alive.
The rest of the block
- FFN:
FFN(x) = W₂ · act(W₁x + b₁) + b₂, with hidden dimension 4× d_model by convention. Applied position-wise and identically at every position. It holds roughly ⅔ of the parameters — and mechanistic interpretability work suggests it acts as a key-value memory storing factual associations. - Residual + norm: Pre-LN, as in §08:
x = x + Attn(LN(x)), thenx = x + FFN(LN(x)). The residual stream is the model's working memory that every block reads from and writes to. - Causal mask: set the upper triangle of the score matrix to
−∞before softmax, so position i can't see j > i. This is what makes training parallel and autoregressive — every position's next-token prediction is computed in one forward pass without leaking the future.
Positional information
Attention is a permutation-equivariant set operation — shuffle the tokens and you get the same outputs, shuffled. Position must be injected explicitly.
| Scheme | How | Trade-off |
|---|---|---|
| Sinusoidal | Fixed sin/cos at geometrically spaced frequencies, added to embeddings | No parameters; relative offsets are a linear function of position, so in principle extrapolates. In practice, poorly |
| Learned absolute | A trainable embedding per position (BERT, GPT-2) | Simple and effective, but hard-capped at the trained context length |
| RoPE | Rotates Q and K by an angle proportional to position, in 2D subspaces | The dot product qᵀk then depends only on relative offset m−n. No added parameters, works with KV caching, extends via interpolation (NTK/YaRN). Default in LLaMA, Qwen, Mistral |
| ALiBi | Adds a linear distance penalty −m·|i−j| to the attention scores | Extremely simple, extrapolates to longer contexts than trained on, but biases toward recency |
Complexity, and how it's attacked
- FlashAttention — the important one. It doesn't change the math at all; it changes the memory access pattern. Attention is memory-bandwidth bound, not compute-bound: writing the n×n score matrix to HBM and reading it back dominates. FlashAttention tiles the computation, keeps blocks in SRAM, and uses the online-softmax trick to compute the exact result without ever materialising the full matrix. Result: exact attention, linear memory in n, 2–4× faster.
- Sparse / sliding-window (Longformer, Mistral): each token attends to a local window plus a few global tokens. Stacking layers still propagates information globally, like a receptive field.
- Linear attention: replace softmax with a kernel feature map so you can reassociate
(QKᵀ)V → Q(KᵀV)and get O(n·d²). Cheaper, generally worse quality. Related family: state-space models (Mamba), which recover recurrence with a parallel scan. - GQA / MQA: share K and V heads across multiple Q heads. Barely affects quality and cuts the KV cache by the sharing factor — the dominant memory cost at inference. Nearly universal now.
Walk me through what happens computationally when an LLM generates the 500th token.
Two distinct phases, and naming them is half the answer.
Prefill processed the prompt: all prompt tokens go through the model in one parallel forward pass, and every layer's K and V projections are stored in the KV cache. This phase is compute-bound — big matmuls, high GPU utilisation.
Decode generates one token at a time. For token 500: embed the single previous token, and at each layer compute Q, K, V for that one position only. Append the new K and V to the cache. Attention is then a (1 × 499) score vector against the cached keys, softmax, and a weighted sum of cached values. Through the FFN, then the LM head produces logits over the vocabulary, and a sampling step picks the next token.
Without the cache you would recompute K and V for all 499 previous tokens at every step — O(n²) redundant work per token, O(n³) overall.
The key insight to volunteer: decode is memory-bandwidth bound, not compute-bound. Each step reads the entire weight matrix from HBM to do a single matrix-vector product, so arithmetic intensity is terrible and the GPU sits mostly idle. That's why batching improves throughput enormously at no latency cost (the weights get reused across the batch), why quantization speeds up inference (fewer bytes to move), and why speculative decoding works — a small draft model proposes k tokens and the big model verifies all k in one parallel forward pass, converting k memory-bound steps into one.
KV cache size = 2 · layers · n · d_model · bytes · batch, which for a 70B model at 8k context runs to tens of gigabytes — hence GQA, paged attention (vLLM), and KV-cache quantization.
Why is BERT bidirectional and GPT not? What follows from that?
It's the training objective, and the mask that objective requires.
BERT is trained with masked language modelling: 15% of tokens are replaced and the model predicts them from both directions. No causal mask, so every token attends to every other. This gives richer contextual representations, which is why encoder models remain strong for classification, NER, and retrieval embeddings. But it can't generate autoregressively, and the objective is sample-inefficient — you only get a learning signal on 15% of positions per pass.
GPT is trained on next-token prediction with a causal mask. Every position gives a training signal, and the objective matches the generation task exactly. Predicting the next token forces the model to learn syntax, semantics, facts and reasoning as a side effect of compression — which turns out to scale extraordinarily well.
Consequence: encoders for understanding-and-scoring, decoders for generation. Note that decoder-only models have largely won even on tasks encoders used to own, because instruction-tuned generative models can be prompted to do them — though for pure embedding quality, bidirectional encoders still hold their ground.
Explain tokenization and one problem it causes.
BPE starts from characters (or bytes) and iteratively merges the most frequent adjacent pair, building a vocabulary of subwords. WordPiece is similar but merges by likelihood gain; SentencePiece/Unigram starts from a large vocabulary and prunes. Subwords are the compromise between character-level (long sequences, no OOV) and word-level (short sequences, huge vocab, OOV problems).
Problems it causes: arithmetic and character-level tasks are hard because numbers tokenize inconsistently ("1234" may be one token, "1235" three) and the model can't see letters inside a token — that's the actual cause of the "how many r's in strawberry" failure. Non-English and code are tokenized less efficiently, so the same content costs more tokens and more money. And glitch tokens — strings present in the tokenizer's training corpus but nearly absent from the model's — have essentially untrained embeddings and produce bizarre behaviour.
LLMs: train, align, serve
The pipeline end to end, plus the parts your own research touches. Expect the follow-ups here to go deeper than anywhere else in the interview.
The training stages
| Stage | Objective | Data | What it buys |
|---|---|---|---|
| Pretraining | Next-token cross-entropy | Trillions of tokens of web text, code, books | World knowledge, grammar, reasoning latent in compression. ~99% of the compute |
| SFT | Same objective, loss masked to response tokens only | 10k–1M curated instruction/response pairs | Format and instruction-following. Teaches how to answer, not new facts |
| Preference tuning | RLHF (PPO) or DPO on chosen/rejected pairs | Human or AI preference comparisons | Helpfulness, harmlessness, tone. Optimises what's hard to write as a loss |
| RLVR / reasoning RL | RL against a programmatic reward (test passes, answer matches) | Math, code, verifiable tasks | Long chain-of-thought, self-correction. The o1/R1 line |
DPO removes the reward model and the RL loop entirely. The insight: for the KL-constrained RLHF objective the optimal policy has a closed form, which can be inverted to express the implicit reward in terms of the policy itself. Substituting into the Bradley-Terry loss gives a plain supervised objective:
Parameter-efficient fine-tuning
- Scaling factor
α/rmultiplies the update; keeping α/r constant lets you change r without re-tuning the learning rate. - QLoRA: base weights frozen in 4-bit NF4, LoRA adapters in bf16, plus double quantization and paged optimizers. Puts a 70B fine-tune on a single 48GB GPU.
- Which modules: originally just Q and V projections; current practice targets all linear layers (Q,K,V,O and the FFN), which is usually worth the extra parameters.
- Limits: LoRA adapts style, format and task behaviour well; it's a poor tool for injecting large amounts of new knowledge — full fine-tuning or retrieval is better for that.
Decoding and test-time compute
- Greedy — argmax each step. Deterministic, repetitive, and myopic: locally optimal tokens can lead to globally poor sequences.
- Beam search — keep k partial hypotheses ranked by cumulative log-prob, with length normalization. Good for translation and summarisation where there's a single right answer; produces bland, degenerate text in open-ended generation.
- Temperature — divide logits by T before softmax. T<1 sharpens, T>1 flattens, T→0 is greedy. Note this is the same operation as calibration temperature scaling, used for a different purpose.
- Top-k — sample from the k most likely tokens. Top-p / nucleus — sample from the smallest set whose cumulative probability exceeds p; adaptive, so it uses a wide set when the model is uncertain and a narrow one when it's confident. Usually better than top-k for that reason. Min-p scales the cutoff by the top token's probability.
- Self-consistency — sample n chains of thought, take the majority answer. Reliably improves accuracy on reasoning benchmarks at n× the cost.
- Best-of-n / verifier reranking — sample n, score with a reward or process model, keep the best. Process reward models score intermediate steps rather than only the final answer, giving denser signal.
The framing: you can buy accuracy with training FLOPs or with inference FLOPs, and beyond a point the second is cheaper. Two axes — sequential (longer chains, self-revision) and parallel (sampling many candidates and aggregating). Snell et al. showed the optimal allocation is question-dependent: easy questions benefit from sequential refinement, hard ones from parallel search.
The open problem — and the one your paper sits on — is the stopping rule. Adaptive schemes need a signal for "am I done", and that signal is almost always model confidence. If confidence is uncalibrated, and specifically if it degrades as you scale compute, then the stopping rule fails exactly where it matters: the model becomes more confident while becoming no more correct. Be ready to state that in two sentences.
Serving and efficiency
- Quantization: INT8/INT4 weights. Post-training (GPTQ uses second-order info to minimise layer-wise error; AWQ protects salient channels identified by activation magnitude) vs quantization-aware training (better, costlier). Weight-only quantization helps most because decode is memory-bound. Outlier activation channels are the main difficulty — hence per-channel scaling and mixed precision.
- Distillation: train a student on the teacher's soft output distribution. The "dark knowledge" is in the relative probabilities of the wrong classes — that a 7 looks somewhat like a 1 — which carries far more information per example than a one-hot label. Use temperature T on both sides and scale the gradient by T². When the teacher is API-only and gives no logits, you fall back to sequence-level distillation on sampled outputs, optionally with rationales.
- Speculative decoding: a small draft model proposes k tokens; the large model verifies them in one forward pass; a rejection-sampling rule guarantees the output distribution is identical to sampling from the large model alone. Pure latency win with no quality cost.
- Batching: continuous/in-flight batching (vLLM, TGI) plus PagedAttention, which stores the KV cache in non-contiguous blocks like OS virtual memory to eliminate fragmentation.
- MoE: replace the FFN with N experts and route each token to the top-k (usually 2). Parameter count scales without proportional FLOPs. Requires a load-balancing auxiliary loss or a few experts get all the traffic; and all experts must be in memory even though only k are active, so it trades memory for compute.
RAG and agents
RAG: embed the corpus, retrieve top-k chunks by vector similarity for a query, put them in the prompt. Use it when knowledge is dynamic, private, verifiable, or too large to fit; fine-tune instead when you need a behaviour, format, or style. The honest failure modes: retrieval quality caps everything (garbage retrieved, garbage generated); chunking destroys context across boundaries; pure dense retrieval misses exact terms, so hybrid BM25 + dense with a cross-encoder reranker is the strong default; and "lost in the middle" — models attend poorly to content in the centre of a long context, so put the highest-ranked chunk first or last.
Agents: an LLM in a loop with tools, memory, and a termination condition. ReAct interleaves reasoning traces with tool calls. Practical concerns you should raise unprompted: error compounding over steps (95% per-step reliability is 60% over 10 steps), the need for schema-constrained output to make tool calls parseable, deterministic guardrails around irreversible actions rather than trusting the model, and evaluation being genuinely hard — trajectory-level judging, not just final-answer accuracy.
Why do LLMs hallucinate, and what actually reduces it?
Root causes, in order of importance:
- The objective. Next-token prediction rewards plausible continuations, not true ones. Nothing in pretraining distinguishes a fact from a fluent-sounding invention with similar statistics.
- No mechanism for "I don't know." The model must place probability mass somewhere; a confident-sounding answer has higher likelihood under the training distribution than hedging, because human text is mostly written by people who knew the answer.
- Alignment can make it worse. If human raters prefer confident, complete answers, RLHF trains the model to be confident and complete even when it shouldn't be — a direct reward for overconfidence.
- Long-tail facts are seen a handful of times and stored as weak, interfering associations.
What helps: retrieval grounding with citations the user can check; abstention thresholds on a calibrated confidence signal; consistency checks (sample multiple times — disagreement across samples correlates with error, which is the basis of semantic entropy methods); training on data that includes appropriate uncertainty; and verifier models or constrained decoding for structured outputs.
The line to end on: "None of these eliminate it, because the failure is in the objective, not the implementation. The practical target is making the model's uncertainty usable — which is a calibration problem, and calibration is what my own research measures."
Chain-of-thought: what's the actual mechanism?
The computation available for a single token is fixed — a constant number of layers. CoT lets the model spend more forward passes on a problem by writing intermediate results into the context, where they're re-readable. It converts a depth-limited computation into a longer sequential one, and it decomposes a hard mapping into steps each of which is well-represented in the training distribution.
Complexity-theoretic framing worth knowing: a fixed-depth transformer is in a limited circuit class, but with a CoT of polynomial length it can simulate a polynomial-time Turing machine. The scratchpad is genuine extra computation, not just prompting flavour.
Important caveats to volunteer: the stated reasoning is not guaranteed to be the model's actual computation — faithfulness studies show models reaching an answer for reasons the trace doesn't mention. CoT mainly helps on multi-step problems and can hurt on tasks where deliberation interferes. And it emerges with scale — small models often produce fluent reasoning that doesn't improve accuracy at all, which is a nice bridge to your own work.
How would you evaluate an LLM feature you built?
Layered, cheapest first:
- A golden set of 100–300 hand-curated cases covering the real distribution, including known failure categories. Small enough to inspect, large enough to move.
- Deterministic checks where possible — schema validity, exact match on extractable fields, tool-call correctness, latency, cost per request. Never use an LLM judge for something a regex can decide.
- LLM-as-judge for the subjective remainder, with a written rubric, pairwise comparison rather than absolute scoring (more reliable), position-swapping to cancel order bias, and validation against human labels on a subsample so you know the judge's agreement rate.
- Ablations — remove one component at a time to show what's actually carrying the performance. Without this, you have a number, not a finding.
- Online: A/B test with a guardrail metric, plus logging for regression detection. Offline metrics are a proxy; user behaviour is the truth.
This maps directly onto your Samsung benchmark work — say so, with the numbers.
Applied ML & system design
The round where they find out whether you've shipped anything. The scoring rubric is almost always: did the candidate clarify requirements before designing, and did they mention data before models?
The framework — use it verbatim
- Clarify (2–3 min, always). What's the business objective, and what proxy metric represents it? What's the scale — QPS, users, items? Latency budget? Is it online or batch? What data exists already? Any fairness or regulatory constraints?
- Frame as an ML problem. What's the input, what's the output, what's a training example, and where does the label come from? Is ML even necessary — would a heuristic get 80% of the value?
- Data. Sources, volume, labelling strategy, the train/val/test split (and why it's split that way), class balance, leakage risks.
- Features. Include a baseline. Say which are computed offline vs at request time, and how you'd avoid train/serve skew.
- Model. Start with the simplest thing that works — logistic regression or gradient boosting — and justify complexity only against a measured gap.
- Evaluation. Offline metric, online metric, and the guardrail metric you'd refuse to regress.
- Serving. Latency, batching, caching, precomputation, model size.
- Monitoring. Drift, staleness, retraining trigger, rollback plan.
- Preprocessing before splitting — fitting a scaler, imputer, or target encoder on the full dataset leaks test statistics into training. Fit on train, transform test. Use a Pipeline so it's structurally impossible.
- Temporal leakage — any feature computed with information unavailable at prediction time. "Total lifetime purchases" computed over the whole history leaks the future.
- Duplicate or near-duplicate rows straddling the split — extremely common in scraped and augmented datasets.
- Group leakage — the same user/patient/document in both splits.
- Target leakage — a feature that's a downstream consequence of the label. "Was assigned to the fraud review queue" predicts fraud perfectly and doesn't exist at inference.
Imbalanced data
Order of preference: (1) change the metric and threshold — often nothing else is needed; (2) class weights in the loss — cheap, principled, no data distortion; (3) resampling — undersample the majority if you have plenty of data, oversample or SMOTE the minority if you don't (SMOTE interpolates between minority neighbours; it's unreliable in high dimensions and can create points inside the majority region); (4) reframe as anomaly detection if the positive rate is extreme (<0.1%). Critical detail: resample only the training fold, never the validation or test set, or your metrics describe a distribution that doesn't exist.
Drift and monitoring
- Data/covariate drift: P(x) changes, P(y|x) doesn't. Detect with PSI, KS test, or KL between reference and live feature distributions. Often survivable.
- Concept drift: P(y|x) changes — the relationship itself moved. Requires retraining. Detectable only once labels arrive, which is why proxy metrics matter.
- Label lag: if ground truth takes 30 days, monitor prediction distribution and feature drift as leading indicators.
- Training–serving skew: the top cause of "great offline, bad online". Same feature code path for both, or a feature store, ideally with logging of the exact features used at inference.
Design a recommendation system for a video platform.
Clarify first: objective — watch time, or long-term retention? Those diverge, and optimising raw watch time gives you clickbait. Catalogue size, DAU, latency budget (~100ms), cold start requirements.
Two-stage architecture, which is the expected answer:
- Candidate generation — millions of items to a few hundred, in milliseconds. Multiple sources: a two-tower neural retriever (user tower and item tower trained with in-batch negatives and sampled softmax, item embeddings indexed in ANN), collaborative filtering, trending, and recent-subscription items. Cheap, high recall, no expensive features.
- Ranking — a few hundred to a page, with an expensive model. Gradient-boosted trees or a deep model over user × item × context features, trained on logged impressions with multi-task heads (P(click), P(complete), P(like)) combined into one utility score.
- Re-ranking — diversity (MMR or determinantal point process), freshness, business rules, filtering already-watched.
The parts that show experience: position bias in the logged data — items shown at the top get clicked more regardless of quality, so train with position as a feature and set it to a constant at inference, or use inverse-propensity weighting. Feedback loops — the model only sees data from what it recommended, so add exploration. Cold start — content features and embeddings for new items, popularity fallback for new users. Evaluation — offline NDCG/recall@k for iteration, but the decision comes from an online A/B test with a retention guardrail.
How do you know an A/B test result is real?
Compute the required sample size before starting, from the minimum detectable effect, baseline variance, α = 0.05, and power = 0.8. Fix the duration in advance and run at least a full weekly cycle to absorb day-of-week effects.
Peeking is the classic error: checking daily and stopping when p < 0.05 inflates the false positive rate far above 5%, because you get many chances to cross the line by chance. Fixes: fixed horizon, sequential testing with alpha spending, or always-valid inference.
Other things to raise: randomisation unit must match the interference structure (user-level, not session-level, if the treatment persists); check for sample ratio mismatch as an assignment-bug detector; multiple-comparison correction if you're testing many metrics; novelty effects that decay after a week; and network effects in social or marketplace settings, which break the independence assumption and need cluster randomisation.
Code it from scratch
The implementations most often asked for on a shared screen. Read them until the shapes feel obvious — then close the page and write one from memory.
Numerically stable softmax + cross-entropy
# The max subtraction is the whole point of the question. def softmax(z): # z: (N, C) z = z - z.max(axis=1, keepdims=True) # shift-invariant: same result, no overflow e = np.exp(z) return e / e.sum(axis=1, keepdims=True) def cross_entropy(logits, y): # y: (N,) integer labels p = softmax(logits) N = logits.shape[0] return -np.log(p[np.arange(N), y] + 1e-12).mean() def ce_grad(logits, y): # dL/dlogits — the (p - y)/N result p = softmax(logits); N = logits.shape[0] p[np.arange(N), y] -= 1.0 return p / N
Linear regression by gradient descent
def linreg(X, y, lr=0.01, epochs=1000, l2=0.0): n, d = X.shape w, b = np.zeros(d), 0.0 for _ in range(epochs): pred = X @ w + b err = pred - y # (n,) dw = (X.T @ err) / n + l2 * w # (d,) ridge term db = err.mean() w -= lr * dw; b -= lr * db return w, b # Closed form, for comparison: w = np.linalg.solve(X.T@X + l2*np.eye(d), X.T@y) # Use solve(), never inv() — better conditioned and ~2x faster.
Logistic regression
def sigmoid(z): return np.where(z >= 0, 1/(1+np.exp(-z)), np.exp(z)/(1+np.exp(z))) # stable both tails def logreg(X, y, lr=0.1, epochs=2000, l2=0.0): n, d = X.shape w, b = np.zeros(d), 0.0 for _ in range(epochs): p = sigmoid(X @ w + b) g = p - y # identical form to linreg — GLM w -= lr * ((X.T @ g)/n + l2*w) b -= lr * g.mean() return w, b
k-means with k-means++ init
def kmeans(X, k, iters=100, seed=0): rng = np.random.default_rng(seed) C = [X[rng.integers(len(X))]] for _ in range(k-1): # k-means++ seeding d2 = ((X[:,None,:] - np.array(C)[None,:,:])**2).sum(-1).min(1) C.append(X[rng.choice(len(X), p=d2/d2.sum())]) C = np.array(C) for _ in range(iters): d2 = ((X[:,None,:] - C[None,:,:])**2).sum(-1) # (n, k) lab = d2.argmin(1) newC = np.array([X[lab==j].mean(0) if (lab==j).any() else C[j] for j in range(k)]) if np.allclose(newC, C): break # converged C = newC return C, lab
Two-layer network, full forward + backward
def train_step(X, y, W1, b1, W2, b2, lr): N = X.shape[0] # ---- forward z1 = X @ W1 + b1; a1 = np.maximum(0, z1) # ReLU z2 = a1 @ W2 + b2 loss = cross_entropy(z2, y) # ---- backward d2 = ce_grad(z2, y) # (N, C) dW2 = a1.T @ d2; db2 = d2.sum(0) # (h,C), (C,) d1 = (d2 @ W2.T) * (z1 > 0) # ReLU' is the mask dW1 = X.T @ d1; db1 = d1.sum(0) # (d,h), (h,) # ---- update for p, g in ((W1,dW1),(b1,db1),(W2,dW2),(b2,db2)): p -= lr * g return loss
Self-attention with causal mask
def attention(X, Wq, Wk, Wv, causal=True): Q, K, V = X @ Wq, X @ Wk, X @ Wv # each (n, d_k) d_k = Q.shape[-1] S = Q @ K.T / np.sqrt(d_k) # (n, n) — the sqrt matters if causal: n = S.shape[0] S = np.where(np.tril(np.ones((n,n))) == 0, -1e9, S) A = softmax(S) # rows sum to 1 return A @ V, A # (n, d_v) def multi_head(X, Wq, Wk, Wv, Wo, h): # Wq: (d, d) n, d = X.shape; dk = d // h Q = (X@Wq).reshape(n,h,dk).transpose(1,0,2) # (h, n, dk) K = (X@Wk).reshape(n,h,dk).transpose(1,0,2) V = (X@Wv).reshape(n,h,dk).transpose(1,0,2) S = Q @ K.transpose(0,2,1) / np.sqrt(dk) # (h, n, n) S = np.where(np.tril(np.ones((n,n))) == 0, -1e9, S) A = np.exp(S - S.max(-1, keepdims=True)) A = A / A.sum(-1, keepdims=True) O = (A @ V).transpose(1,0,2).reshape(n, d) # concat heads return O @ Wo
LayerNorm and RMSNorm
def layer_norm(x, g, b, eps=1e-5): # normalise the FEATURE axis mu = x.mean(-1, keepdims=True) var = x.var(-1, keepdims=True) return g * (x - mu) / np.sqrt(var + eps) + b def rms_norm(x, g, eps=1e-6): # no mean subtraction, no bias return g * x / np.sqrt((x**2).mean(-1, keepdims=True) + eps)
ROC-AUC without sklearn (the rank trick)
def auc(y, s): # AUC = P(score of a random positive > score of a random negative) order = np.argsort(s) ranks = np.empty(len(s)); ranks[order] = np.arange(1, len(s)+1) n_pos, n_neg = y.sum(), len(y) - y.sum() return (ranks[y==1].sum() - n_pos*(n_pos+1)/2) / (n_pos*n_neg) # This is the Mann-Whitney U statistic. Handles ties poorly — use # average ranks (scipy.stats.rankdata) if ties are common.
Best-of-n with a confidence threshold
# Adaptive test-time compute: stop early when calibrated confidence clears tau. def adaptive_sample(model, prompt, n_max=16, tau=0.9, T=1.0): answers, budget = [], 0 for i in range(n_max): out, logits = model.generate(prompt) conf = softmax(logits / T).max() # T fitted on a HELD-OUT split answers.append(out); budget += 1 agree = sum(a == out for a in answers) / len(answers) if i >= 2 and conf > tau and agree > tau: break # the stopping rule — only as good return majority(answers), budget # as the calibration underneath
Say the shapes out loud as you go — "X is N by d, W1 is d by h, so z1 is N by h". It catches your own bugs, it's the single clearest signal that you've done this before, and it gives the interviewer something to agree with while you think.
Rapid-fire bank
Short questions, short answers. Read the question, answer in your head, then open. Anything you fumble goes on the revisit list.
Fundamentals
Parametric vs non-parametric?
Parametric models have a fixed number of parameters regardless of dataset size (linear regression, neural nets), so training data can be discarded afterwards. Non-parametric models grow with the data (kNN, kernel SVM, decision trees, Gaussian processes) — more flexible, more data-hungry, more expensive at inference.
Why standardise features? Which models don't need it?
Anything based on distances (kNN, k-means, SVM-RBF, PCA) or on gradient descent over a shared learning rate needs it — otherwise features with larger ranges dominate the distance metric or produce badly-conditioned loss surfaces with elongated contours that gradient descent zig-zags across. Also required for meaningful L1/L2 penalties, since the penalty is scale-dependent. Tree-based models don't need it — splits are threshold comparisons, invariant to any monotone transform.
How do you handle missing values?
First ask why they're missing — MCAR, MAR, or MNAR (missing not at random, where missingness depends on the unobserved value itself and any imputation biases the model). Options: drop rows if rare and MCAR; drop the column if mostly missing; impute with median (robust) or mode; model-based imputation (kNN, MICE, iterative); add a binary "was_missing" indicator, which is often the most informative feature since missingness itself carries signal; or use a model with native support — LightGBM and XGBoost learn a default split direction for missing values. Always fit the imputer on the training fold only.
Encode a categorical feature with 10,000 levels.
One-hot is out — 10,000 sparse columns. Options: target/mean encoding with out-of-fold computation and smoothing toward the global mean for rare levels (leaks badly if done naively — this is the point of the question); frequency encoding; hashing to a fixed number of buckets (collisions, but bounded memory and handles unseen levels); learned embeddings if you're using a neural net; or CatBoost, whose ordered target statistics solve the leakage properly. Also consider grouping the long tail into "other".
Type I vs Type II error, in ML terms?
Type I = false positive = you flagged something that wasn't there (controlled by α). Type II = false negative = you missed something real (controlled by 1 − power). Precision is degraded by Type I, recall by Type II. Which one matters is a product decision: a cancer screen tolerates false positives to avoid misses; a spam filter does the reverse.
What is the kernel trick, in one sentence?
Computing inner products in a high- or infinite-dimensional feature space without ever constructing the feature vectors, by replacing φ(x)ᵀφ(x') with a kernel function K(x,x') — valid whenever K is symmetric positive semi-definite (Mercer's condition).
Deep learning
Batch size: what does increasing it do?
Lower-variance gradient estimates, better hardware utilisation, faster wall-clock epochs. But: less gradient noise means less implicit regularization, and large-batch training is known to converge to sharper minima that generalise worse unless you scale the learning rate (linear or √ rule) and use warmup. Memory scales with batch size, so gradient accumulation simulates a large batch on small hardware. Very small batches make BatchNorm statistics unreliable.
Epoch, batch, iteration — define all three.
Iteration = one parameter update from one batch. Batch = the group of examples in that update. Epoch = one full pass over the training set = ceil(n / batch_size) iterations. Note LLM pretraining often runs less than one epoch on the full corpus.
What are skip connections for, besides gradients?
Three things: (1) a gradient highway — the Jacobian is I + ∂F/∂x, so a magnitude-1 path exists to every earlier layer; (2) they make the identity function trivially representable, which fixes the degradation problem where deeper plain networks had higher training error; (3) they create an implicit ensemble of paths of different depths, so the network behaves like an ensemble of shallower networks. In transformers they also define the residual stream — a shared communication channel that every block reads from and writes to.
Transfer learning: freeze or fine-tune?
Depends on target dataset size and domain similarity. Small + similar: freeze the backbone, train the head only. Large + similar: fine-tune everything at a low learning rate. Small + different: freeze early layers (generic edges/textures transfer), fine-tune later ones. Large + different: fine-tune everything, or reconsider whether pretraining helps. Use discriminative learning rates — lower for earlier layers — and watch for catastrophic forgetting, which is what LoRA and adapters mitigate by leaving the base weights untouched.
Why does label smoothing help?
Replace the one-hot target with (1−ε) on the true class and ε/(K−1) elsewhere. A hard one-hot target is only reachable with infinite logits, so the model is pushed to ever-larger logit gaps — overconfidence with no accuracy gain. Smoothing gives a finite optimum, improves calibration and usually generalisation. Trade-off worth naming: it compresses the representation of each class into a tighter cluster, which can hurt when you want the penultimate embeddings for retrieval or distillation.
Contrastive learning — what's the objective?
Pull representations of positive pairs (two augmentations of the same image, a query and its relevant document) together and push negatives apart. InfoNCE: −log[ exp(sim(z,z⁺)/τ) / Σ exp(sim(z,zⱼ)/τ) ] — a softmax over one positive and many negatives, so it's really cross-entropy over a similarity-based classification. τ controls how much the loss focuses on hard negatives. The number and difficulty of negatives is the main quality lever, which is why in-batch negatives with large batches (SimCLR) or a momentum queue (MoCo) matter, and why hard-negative mining dominates retrieval training.
LLMs
What do scaling laws say, and what changed with Chinchilla?
Loss falls as a power law in parameters, data, and compute, over many orders of magnitude. The original Kaplan et al. work implied that given more compute you should mostly grow the model. Chinchilla (Hoffmann et al.) corrected this: for a fixed compute budget, parameters and tokens should scale roughly equally — about 20 tokens per parameter. Existing large models were badly under-trained; a 70B model trained on 1.4T tokens beat a 280B model trained on 300B. The later shift: since inference cost is paid forever, it's often optimal to train a smaller model far past Chinchilla-optimal, which is what the small-but-heavily-trained open models do.
What is emergence, and is it real?
The claim: some capabilities appear abruptly at a scale threshold rather than improving smoothly. The important counterargument (Schaeffer et al.) is that apparent emergence is often an artifact of discontinuous metrics — exact-match accuracy on a multi-step task is a step function over a smoothly improving per-token probability. Switch to a continuous metric like token edit distance or log-likelihood and the curve is smooth. Give both sides; the willingness to state the deflationary account is what scores.
Perplexity — define it and say what it hides.
PPL = exp(mean negative log-likelihood per token) — the effective branching factor, i.e. how many equally-likely options the model is choosing among. Lower is better. What it hides: it's tokenizer-dependent, so it's not comparable across models with different vocabularies; it measures fit to a distribution, not usefulness — an aligned model often has worse perplexity than its base model while being far more useful; and it's dominated by common tokens, so it barely moves for large changes in rare-but-important behaviour.
Catastrophic forgetting — cause and mitigations?
Fine-tuning on a narrow distribution moves weights to minimise the new loss with no constraint to preserve old behaviour, so unrelated capabilities degrade. Mitigations: replay a mixture of the original data; regularise toward the base weights (L2-SP, EWC weights the penalty by the Fisher information so important parameters move less); PEFT methods like LoRA that leave the base intact and can be switched off; lower learning rates and fewer epochs; and multi-task rather than sequential training where possible.
Prompt engineering techniques that actually have evidence behind them?
Few-shot examples (format matters more than correctness of the labels, per Min et al.); chain-of-thought for multi-step problems; self-consistency; decomposition into sub-tasks; explicit output schemas; putting the most important content at the start or end of a long context; and giving the model an explicit option to say it doesn't know. Weak or overstated: elaborate persona prompts, emotional appeals, and threats — effects are small and inconsistent across models. Being willing to say which techniques don't replicate is a good signal.
Why can't you just increase context length instead of using RAG?
Four reasons. Cost: attention is quadratic in n, and even with linear-memory kernels you pay for every token in every request, forever. Latency: prefill time grows with context. Quality: retrieval degrades in the middle of long contexts, and irrelevant content actively distracts the model — more context is not monotonically better. And staleness: a long context still has to be assembled from somewhere, which is retrieval by another name. The right framing is that long context and RAG are complementary — retrieval selects, context holds.
Defending your own work
Highest-leverage section on the page. Everything above is shared with every other candidate. This part isn't — and interviewers spend more time here than on any single technical topic.
The 90-second version of the calibration paper
Structure it as problem → why it's hard → what you did → what you found → what's still open. Something like:
"Test-time compute scaling assumes that if you let a model generate more, you can tell when it's converged and stop. That stopping signal is almost always the model's own confidence. So I asked whether small models — under 3B — actually have usable confidence under test-time scaling.
I evaluated seven models from 0.5B to 7B, including instruct and R1-distilled variants, across GSM8K, MATH-500 and MMLU-Pro under a generate-once protocol. The finding is what I call the confidently-wrong trap: as you scale compute, confidence and correctness decouple in the small models — they get more certain without getting more right, which means a confidence-based stopping rule fails precisely where you need it.
The second contribution is a calibration guard that's leak-free — the threshold is fitted on a split disjoint from both the training and evaluation data, because fitting it on the eval set is the obvious way to make these numbers look much better than they are."
Then stop. Let them ask. Talking for four minutes unprompted is the most common self-inflicted wound in a research interview.
The questions they will actually ask
"What's the weakness of this work?"
This is a test of intellectual honesty, and the failure mode is deflecting. Answer it directly, pick a real limitation, and show you've thought about how to address it.
Candidate answers: the generate-once protocol is a deliberate simplification and doesn't capture sequential-revision regimes; three benchmarks is narrow and two of them are math, so it's unclear how far the finding generalises to open-ended tasks; confidence is measured one specific way and alternative signals (semantic entropy, verifier scores, self-consistency agreement) might behave differently; and the model set is open-weights only, so nothing is said about frontier models.
Then convert it forward: "The version I'd want next is a multi-round protocol with several confidence estimators compared head to head, because right now I can show the failure but I can't fully attribute it."
Do not say "no real weaknesses" or list only trivial ones. Interviewers read that as either dishonesty or not having engaged with the work deeply.
"Why does calibration collapse in small models specifically?"
Be honest about what you measured versus what you're hypothesising — that distinction is itself the answer they're grading.
Plausible mechanisms to offer as hypotheses: smaller models have less capacity to represent uncertainty as a separate quantity from the answer; instruction-tuning and preference optimisation reward confident-sounding outputs, and that pressure is relatively larger for a small model with less to fall back on; distilled reasoning models are trained to imitate confident traces from a stronger teacher, inheriting the confidence without the competence behind it; and longer generations give more opportunity for the model to condition on its own earlier tokens, which reinforces an initial commitment regardless of correctness.
End with: "I have the correlational result; I haven't isolated the cause. Distinguishing them would need a controlled ablation over the tuning stages."
"Why should I care? What does this change for someone building a product?"
Concrete consequences: any system that routes between a cheap and an expensive model on confidence will mis-route; any early-exit or adaptive-compute scheme will exit early on exactly the wrong questions; any abstention or escalate-to-human threshold will let confident errors through, which is the most expensive failure class because nobody reviews them. And any of those thresholds tuned on the evaluation set will look fine in the report and fail in production.
Practical recommendation: don't use raw confidence from a sub-3B model as a control signal; use agreement across samples, or a separate verifier, or calibrate on a genuinely held-out split and re-check it after any fine-tuning.
The Samsung project — have these numbers ready
The interviewer wants your decisions, not the org chart. Lead with the design choice and the reason for it.
- Why a knowledge graph rather than dumping state into the prompt? Implicit commands ("it's too dark in here") need resolution against device capabilities and current state. The two-layer split — a static ontology of what devices can do, plus a dynamic layer of what's currently true — keeps the slow-changing schema separate from fast-changing state, and lets you query rather than re-serialise everything into context each turn. Have the KG-vs-JSON ablation ready as evidence.
- Why four modules instead of one call? Device → function → condition → NLG separates concerns that fail differently, so you can evaluate and fix them independently, and each sub-task gets a constrained output schema, which is what makes the output parseable at all.
- The safety gates — grounding, range/type, precondition, impact tier, self-consistency confidence — are deterministic, not model-judged. That's the defensible design position: never let a probabilistic system be the last check before an irreversible action. The 0% false-HIGH-impact rate across the 38-case suite is the number to quote.
- Distillation under a black-box teacher. No logits from an API model, so no classical soft-target KD — you're doing sequence-level distillation on sampled outputs plus rationales. Knowing exactly why the standard method doesn't apply is a strong signal.
- The benchmark: 5 models × 6 categories × 3 context modes, ~4,050 LLM calls, 1,800 scored comparisons, with Gemini as both benchmark and judge. Volunteer the flaw before they find it — using the same model as both a compared system and the judge is a bias you should name, along with how you'd control for it (swap judges, human-validate a subsample, position-swap).
- Be straight about scope: the SFT and GRPO scripts were delivered but not run by the final review, so there are no student checkpoints. Say that plainly if asked. Overstating it is the one thing that can actually cost you the offer.
Your other projects, in one line each
| Project | The one thing to say about it |
|---|---|
| G1-Scale | Tool-augmented LLM for large graph reasoning — a ReAct agent with a NetworkX traversal tool, SFT on oracle traces then RLVR with PPO, benchmarked across 0.5B–7B. The interesting claim: giving a small model a reliable tool beats scaling the model for structured tasks |
| RAG CS QA system | Five-subject retrieval agent with topic routing and claim-level fact verification, LangChain + ChromaDB. Built end to end in a 1.5-hour timed exam — mention the constraint, it reframes the scope as a strength |
| ViT image retrieval | 97.8% accuracy — be ready to say what the baseline was and what split it's measured on, or the number invites scepticism |
| Deep Researcher | Multi-step research agent — good material for a discussion about error compounding and trajectory evaluation |
"We" instead of "I". On team projects, state your specific contribution explicitly and generously credit the rest. Vagueness here reads as inflation.
Unearned numbers. Any metric you quote will be followed by "compared to what?" and "how did you measure it?". If you can't answer both, don't lead with the number.
Walk-in checklist
Read this one in the morning. Nothing new here — just the things that are easy to forget while nervous.
- Restate the question before answering. Buys thinking time, catches misunderstandings early.
- Three beats: definition → mechanism → when it breaks.
- Think out loud. Silence is unscoreable.
- Say shapes and units when doing math.
- If corrected, update immediately and say so. Defending a wrong answer costs more than the wrong answer did.
- Ask about constraints before designing anything.
- Bluff. Say what you'd expect from first principles and mark it as a guess.
- Lead with the most complex solution. Baseline first, always.
- Recite definitions without the mechanism.
- Quote a metric you can't contextualise.
- Talk for four minutes uninterrupted.
- Say a project was "basically done" when scripts weren't run.
Ask them something real
- How do you decide when a model is good enough to ship — what's the bar, and who sets it?
- What does the evaluation setup look like day to day? Offline suites, human review, online?
- Where does the team spend most of its time — data, modelling, or infrastructure?
- What's a recent result that surprised the team, or changed your approach?
- If I joined, what would the first project be?
Questions about evaluation and data always land well, because they signal you've worked on real systems rather than notebooks.
Revisit before you walk in
Everything you marked shaky, in order. If this list is long, work top-down — the earlier sections are more likely to come up.
- Nothing marked yet. Work through the drills above and mark each one honestly — the list builds itself.