MLOps & Production
20 questions
What is model drift? How do you monitor for it?▶
Types:
- Data drift (covariate shift): P(X) changes but P(Y|X) stays the same. Features look different. e.g., new user demographics.
- Label shift: P(Y) changes. Distribution of outcomes changes. e.g., more fraud during holiday season.
- Concept drift: P(Y|X) changes. The relationship between features and labels changes. Most dangerous. e.g., "COVID" was never used before 2020 but suddenly matters for medical NLP.
Monitoring:
- Input feature distribution: track mean, variance, quantiles per feature. Statistical tests: Kolmogorov-Smirnov, PSI (Population Stability Index).
- Output distribution: track label distribution of predictions. PSI on output scores.
- Model performance: track accuracy, AUC, F1 against ground truth (when labels arrive with delay).
- Alerting: PSI > 0.2 = significant drift, trigger investigation. PSI > 0.25 = retrain.
Production tip: log every prediction with features (sample 5-10% for storage). You'll need these for drift analysis and retraining datasets.
What is LLMOps? How does it differ from classical MLOps?startup▶
- Prompt versioning: prompts are code - version-control them, A/B test them, roll back bad prompt changes. Classical MLOps versions model weights; LLMOps also versions prompts.
- Evaluation pipeline: LLM outputs are non-deterministic and hard to evaluate (no clear ground truth). Need LLM-as-judge, human eval, golden datasets. Classical MLOps has clear metrics (AUC, RMSE).
- Cost management: token cost per query is explicit. Need to track per-user, per-feature costs. Optimize prompt length, cache repeated calls.
- Observability: trace complete LLM chains (LangSmith, LangFuse). Log: prompt, context, output, latency, token count, cost per call.
- Model catalog: swap between providers (OpenAI → Anthropic → local) behind an abstraction layer (LiteLLM). Can't do this with classical ML.
How do you implement A/B testing for LLM applications?startup▶
Challenge: LLM outputs are text - no simple binary metric. How do you know if version B is better than version A?
Approach:
- Define metrics: task-completion rate, user satisfaction score (thumbs up/down), downstream business metrics (conversion, session length)
- Traffic split: route X% of users to model A, (100-X)% to model B. Use consistent hashing on user_id for stable assignment.
- Guardrail metrics: latency, error rate, cost per request must not degrade
- Sample size: calculate minimum sample size for desired power (80%) and significance (p<0.05). Typically need 10K-100K sessions per variant.
- LLM-as-judge: offline evaluation - sample 500 outputs from each variant, have GPT-4 compare them pairwise. Fast, cheap, but biased toward GPT-4 style.
- Shadow mode: run B in parallel without serving to users. Compare outputs offline before exposing users.
What is CI/CD for ML models?▶
Continuous Integration / Continuous Deployment adapted for ML workflows.
CI for ML:
- Automated tests on every commit: data validation tests (schema, distributions), unit tests for preprocessing/feature engineering, model training smoke tests (1 epoch, small data)
- Model quality gates: hold-out evaluation must pass minimum thresholds before merge
CD for ML:
- Triggered by: new model trained, new data available, performance degradation alert
- Stages: train → evaluate → shadow deploy → canary (5% traffic) → full deploy
- Rollback: if quality drops, revert to previous checkpoint. For LLMs: revert prompt config or model version.
Tools: MLflow (tracking), DVC (data versioning), Weights & Biases (experiment tracking), GitHub Actions + Argo Workflows (pipeline orchestration), Seldon / BentoML / TorchServe (serving).
How do you handle PII in ML pipelines?▶
- Detection: NER models or regex patterns to identify PII (names, emails, SSNs, phone numbers) before storing/training
- Anonymization: replace PII with tokens ("PERSON_1", "EMAIL_1") - preserves structure for NLP models
- Pseudonymization: hash PII consistently so same entity gets same pseudonym (enables entity-level analysis without real data)
- Differential privacy: add calibrated Gaussian/Laplace noise to gradients during training. Provides formal privacy guarantees. Cost: accuracy loss.
- Federated learning: train on device, only send gradient updates (not raw data). Used in Gboard, Apple's keyboard.
- Access controls: RBAC on raw data, audit logs for data access
- GDPR: right to erasure - must be able to delete user's data and retrain or show data influence was removed
GPTQ vs AWQ - how do post-training quantization methods differ?▶
Both quantize a pretrained model's weights to low bit-width (commonly 4-bit) without retraining - but they pick which information to preserve differently.
GPTQ: quantizes layer-by-layer, column-by-column, using second-order (Hessian) information to choose the rounding that minimizes the resulting output error, and updates remaining unquantized columns to compensate for the error just introduced. Calibration is relatively expensive (requires Hessian computation) but the method is general-purpose.
AWQ (Activation-aware Weight Quantization): observes that not all weight channels matter equally - channels that interact with large-magnitude activations are disproportionately important for output quality. AWQ identifies a small % of these salient channels using activation statistics (not weight magnitude), scales them up before quantization to preserve their precision, and rescales back down at inference.
Practical difference: AWQ skips the expensive Hessian-based calibration GPTQ needs, is faster to apply, and tends to preserve quality better at very low bit-widths - which is why it's become the more common default for serving quantized open-weight LLMs.
ONNX, TorchScript, Triton Inference Server - when do you use each?▶
ONNX (Open Neural Network Exchange): an IR (intermediate representation) that captures a model's computation graph in a framework-agnostic format. Export once (PyTorch → ONNX), run anywhere (TensorRT, ONNX Runtime, Core ML, OpenVINO). Use when you need to deploy to a different runtime than you trained in (e.g., ONNX Runtime for CPU serving, TensorRT for GPU).
TorchScript: serialises a PyTorch model into a graph format that runs without the Python interpreter - enables deployment in C++ environments (mobile, embedded, latency-critical services). Two modes: tracing (run once, record operations - fails with dynamic control flow) and scripting (compile the code - handles conditionals but requires type annotations). Use when staying in the PyTorch ecosystem but need to eliminate Python overhead.
Triton Inference Server (NVIDIA): a model-serving framework that sits above the model format layer - it handles HTTP/gRPC endpoints, dynamic batching, concurrent model execution, multiple model instances, and GPU/CPU scheduling. Supports ONNX, TensorRT, PyTorch, TensorFlow backends. Use when you need a production inference server with batching, multi-model serving, and SLAs - it's what powers most large-scale GPU serving at Beltech and similar shops.
Typical stack: PyTorch training → export to ONNX or TensorRT → serve via Triton. TorchScript is a shortcut if they are already on PyTorch everywhere and don't need TRT optimisation.
TTFT vs tokens/sec - which do you optimize for, and why?frontierstartup▶
TTFT (time to first token): prefill phase - process the full prompt. Dominated by prompt length and weight bandwidth. Users feel this as "how long until it starts typing."
TPS (tokens/sec): decode phase - one token at a time. Memory-bandwidth bound at batch=1.
- Chat UX: TTFT < 500ms matters more than peak TPS
- Batch API / summarization: optimize aggregate TPS
- Prefix caching / prompt KV reuse cuts TTFT for repeated system prompts
vLLM vs TensorRT-LLM vs plain PyTorch serve - when to pick each?startup▶
- PyTorch eager: prototyping only. No continuous batching, poor GPU util.
- vLLM: PagedAttention + continuous batching. Best default for multi-tenant LLM APIs on NVIDIA GPUs.
- TensorRT-LLM: compiled graphs, kernel fusion, FP8. Highest throughput when you can afford compile time and fixed model arch.
- ONNX Runtime: cross-platform, good for smaller models / CPU fallback.
For CV at production edge-deploy edge deploy: TensorRT + Triton. For LLM SaaS: vLLM or TRT-LLM behind a gateway.
Walk through exporting a YOLO detector to TensorRT for production.startup▶
- Train in PyTorch → export ONNX (opset 17+, simplify graph)
- Build TRT engine: FP16, fixed input shape or dynamic batch
- Calibrate INT8 if edge GPU (Jetson) - need 500–1000 representative images
- Validate mAP on val set: FP16 usually <0.5% drop; INT8 can cost 1–2% if miscalibrated
- Serve via Triton or DeepStream with dynamic batching
Know the numbers: YOLOv10-n ~2ms/img on Orin NX FP16 vs ~15ms PyTorch CPU.
How do you calculate cost per 1M tokens for an LLM API?startup▶
Cost = (input_tokens × price_in + output_tokens × price_out) / 1M.
Self-hosted: GPU_hr × $/hr ÷ tokens_generated_per_hr.
Example: 1× A100, 7B int4, ~4000 tok/s aggregate with batching → 14.4B tok/hr. At $2/GPU-hr → ~$0.14 per 1M tokens (ignoring overhead). Compare to OpenAI GPT-4o-mini ~$0.15/1M input.
What signals do you use to autoscale GPU inference workers?frontier▶
- Request queue depth (best leading indicator)
- GPU utilization < 40% with growing queue → add replicas
- P99 TTFT or TPS breaching SLA
- KV cache memory pressure (vLLM block usage %)
Scale down slowly (cooldown 5–10 min) - GPU cold start + model load is 30s–2min.
Design a canary rollout for a new model version in production.startup▶
- Route 1–5% traffic to new model (consistent hash on user_id)
- Guardrails: error rate, P99 latency, cost/token, safety filter hit rate
- Offline: golden-set eval before any traffic (faithfulness, task accuracy)
- Shadow mode optional: run new model, log outputs, don't serve
- Rollback: flip traffic weight to 0, keep old weights hot
What do you log on every inference request in production?startup▶
- request_id, model_version, prompt_hash (not raw prompt if PII)
- input_tokens, output_tokens, TTFT_ms, total_latency_ms
- GPU_id, batch_size at decode step
- finish_reason (stop, length, error)
- downstream task outcome if available (thumbs up, conversion)
Sample 5–10% for full prompt/response storage with retention policy.
Sarvam / Krutrim-style: what breaks when you serve Indic LLMs at scale?startupfrontier▶
- Tokenization tax: Devanagari/Tamil text → 2–4× more tokens than English → higher cost, shorter effective context
- Code-mixing: Hinglish needs tokenizer trained on mixed corpora, not Hindi-only
- ASR→LLM pipeline latency: streaming ASR + LLM must stay <300ms for voice UX
- Eval gap: MMLU doesn't measure Indic quality - need IndicQA, IndicGLUE, human eval
- Data sovereignty: weights + logs stay in-region (IndiaAI mission requirement)
Frontier lab inference round: "7B model, 10K concurrent users, 2s P99 latency - design it."frontier▶
Back-of-envelope:
- Assume 50 tok avg output, need ~25 tok/s per user → unrealistic on single GPU
- Reality: async API, streaming, most users don't need 50 tok instantly
- Peak RPS × avg_output_tokens / aggregate_TPS = GPUs needed
- Architecture: LB → API gateway (rate limit) → vLLM pool with continuous batching → semantic cache for repeated queries
- Quantize to AWQ int4 to double effective throughput
AIOps for ML: how is it different from classical DevOps monitoring?startupfrontier▶
Classical DevOps watches CPU, memory, 5xx. AIOps adds:
- Data drift: feature/embedding distribution shift (PSI > 0.2)
- Model quality: online accuracy proxy when labels are delayed
- LLM-specific: hallucination rate, refusal rate, toxicity score, cost/token trend
- Automated remediation: rollback model version, route to fallback, alert on-call
Tools: Evidently, WhyLabs, Arize, Langfuse for LLM traces.
early-career CV/ML at a startup: what will they actually ask in a technical round?startup▶
Based on Sarvam, Krutrim, Yellow.ai, Uniphore, and similar:
- Walk through a project end-to-end (data → train → deploy → metrics)
- Fine-tune a 7B model: LoRA vs full FT, dataset format, eval
- Build or debug a RAG pipeline live
- YOLO/detector tradeoffs, mAP, TensorRT export (the Beltech story)
- Python + PyTorch fluency: write attention or dataloader on the spot
- System sense: "how would you serve this to 100 QPS?"
Less theory than frontier labs; more "show you can ship."
What is prefix caching? When does it help and when does it not?▶
Prefix caching (a.k.a. prompt caching) stores the KV cache for a shared prompt prefix so it is computed once and reused across requests, skipping the prefill for those tokens.
Big win when: many requests share a long fixed prefix - a system prompt, a few-shot template, or a long document in a multi-turn chat. The expensive prefill over those tokens happens once.
No help when: prompts are unique with no shared prefix, or the shared part is short. The cache is keyed on an exact token-prefix match, so any change near the start invalidates it.
Design note: put the stable, reused content at the front of the prompt and the variable content at the end to maximize cache hits. vLLM and most serving stacks support this automatically.
Explain disaggregated prefill/decode serving. Why split them?frontier▶
LLM serving has two phases with opposite hardware profiles. Prefill processes the whole prompt in parallel - compute-bound, high GPU utilization. Decode generates one token at a time - memory-bandwidth-bound, low compute utilization.
The problem: on one GPU they interfere. A long prefill blocks ongoing decodes (latency spikes), and decodes waste the compute a prefill would use.
Disaggregation: run prefill on one pool of GPUs and decode on another, transferring the KV cache between them. Each pool is tuned for its phase, so you hit better TTFT and steadier tokens/sec under mixed load.
Cost: the KV-cache transfer over the interconnect, and more orchestration. Worth it at scale (this is how frontier serving stacks hit their SLAs).