Skip to content
samit
Interviews/Projects to Do

Projects to Do

12 questions

12 questions
Frontier Personal Projects
MathLM (the repo) - train and evaluate the own mini LLM stackfrontier▶

Goal: push the mathLM project into a strong interview narrative: data, training, eval, and deployment tradeoffs.

Milestones:

  • Document tokenizer/data pipeline and why one chose it
  • Run ablations: context length, model size, optimizer settings
  • Track evals on math/coding style prompts with failure analysis
  • Add inference benchmarks (TTFT, tok/s, memory) across FP16/INT8
  • Write a one-page architecture + lessons learned note for interviews

Why it matters: this is the most credible frontier proof because you built it yourself and can defend every decision.

vLLM Serving Lab - deploy an open model with batching, KV cache, and latency budgetsfrontier▶

Goal: productionize one open LLM endpoint with vLLM and measure real serving metrics.

Milestones:

  • Serve a 7B model with vLLM + OpenAI-compatible API
  • Add load testing and capture P50/P95 TTFT + tok/s
  • Compare FP16 vs AWQ/GPTQ quantized variants
  • Add request logging, rate limits, and fallback behavior
  • Publish a benchmark note with cost per 1M tokens
Multimodal CV Agent - combine detector + VLM + retrieval for practical QAfrontier▶

Goal: build a small multimodal agent for image/video question answering using the CV background.

Milestones:

  • Use an object detector/segmenter to extract structured scene context
  • Fuse with a VLM for natural-language responses
  • Add retrieval over docs/specs for grounded answers
  • Evaluate factuality and failure cases (hallucinations, missed objects)
  • Deploy demo with clear latency and quality metrics
Systems & Kernels
Triton FlashAttention - write fused attention as a GPU kernelfrontier▶

Goal: a fused attention kernel in Triton that uses tiling + the online-softmax trick so the N×N score matrix never touches HBM.

What it teaches: the GPU programming model (program ids, block pointers), why attention is IO-bound not compute-bound, the online-softmax rescaling, and tiling Q/K/V into SRAM.

Milestones:

  • Triton vector-add and softmax kernels to learn the model
  • Naive attention kernel (materialise scores) as a baseline
  • Tiled forward with online softmax
  • Match F.scaled_dot_product_attention numerically (atol/rtol)
  • Benchmark vs PyTorch eager; plot the speedup vs sequence length

Note: "I wrote FlashAttention in Triton" signals you understand the IO-bound nature of attention, not just the math. Be ready to explain the online softmax and why no N×N matrix hits HBM.

Resources: Triton fused-attention tutorial, FlashAttention (Dao et al., 2022).

Transformers in CUDA - a GPT forward/backward in raw C++/CUDAfrontier▶

Goal: a GPT block (ideally a full training step) in raw CUDA/C++, llm.c style - no PyTorch.

What it teaches: writing kernels for matmul, LayerNorm, softmax, GELU; memory coalescing, shared memory, warps and occupancy; what cuBLAS does for one; and moving from fp32 to mixed precision.

Milestones:

  • CPU reference forward in plain C
  • Port matmul to CUDA, verify against the reference
  • LayerNorm / softmax / GELU kernels
  • Wire a full forward pass, match PyTorch logits
  • Backward pass + a training step on Shakespeare
  • Profile with Nsight, chase occupancy and memory throughput

Note: the deepest "do you know the machine" flex. Even finishing the forward pass is strong. Karpathy's llm.c is the reference implementation.

Resources: Karpathy llm.c, CUDA C++ Programming Guide, the PMPP book.

Contribute to tinygrad - read and modify a real DL frameworkfrontier▶

Goal: explore George Hotz's tinygrad (a deliberately tiny autograd/DL framework) and land a real PR.

What it teaches: lazy evaluation, the kernel-fusion scheduler, how ops lower to a small instruction set, and how a framework is built end to end - not just used.

Milestones:

  • Read the codebase; run the MNIST and llama examples
  • DEBUG=2 to watch kernels fuse and schedule
  • Take a "good first issue" or add a missing op + test
  • Attempt a bounty PR (tiny corp pays for some)
  • Implement a model in tinygrad to stress the framework

Note: contributing to tinygrad is rare signal - it says you can read and modify a DL compiler, not just call .fit. Strong talking point for Hotz / comma / tiny corp style teams.

Resources: github.com/tinygrad/tinygrad, the docs/ and examples/, the Discord bounty board.

Models From Scratch
Llama 3 from scratch - implement the architecture and load real weightsfrontier▶

Goal: the Llama-3 decoder from scratch and load the released weights to run inference.

What it teaches: the modern decoder stack - RoPE, RMSNorm, SwiGLU, GQA, KV cache - plus the exact tensor shapes and why each replaced its predecessor (RoPE vs learned positions, RMSNorm vs LayerNorm, SwiGLU vs ReLU, GQA vs MHA).

Milestones:

  • RMSNorm, RoPE, and SwiGLU as standalone modules
  • Attention with grouped-query attention + KV cache
  • Assemble the block, load HF/Meta weights, match logits on a prompt
  • Sampling (temperature, top-p) and free-form generation
  • Stretch: paged KV cache or quantized inference

Note: covers half the "modern LLM internals" surface in one project. Any "why does Llama use X" becomes easy once you have built it. Ties directly to the MathLM/MiniTorch work.

Resources: the Llama 3 paper, Karpathy/Meta reference impls.

Stable Diffusion - read the paper and implement a text-to-image sampler▶

Goal: read the Latent Diffusion paper and implement a minimal text-to-image pipeline.

What it teaches: VAE encode/decode to a latent space, the U-Net denoiser, DDPM vs DDIM sampling, classifier-free guidance, and CLIP text conditioning via cross-attention.

Milestones:

  • DDPM on MNIST/CIFAR in pixel space (forward noising + reverse denoising)
  • Add a U-Net with a time embedding
  • DDIM sampler for fast (≈20-50 step) generation
  • Classifier-free guidance
  • Latent diffusion: VAE + text conditioning (or load SD weights and write only the sampler)

Note: diffusion is the other half of generative modeling (vs LLMs). Implementing the sampler proves you understand the forward/reverse process, not just the high-level idea.

Resources: DDPM (Ho et al., 2020), DDIM (Song et al., 2020), Latent Diffusion (Rombach et al., 2022).

Multimodal Vision-Language Model - code a VLM from scratch (PaliGemma-style)frontier▶

Goal: a vision-language model from scratch - a SigLIP/ViT vision encoder + a projection layer + a Gemma decoder - following Umar Jamil's end-to-end walkthrough.

What it teaches: how image patches become tokens the LLM consumes, the projection that bridges vision and text embedding spaces, contrastive (SigLIP/CLIP) vs generative training, and the two-stage train (align the projector, then instruction-tune).

Milestones:

  • SigLIP/ViT patch encoder
  • Linear/MLP projector into the LLM embedding space
  • Wire vision tokens + text tokens into the decoder
  • Load PaliGemma weights, caption an image
  • Visual question answering / instruction following

Note: multimodal is where frontier labs are pushing. Building one shows you can connect modalities, and it ties straight to the CV background.

Resources: Umar Jamil, "Coding a multimodal (vision) language model from scratch" (youtube.com/watch?v=vAmKB7iPkWw); PaliGemma, SigLIP, and LLaVA papers.

Applied & Agents
Coding harness - build an agentic plan/edit/run/test loop▶

Build: a coding agent harness that takes a task, plans, edits files, runs the code and tests, reads the output, and iterates until tests pass.

What it teaches: tool-use loops, structured output and function calling, context management across turns, sandboxed execution, and how to keep an agent from looping forever.

Milestones:

  • Single tool call: model proposes a shell command, you run it, feed back stdout
  • File read/write tools with a diff-based edit format
  • Plan then act: a scratchpad of steps before editing
  • Run tests, parse failures, retry with the error in context
  • Guardrails: step budget, timeout, and a stop condition
  • Score it on a small set of toy bugs (your own mini SWE-bench)

Ask: what are the failure modes of agentic systems and how do you bound them? You will have lived the answer.

RAG system (pi.dev) - retrieval over your own data, done properly▶

Build: a real retrieval-augmented generation system over your own corpus (notes, blog, papers), the kind you would ship on pi.dev.

What it teaches: chunking strategy, embeddings and vector search, hybrid dense+sparse retrieval, reranking, the "lost in the middle" problem, and how to actually evaluate retrieval quality instead of eyeballing it.

Milestones:

  • Ingest + chunk (try fixed vs semantic vs recursive) and embed
  • Dense vector search baseline, then add BM25 hybrid
  • Add a reranker over the top-k
  • Generation with citations grounded in retrieved chunks
  • Eval harness: retrieval recall@k and answer faithfulness, not vibes
  • Ship it behind a small API

Ask: when do you choose RAG over fine-tuning, and how do you measure a RAG system? You will have the numbers.

Ship one AI app end to end that people actually usestartup▶

Build: one complete AI product - not a notebook - that a real user can hit: a UI, a model behind an API, eval, and a deploy.

What it teaches: the unglamorous 90% that interviews probe: latency budgets, cost per request, prompt/version management, handling bad output, monitoring, and iterating on real usage.

Milestones:

  • Pick a narrow problem with a clear user and a clear success metric
  • Thin vertical slice: input - model - output, deployed
  • Add eval and logging on every request
  • Handle the long tail: refusals, timeouts, garbage input
  • Measure cost and latency, then optimize the worst one
  • Put it in front of real users and iterate on what breaks

Ask: startup rounds care that you can ship. A live URL beats any take-home.