Projects to Do
12 questions
MathLM (the repo) - train and evaluate the own mini LLM stackfrontier▶
Goal: push the mathLM project into a strong interview narrative: data, training, eval, and deployment tradeoffs.
Milestones:
- Document tokenizer/data pipeline and why one chose it
- Run ablations: context length, model size, optimizer settings
- Track evals on math/coding style prompts with failure analysis
- Add inference benchmarks (TTFT, tok/s, memory) across FP16/INT8
- Write a one-page architecture + lessons learned note for interviews
Why it matters: this is the most credible frontier proof because you built it yourself and can defend every decision.
vLLM Serving Lab - deploy an open model with batching, KV cache, and latency budgetsfrontier▶
Goal: productionize one open LLM endpoint with vLLM and measure real serving metrics.
Milestones:
- Serve a 7B model with vLLM + OpenAI-compatible API
- Add load testing and capture P50/P95 TTFT + tok/s
- Compare FP16 vs AWQ/GPTQ quantized variants
- Add request logging, rate limits, and fallback behavior
- Publish a benchmark note with cost per 1M tokens
Multimodal CV Agent - combine detector + VLM + retrieval for practical QAfrontier▶
Goal: build a small multimodal agent for image/video question answering using the CV background.
Milestones:
- Use an object detector/segmenter to extract structured scene context
- Fuse with a VLM for natural-language responses
- Add retrieval over docs/specs for grounded answers
- Evaluate factuality and failure cases (hallucinations, missed objects)
- Deploy demo with clear latency and quality metrics
Triton FlashAttention - write fused attention as a GPU kernelfrontier▶
Goal: a fused attention kernel in Triton that uses tiling + the online-softmax trick so the N×N score matrix never touches HBM.
What it teaches: the GPU programming model (program ids, block pointers), why attention is IO-bound not compute-bound, the online-softmax rescaling, and tiling Q/K/V into SRAM.
Milestones:
- Triton vector-add and softmax kernels to learn the model
- Naive attention kernel (materialise scores) as a baseline
- Tiled forward with online softmax
- Match
F.scaled_dot_product_attentionnumerically (atol/rtol) - Benchmark vs PyTorch eager; plot the speedup vs sequence length
Note: "I wrote FlashAttention in Triton" signals you understand the IO-bound nature of attention, not just the math. Be ready to explain the online softmax and why no N×N matrix hits HBM.
Resources: Triton fused-attention tutorial, FlashAttention (Dao et al., 2022).
Transformers in CUDA - a GPT forward/backward in raw C++/CUDAfrontier▶
Goal: a GPT block (ideally a full training step) in raw CUDA/C++, llm.c style - no PyTorch.
What it teaches: writing kernels for matmul, LayerNorm, softmax, GELU; memory coalescing, shared memory, warps and occupancy; what cuBLAS does for one; and moving from fp32 to mixed precision.
Milestones:
- CPU reference forward in plain C
- Port matmul to CUDA, verify against the reference
- LayerNorm / softmax / GELU kernels
- Wire a full forward pass, match PyTorch logits
- Backward pass + a training step on Shakespeare
- Profile with Nsight, chase occupancy and memory throughput
Note: the deepest "do you know the machine" flex. Even finishing the forward pass is strong. Karpathy's llm.c is the reference implementation.
Resources: Karpathy llm.c, CUDA C++ Programming Guide, the PMPP book.
Contribute to tinygrad - read and modify a real DL frameworkfrontier▶
Goal: explore George Hotz's tinygrad (a deliberately tiny autograd/DL framework) and land a real PR.
What it teaches: lazy evaluation, the kernel-fusion scheduler, how ops lower to a small instruction set, and how a framework is built end to end - not just used.
Milestones:
- Read the codebase; run the MNIST and llama examples
DEBUG=2to watch kernels fuse and schedule- Take a "good first issue" or add a missing op + test
- Attempt a bounty PR (tiny corp pays for some)
- Implement a model in tinygrad to stress the framework
Note: contributing to tinygrad is rare signal - it says you can read and modify a DL compiler, not just call .fit. Strong talking point for Hotz / comma / tiny corp style teams.
Resources: github.com/tinygrad/tinygrad, the docs/ and examples/, the Discord bounty board.
Llama 3 from scratch - implement the architecture and load real weightsfrontier▶
Goal: the Llama-3 decoder from scratch and load the released weights to run inference.
What it teaches: the modern decoder stack - RoPE, RMSNorm, SwiGLU, GQA, KV cache - plus the exact tensor shapes and why each replaced its predecessor (RoPE vs learned positions, RMSNorm vs LayerNorm, SwiGLU vs ReLU, GQA vs MHA).
Milestones:
- RMSNorm, RoPE, and SwiGLU as standalone modules
- Attention with grouped-query attention + KV cache
- Assemble the block, load HF/Meta weights, match logits on a prompt
- Sampling (temperature, top-p) and free-form generation
- Stretch: paged KV cache or quantized inference
Note: covers half the "modern LLM internals" surface in one project. Any "why does Llama use X" becomes easy once you have built it. Ties directly to the MathLM/MiniTorch work.
Resources: the Llama 3 paper, Karpathy/Meta reference impls.
Stable Diffusion - read the paper and implement a text-to-image sampler▶
Goal: read the Latent Diffusion paper and implement a minimal text-to-image pipeline.
What it teaches: VAE encode/decode to a latent space, the U-Net denoiser, DDPM vs DDIM sampling, classifier-free guidance, and CLIP text conditioning via cross-attention.
Milestones:
- DDPM on MNIST/CIFAR in pixel space (forward noising + reverse denoising)
- Add a U-Net with a time embedding
- DDIM sampler for fast (≈20-50 step) generation
- Classifier-free guidance
- Latent diffusion: VAE + text conditioning (or load SD weights and write only the sampler)
Note: diffusion is the other half of generative modeling (vs LLMs). Implementing the sampler proves you understand the forward/reverse process, not just the high-level idea.
Resources: DDPM (Ho et al., 2020), DDIM (Song et al., 2020), Latent Diffusion (Rombach et al., 2022).
Multimodal Vision-Language Model - code a VLM from scratch (PaliGemma-style)frontier▶
Goal: a vision-language model from scratch - a SigLIP/ViT vision encoder + a projection layer + a Gemma decoder - following Umar Jamil's end-to-end walkthrough.
What it teaches: how image patches become tokens the LLM consumes, the projection that bridges vision and text embedding spaces, contrastive (SigLIP/CLIP) vs generative training, and the two-stage train (align the projector, then instruction-tune).
Milestones:
- SigLIP/ViT patch encoder
- Linear/MLP projector into the LLM embedding space
- Wire vision tokens + text tokens into the decoder
- Load PaliGemma weights, caption an image
- Visual question answering / instruction following
Note: multimodal is where frontier labs are pushing. Building one shows you can connect modalities, and it ties straight to the CV background.
Resources: Umar Jamil, "Coding a multimodal (vision) language model from scratch" (youtube.com/watch?v=vAmKB7iPkWw); PaliGemma, SigLIP, and LLaVA papers.
Coding harness - build an agentic plan/edit/run/test loop▶
Build: a coding agent harness that takes a task, plans, edits files, runs the code and tests, reads the output, and iterates until tests pass.
What it teaches: tool-use loops, structured output and function calling, context management across turns, sandboxed execution, and how to keep an agent from looping forever.
Milestones:
- Single tool call: model proposes a shell command, you run it, feed back stdout
- File read/write tools with a diff-based edit format
- Plan then act: a scratchpad of steps before editing
- Run tests, parse failures, retry with the error in context
- Guardrails: step budget, timeout, and a stop condition
- Score it on a small set of toy bugs (your own mini SWE-bench)
Ask: what are the failure modes of agentic systems and how do you bound them? You will have lived the answer.
RAG system (pi.dev) - retrieval over your own data, done properly▶
Build: a real retrieval-augmented generation system over your own corpus (notes, blog, papers), the kind you would ship on pi.dev.
What it teaches: chunking strategy, embeddings and vector search, hybrid dense+sparse retrieval, reranking, the "lost in the middle" problem, and how to actually evaluate retrieval quality instead of eyeballing it.
Milestones:
- Ingest + chunk (try fixed vs semantic vs recursive) and embed
- Dense vector search baseline, then add BM25 hybrid
- Add a reranker over the top-k
- Generation with citations grounded in retrieved chunks
- Eval harness: retrieval recall@k and answer faithfulness, not vibes
- Ship it behind a small API
Ask: when do you choose RAG over fine-tuning, and how do you measure a RAG system? You will have the numbers.
Ship one AI app end to end that people actually usestartup▶
Build: one complete AI product - not a notebook - that a real user can hit: a UI, a model behind an API, eval, and a deploy.
What it teaches: the unglamorous 90% that interviews probe: latency budgets, cost per request, prompt/version management, handling bad output, monitoring, and iterating on real usage.
Milestones:
- Pick a narrow problem with a clear user and a clear success metric
- Thin vertical slice: input - model - output, deployed
- Add eval and logging on every request
- Handle the long tail: refusals, timeouts, garbage input
- Measure cost and latency, then optimize the worst one
- Put it in front of real users and iterate on what breaks
Ask: startup rounds care that you can ship. A live URL beats any take-home.