Skip to content
samit
Interviews/Computer Vision

Computer Vision

14 questions

14 questions
Detection & Segmentation
Two-stage (Faster R-CNN) vs one-stage (YOLO) detectors.▶

Two-stage (Faster R-CNN):

  • Stage 1: Region Proposal Network (RPN) generates ~2000 candidate bounding boxes
  • Stage 2: RoI pooling + classification + bbox regression for each proposal
  • Higher accuracy, especially for small/dense objects
  • Slower (~5 FPS on CPU, ~50 FPS with GPU)

One-stage (YOLO):

  • Divide image into S×S grid. Each cell predicts B boxes + class probabilities directly
  • Single forward pass → predictions
  • YOLOv8/v11: multi-scale prediction (FPN), anchor-free, ~100+ FPS
  • Lower accuracy on small/dense objects vs two-stage

DETR: Transformer-based. No anchors, no NMS. N learnable object queries attend to CNN features. Bipartite matching loss (Hungarian algorithm) for training. Cleaner but slower to train.

What is mAP (mean Average Precision)?▶

Standard metric for object detection evaluation.

IoU (Intersection over Union): overlap between predicted and ground-truth box. TP if IoU > threshold (e.g., 0.5).

Per-class Average Precision (AP):

  1. Sort all predictions by confidence score
  2. Compute precision and recall at each threshold (cumulative TP/FP)
  3. AP = area under the Precision-Recall curve (interpolated)

mAP: mean of AP across all classes.

COCO mAP: averages over IoU thresholds from 0.5 to 0.95 (step 0.05). More strict than PASCAL VOC (which uses IoU=0.5 only). mAP50 is the IoU=0.5 version.

How does SAM (Segment Anything Model) work?frontier▶

SAM (Kirillov et al., Meta 2023) is a promptable segmentation model that generalizes to any object category.

Architecture:

  • Image encoder: ViT-H extracts image embedding (16× downsampled)
  • Prompt encoder: encodes points, boxes, masks, or text as dense/sparse embeddings
  • Mask decoder: lightweight Transformer that attends between prompt embeddings and image embedding → outputs 3 mask candidates + quality scores

Training: SA-1B dataset - 1B masks on 11M images, mostly auto-generated via SAM's own predictions + human correction.

Prompts: click a point → segment object containing that point. Draw a box → segment object inside. Provide a mask → refine it. No fine-tuning needed for new objects.

SAM 2 (2024): extends to video segmentation with streaming memory.

Vision-Language & Multimodal
How does CLIP work? Why is it powerful?frontier▶

Contrastive Language-Image Pretraining (Radford et al., OpenAI 2021).

Training:

  • Batch of N (image, text) pairs from internet
  • Image encoder (ViT or ResNet) → image embeddings
  • Text encoder (Transformer) → text embeddings
  • Contrastive loss: maximize similarity of N correct pairs, minimize similarity of N²-N incorrect pairs
  • Trained on 400M (image, text) pairs

Zero-shot classification: embed image. Embed "a photo of a [class]" for each class. Predict class with highest cosine similarity. Works on ImageNet with ~76% accuracy - without any labeled training data!

Why powerful: internet text is rich in image descriptions. CLIP learns visual concepts by their linguistic descriptions, not predefined class labels. Generalizes to any concept that can be described in text.

Applications: image-text retrieval, zero-shot classification, image generation guidance (DALL-E uses CLIP), VQA foundation.

How does Whisper process audio? What happens to long audio?▶

Whisper (Radford et al., OpenAI 2022) is an encoder-decoder Transformer trained on 680K hours of diverse multilingual audio.

Preprocessing:

  • Resample to 16kHz mono
  • Compute log-Mel spectrogram (80 mel bins, 25ms windows, 10ms hop)
  • Pad or trim to exactly 30-second chunks - fixed context window

Architecture: Conv encoder (reduces temporal resolution) → Transformer encoder → Transformer decoder with cross-attention → token-by-token transcript

Long audio handling: split into 30-second segments (with overlap), transcribe each independently, merge transcripts. In JAX implementations: parallel processing of segments via vmap across batch dimension. Special tokens encode task (transcribe/translate), language, timestamps.

JAX arch question context: likely referring to the whisper-jax implementation that uses XLA compilation and sharding via pmap for GPU/TPU parallelism across chunks.

How does ViT tokenize images? Compare to CNNs.▶

Split 224×224 into 16×16 patches → 196 tokens. Flatten each patch (16×16×3=768), linear project to d_model, add position embeddings, run a Transformer encoder. [CLS] token (or global pool) for classification.

ViT vs CNN:

  • CNNs: locality + translation equivariance built in; data-efficient
  • ViT: weak inductive bias; needs large data (JFT-300M) but scales better - global context from layer 1

MAE: mask 75% of patches, reconstruct pixels - strong self-supervised pretrain on unlabeled images.

Explain the LLaVA architecture - how do vision encoders connect to LLMs?▶

LLaVA-style VLMs glue three pretrained pieces together:

1. Vision encoder (typically a frozen CLIP ViT) turns an image into a grid of patch embeddings.

2. Projection layer - a small linear layer or 2-layer MLP that maps the vision encoder's embedding space into the LLM's token embedding space, so each image patch becomes a "visual token" the LLM can attend to like any text token.

3. LLM (e.g. Vicuna, Llama) - receives the projected visual tokens concatenated with text tokens and generates a response autoregressively, same as a normal language model.

Training is staged:

Stage 1 (feature alignment): freeze both the vision encoder and the LLM, train only the projection layer on image-caption pairs so visual tokens land in a region of embedding space the LLM already understands.

Stage 2 (visual instruction tuning): unfreeze the LLM (often with LoRA) and fine-tune end-to-end on instruction-following data that mixes images and text, teaching the model to actually reason over visual content rather than just caption it.

Frontier CV
How do you use SAM at inference? Point, box, and mask prompts.frontier▶

Pipeline: ViT image encoder (runs once per image) → prompt encoder (points/boxes/masks) → lightweight mask decoder outputs 3 mask candidates + IoU quality scores.

  • Point: foreground (+) / background (−) clicks steer which object to segment
  • Box: tight bbox as prompt - often best for detection→segment pipelines
  • Mask: coarse mask input for refinement

Pick the mask with highest predicted IoU. No fine-tuning for new classes - prompt is the interface.

SAM 2 for video - how does it track objects across frames?frontier▶

SAM 2 adds a streaming memory module: past frame embeddings and mask predictions are stored in a compact memory bank. Each new frame attends to memory + current image features to propagate masks without re-encoding the whole video.

Enables interactive video segmentation (click frame 1, track through clip) and improves temporal consistency vs running SAM independently per frame.

Open-vocabulary detection - how do Grounding DINO and YOLO-World work?frontier▶

Detect arbitrary categories described in text at inference, not just fixed training classes.

  • Grounding DINO: fuses a text encoder with a DINO-style detector; cross-attention between language tokens and image features localizes phrases
  • YOLO-World: image-text pretrain + prompt-then-detect - embed class names, score boxes against text embeddings at runtime

Use when the label set changes often (retail SKUs, traffic signs) without retraining the full detector.

DINOv2 - why use it over supervised ImageNet pretrain?frontier▶

DINOv2 (self-distilled ViT on curated web images) produces general visual features without labels.

  • Strong linear probes and k-NN on downstream tasks with few labels
  • Better transfer to domains ImageNet doesn't cover (satellite, medical, industrial)
  • Typical workflow: frozen DINOv2 backbone + small task head; optional partial unfreeze with LoRA on last blocks
Video understanding - how do models capture temporal information?frontierstartup▶

Options:

  • 3D convolutions (I3D, SlowFast): explicit spatiotemporal filters - good inductive bias, heavier compute
  • Frame sampling + Transformer (TimeSformer, ViViT): treat frames or tubelets as tokens; divided space-time attention
  • VideoMAE: mask spacetime patches, reconstruct - self-supervised pretrain like MAE
  • SAM 2 / streaming memory: temporal propagation for segmentation/tracking

For surveillance/tracking : detection per frame + DeepSORT/ByteTrack often beats end-to-end video models on cost and debuggability.

When would you fine-tune SAM vs use zero-shot prompts vs train a U-Net?startup▶
  • Zero-shot SAM: novel objects, few labels, interactive labeling tool
  • Fine-tune SAM (LoRA on decoder): consistent domain (medical slices, satellite) where prompts alone miss boundaries
  • U-Net / dedicated seg model: fixed classes, need max FPS on edge, dataset >5K masks - simpler deploy than ViT-H
Contrastive learning for re-identification - what matters for multi-camera tracking?startup▶

Learn embeddings where same vehicle/person across cameras is closer than different identities.

  • Loss: triplet loss or supervised contrastive with batch-hard mining
  • Augmentation: strong color jitter (cameras differ), random erasing
  • At inference: cosine distance in gallery; combine with motion (Kalman) in DeepSORT

Metric quality (mAP on re-id benchmark) often matters more than detector mAP for end-to-end MOTA.