PyTorch / TensorFlow
The two dominant deep learning frameworks — used to build, train, and deploy neural networks for computer vision, NLP, and generative AI applications.
PyTorch and TensorFlow are the two foundational deep learning frameworks that power virtually all production neural network development. PyTorch (Meta) has become the dominant framework in research and is rapidly taking over industry; TensorFlow (Google) remains widely deployed in production systems and mobile (TensorFlow Lite). Together they underpin the AI revolution — large language models, image generators, speech recognition, and recommendation systems are all built on one or both. ML Engineer, AI Engineer, and senior Data Scientist roles almost universally require fluency with at least one, and increasingly both.
Typical time to job-readiness: ~4 months.
Learning PyTorch / TensorFlow
Beginner
Start with PyTorch — it is the more Pythonic of the two and dominant in research. Follow the official PyTorch tutorials and build a basic image classifier on MNIST or CIFAR-10. Understand tensors, autograd, and the training loop (forward pass, loss, backward pass, optimizer step).
Intermediate
Build CNNs for vision and transformer-based models for text, learn transfer learning with pretrained models from HuggingFace, and use PyTorch Lightning to structure training code cleanly. For TensorFlow, learn the Keras high-level API — it handles most production use cases without dropping to the raw TF layer.
Advanced
Custom training loops, mixed precision training, distributed training across GPUs with PyTorch DDP or TensorFlow's MirroredStrategy, model optimization for inference (quantization, pruning, ONNX export), and MLflow or Weights & Biases for experiment tracking. ML engineer interviews include both coding (implement a training loop, explain backprop) and system design (design a model serving pipeline for 10M daily users).
Key concepts
- Tensors: multi-dimensional arrays that represent data and model parameters; the core data structure in both frameworks
- Autograd (automatic differentiation): computes gradients automatically for backpropagation
- Training loop: forward pass (predictions) → loss computation → backward pass (gradients) → optimizer step (update weights)
- Transfer learning: start from a pretrained model (ResNet, BERT) and fine-tune on your data — standard for most production tasks
- PyTorch's define-by-run (eager) vs TensorFlow's graph mode — PyTorch is more Pythonic; TF's graph mode is faster for production
- Model serving: exporting trained models for inference (ONNX, TorchScript, TF SavedModel) and serving via REST or gRPC
Common interview topics
- Explain how backpropagation works in neural networks
- What is the difference between PyTorch and TensorFlow — when would you choose each
- How does transfer learning work and why is it preferred over training from scratch
- What is a gradient and how does an optimizer use it to update weights
- Design a system to serve a PyTorch model to 1 million users per day