The homework requires Python 3.12. You do not need to install it yourself: the homework ships with a pyproject.toml pinned to 3.12, and uv sync (below) downloads that interpreter if it is missing. To confirm, run uv run python --version inside the homework directory after syncing.
We recommend using uv as it's much faster than pip and conda for managing Python environments and packages.
uv is a modern, Rust-based package + project manager for Python. It keeps the familiar pip workflow but re-implements the engine for speed and reliability. Concretely: it creates a venv, resolves and installs dependencies with its own fast installer, and deduplicates files via a global cache (copy-on-write on macOS, hardlinks on Linux/Windows). It can also manage Python versions per project (e.g., pin 3.12) so each assignment uses a clean, reproducible interpreter. Think "pip + virtualenv + pip-tools + pyenv/pipx".
Please refer to the official uv installation documentation for the most up-to-date installation instructions for your platform.
Create and activate a virtual environment with the required dependencies:
# Install uv once curl -LsSf https://astral.sh/uv/install.sh | sh # Optional: `uv` binary by default goes to `$HOME/.local/bin` on Linux/macOS, # so you may need to add it to your PATH (uv may have done this for you): export PATH="$HOME/.local/bin:$PATH"
# Install uv once powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex" # Then open a NEW terminal window so that `uv` is on your PATH
# Download the homework zip and unzip into `hw1_foundations/` # In your hw directory (it ships with a pyproject.toml pinned to Python 3.12) uv sync # Create .venv, fetch Python 3.12 if needed, install numpy + einops uv run python grader.py # Run the local grader
If you cannot run the assignment on your laptop or need additional computing resources, Stanford provides FarmShare, a community computing environment for coursework and unsponsored research. Please follow the instructions at https://docs.farmshare.stanford.edu/ to get started with the computing environment.
Welcome to your first CS221 assignment!
The goal of this assignment is to sharpen your math, programming, and ethical analysis skills
needed for this class. If you meet the prerequisites, you should find these
problems relatively innocuous. Some of these problems will occur again
as subproblems of later homeworks, so make sure you know how to do them.
If you're unsure about them or need a refresher,
we recommend going through our prerequisites module or other resources on the Internet,
or coming to office hours.
Before you get started, please read the Homeworks section on the course website thoroughly.
We've created a LaTeX template (hw1_foundations_template.tex in the same folder as this homework) for you to use that contains the prompts for each question.
Linear algebra forms the foundation of modern AI and machine learning. In this problem, you'll work with vectors and matrices using NumPy, which is the standard library for numerical computing in Python and provides efficient implementations of vector and matrix operations that are essential for AI. Understanding these operations is crucial for implementing neural networks, optimization algorithms, and data processing pipelines. For example, the bulk of modern LLMs are just dense matrix multiplications, and NumPy is the first step towards being able to manipulate these matrices. You'll practice basic operations like dot products, matrix multiplication, and distance calculations that appear everywhere in machine learning algorithms.
You'll also learn about Einstein summation notation (einsum), which we'll use through the einops library. Einsum notation allows you to express complex array/matrix/tensor operations concisely, including matrix multiplications, tensor contractions, array reshaping, and pretty much everything used in the modern AI models. einops is a library that allows you to implement einsum notation in Python (from einops import einsum, rearrange). Understanding einsum allows you to code more readable code. If you're new to einsum, this YouTube video and the einops basics tutorial may be helpful.
einsum string for $X\mathbf w$; (ii) for the pairwise dot-product matrix $XX^\top$; (iii) for $\operatorname{diag}(X^\top X)$ (column-wise squared norms). For each string, briefly explain why it computes the requested quantity.
einops format (e.g., 'n d, d m -> n m') rather than the np.einsum format (e.g., 'nd, dm -> nm').einsum (from einops) for the matrix multiplication and NumPy broadcasting for the bias; do not use Python loops.
linear_project(x, W, b) in submission.py with shapes: x:(B,D_in), W:(D_in,D_out), b:(D_out,) → (B,D_out).einops.rearrange pattern string that performs this reshape. Do not perform the reshape yourself; just return the pattern string. It is fine to assume D % G == 0.
split_last_dim_pattern() in submission.py returning a rearrange pattern string. The autograder will apply it with the appropriate g=num_groups (note that the group axis must be named exactly g).einsum (from einops) without loops. If normalize=True, divide the result by $\sqrt{D}$.
normalized_inner_products(A, C, normalize=True) in submission.py with shapes: A:(B,M,D), C:(B,N,D) → (B,M,N).-np.inf, leaving other entries unchanged. Construct the mask using NumPy broadcasting; do not use loops. Do not use np.triu, np.tril, np.triu_indices, np.tril_indices, or similar built-in triangular helpers; build the mask yourself from index arrays (e.g., np.arange) and broadcasting.
mask_strictly_upper(scores) in submission.py for scores:(B,L,L), returning a masked array of the same shape.einops.einsum pattern string that computes the weighted sums $out[b,:] = \sum_{j=1}^N P[b,j]\, V[b,j,:]$, yielding $out \in \mathbb{R}^{B\times D}$.
prob_weighted_sum_einsum() in submission.py that returns the einsum string; the autograder will apply it to P and V.
Gradients are essential for training machine learning models through optimization algorithms like gradient descent. In this problem, you'll practice computing gradients analytically and verify your results using numerical methods (finite differences).
The textbook for MATH 51 may be useful for the gradient problems here, specifically the sections "Gradients, Local Approximations, and Gradient Descent" (p. 209).
gradient_warmup(w, c) in submission.py that takes vectors w and c and returns the gradient vector.np.ones or np.repeat functions to create the gradient matrices.
matrix_grad(A, B) in submission.py that returns a tuple (grad_A, grad_B) where each gradient matrix has the same shape as the corresponding input matrix.Image from Wikipedia
np.eye(d) function to create the unit vectors. Note that b has shape (n,), so it may need an extra axis (b[:, None]) to broadcast against an (n, d) array.
submission.py, implement least_squares_grad(w, A, b) (analytic gradient) and least_squares_finite_diff_grad(w, A, b, epsilon=1e-5) (central-difference gradient). The autograder will compare them on random instances.
Optimization is central to AI - we cast many AI problems as finding the best solution in a rigorous mathematical sense. In this problem, you'll work with analytical optimization techniques and implement them using NumPy to verify your mathematical solutions computationally.
The programming component will help you understand how theoretical optimization translates to practical implementations. You'll derive the minimizer of a weighted quadratic analytically and then use gradient descent to find it numerically.
The textbook for MATH 51 may be useful for the optimization problems here, specifically the sections "Maxima, Minima, and Critical Points" (p. 186).
gradient_descent_quadratic(x, w, theta0, lr, num_steps) in submission.py to minimize the scalar objective
$f(\theta) = \sum_{i=1}^n w_i (\theta - x_i)^2$, where $\theta\in\mathbb{R}$ (recall problem 3a).
The function should return the final scalar iterate after num_steps gradient steps.
num_steps iterations of gradient descent starting from theta0 with learning rate lr.
One of the goals of this course is to teach you how to tackle real-world problems with tools from AI. But real-world problems have real-world consequences. Along with technical skills, an important skill every practitioner of AI needs to develop is an awareness of the ethical issues associated with AI. The purpose of this exercise is to practice spotting potential ethical concerns in applications of AI - even seemingly innocuous ones.
In this question, you will explore the ethics of four different real-world scenarios using the ethics guidelines produced by a machine learning research venue, the NeurIPS conference. The NeurIPS Ethical Guidelines list seventeen non-exhaustive concerns under General Ethical Conduct and Potential Negative Social Impacts (the numbered lists). For each scenario, you will write a potential negative impacts statement. To do so, you will first determine if the algorithm / dataset / technique could have a potential negative
social impact or violate general ethical conduct (again, the seventeen numbered items taken from the NeurIPS Ethical Guidelines page). If the scenario does violate ethical conduct or has potential negative social impacts, list one concern it violates and justify why you think that concern applies to the scenario. If you do not think the scenario has an ethical concern, explain how you came to that decision.
Unlike earlier problems in the homework there are many possible good answers. If you can justify your answer, then you should feel confident that you have answered the question well.
Each of the scenarios is drawn from a real AI research paper; you should think about why the researchers may have chosen for the algorithms to behave in the way as described in the scenario. The ethics of AI research closely mirror the potential real-world consequences of deploying AI, and the lessons you'll draw from this exercise will certainly be applicable to deploying AI at scale. As a note, you are not required to read the original papers, but we have linked to them in case they might be useful. Furthermore, you are welcome to respond to anything in the linked article that's not mentioned in the written scenario, but the scenarios as described here should provide enough detail to find at least one concern.
Submission is done on Gradescope.
Written: When submitting the written parts, make sure to select all the pages
that contain part of your answer for that problem, or else you will not get credit.
To double check after submission, you can click on each problem link on the right side, and it should show
the pages that are selected for that problem.
Programming: After you submit, the autograder will take a few minutes to run. Check back after
it runs to make sure that your submission succeeded. If your autograder crashes, you will receive a 0 on the
programming part of the assignment. Note: the only file to be submitted to Gradescope is submission.py.
More details can be found in the Submission section on the course website.
[3] Imperial College London. Loan Default Prediction Dataset. 2014.
[4] Caliskan-Islam et. al. De-anonymizing programmers via code stylometry. 2015.