Efficient AI

Pruning / Sparsity (II)

Aditya Desai

Many thanks to all the resources listed in References on the main website. Almost all figures are taken from these resources; some were generated with Gemini. Some images are taken from Google Images and still need proper citation (in progress).

Logistics : Audit

  • Currently 8 Students are enrolled for auditing the class
  • Audit Students are expected to attend all classes (slack = 3 classes).
  • Sign in sheet from next class (TA)

Logistics: Seminar

  • Seminar starts on 21st August 2026
  • Seminar will happen on every Friday.
  • Groups will be created by TAs. We will announce the groups in the next class.
  • First Seminar is on Sparse Attention(Decoding)

Logistics: Exams

  • Exams are more of a tool to encourage engagement
  • Seminar is part of the syllabus. Projects are part as well.
  • If you are attending classes and following them, you should be good for exams

Logistics: Project

  • Hope you have started working on your projects!
  • Submission portal opens next week and first leaderboard will be released a week after that.
  • Assignment 1: Get on the leaderboard : Make a submission (no memory constraint)
  • Grade: 1% of the entire grade (taken from project )
  • Deadline: 14th August 2026

Logistics: Project Evaluation

  • Compression of Interest: $10\%, 20\%$ and $40\%$ of BF16 checkpoints
  • if a Submission has memory $k\%$ it goes in all heads that are larger than $k$

Hessian complexity

hessian_complexity.html

AutoML

How much to prune in each layer? Unstructured Pruning

    Natural alternatives,
    1. OBS, OBD : iterative pruning already chooses weights dynamically. Generally impractical for big networks.
    2. IMP: We can do global pruning by choosing weights across all layers. This can lead to degenerate pruning.
    3. Do pruning layer-wise
      • How to distribute the target sparsity across layers?

AutoML for Pruning

    AutoML for Pruning

use RL to learn sparsity predictors for each layer

Structured Sparsity

Unstructured Sparsity vs. Structured Sparsity

Unstructured Pruning

  • Quality?
  • Efficiency?

Quality

Quality
  • # possible patterns ( # subspaces for projection)
  • # Possiblity of finding a nearby projection

Efficiency?

Unstructured sparsity:

  • does not have contiguous memory access (overheads). i.e. low computational intensity
  • cannot use tensor cores.
  • 50% sparsity needs n/2 floats + index memory (n/2 log(n) bits)

Making Unstructured Sparsity Efficient?

Some work tries to make unstructured sparsity efficient:

When do we see a breakthrough for speed?

When do we see a breakthrough for speed?
  • One example from : Sparse GPU Kernels for Deep Learning Gale.et.al
  • Getting speedups is HARD
  • As GPU hardware improves, dense GEMMs are super optimized for. Gets even difficult.

Promise of Structured Sparsity

  • Reduce the "shape" of the workload
  • reuse existing optimized dense kernels
  • Challenge: Maintaining Quality
Structured Pruning

Pruning a Linear layer -- a hardware friendly way

GEMM Pruning - Linear Layer Hardware Friendly
  1. What is the algorithm for MM (matrix multiplication) on GPU?
  2. Why is tiling important?
  3. How would you prune?

Pruning a Linear layer : a hardware friendly way

GEMM Pruning - Linear Layer Hardware Friendly
  • Look at the algorithm where weight is used and make sure that the sparsity does not break contiguous memory access or add any overheads FLOPS reduction should lead to actual reduction of steps in existing kernels.

Pruning a Linear layer : a hardware friendly way

GEMM Pruning - Linear Layer Hardware Friendly
  • Neuron / Channel pruning: Remove columns/rows to reduce the shape of the workload
  • Block Sparsity: Ability to skip certain "blocks" in computing the matrix multiplication — block-wise sparsity

Pruning a convolution : a hardware friendly way

Convolution Pruning - Hardware Friendly

Pruning a convolution : a hardware friendly way

Im2col and GEMM implementation of convolution
  • Im2col and GEMM implementation of convolution (#channels = 1)

Pruning a convolution : a hardware friendly way

Convolution GEMM implementation of convolution

How to prune? (General Recipe for Groups)

  1. Train a dense network to convergence.
  2. Compute the saliency score of each group of weights
  3. Remove group with the smallest saliency scores (or multiple, with / without adjustment of others)
  4. Retrain the pruned network. Return to step 2 (iterate until target sparsity).

How to prune?

  • Extend OBS (sim. OBD) to get saliency and update rule for groups of weights? (Exercise)
  • Iterative Magnitude Pruning (IMP) can be extended to groups of weights through l_p norms
  • The $\ell_p$ norm of a weight vector $w$ is defined as:
    $\displaystyle \| w \|_p = \left( \sum_{i=1}^{n} |w_i|^p \right)^{1/p}$
  • While training, you can use group wise $l_p$ regularizer to induce group sparsity.

Results from the paper

“Learning Structured Sparsity in Deep Neural Networks” by Wen et al.

Unstructured sparsity does not show speedups on CPUs / GPUs

AlexNet GPU speedups from unstructured sparse matrices

Even at high sparsity, irregular sparse execution can be slower than dense GEMM because storage and indexing overheads dominate.

Structured pruning for CNNs

Filter-wise, channel-wise, shape-wise, and depth-wise structured sparsity (Exercise) In GEMM Implementation how do these pruning granularities reduce workload?

Experimental setup

  • Three-step scheme:
    1. Train with group regularization.
    2. Prune groups with the lowest magnitude.
    3. Fine-tune the remaining network.
  • Reported speedups are layer-wise and measured on CPU unless specified.

LeNet: filters, channels, and filter shapes

Table 1 — Penalizing unimportant filters and channels
LeNet # Error Filter # Channel # FLOP Speedup
1 (baseline)0.9%20–501–20100%–100%1.00×–1.00×
20.8%5–191–425%–7.6%1.64×–5.23×
31.0%3–121–315%–3.6%1.99×–7.44×
Table 2 — Learning filter shapes
LeNet # Error Filter size Channel # FLOP Speedup
1 (baseline)0.9%25–5001–20100%–100%1.00×–1.00×
40.8%21–411–28.4%–8.2%2.33×–6.93×
51.0%7–141–11.4%–2.8%5.19×–10.82×

Structured groups preserve LeNet accuracy while reducing layer FLOPs substantially and reaching up to 10.82× measured speedup.

MLP and ConvNet structured sparsity

Figure 4(a) — Learned neurons in MLP
MLP # Error Neurons per layer FLOP per layer
1 (baseline)1.43%784–500–300–10100%–100%–100%
21.34%469–294–166–1035.18%–32.54%–55.33%
31.53%434–174–78–1019.26%–9.05%–26.00%
Table 3 — ConvNet on CIFAR-10
ConvNet # Error Row sparsity Column sparsity Speedup
1 (baseline)17.9%12.5%–0%–0%0%–0%–0%1.00×–1.00×–1.00×
217.9%50.0%–28.1%–1.6%0%–59.3%–35.1%1.43×–3.05×–1.57×
316.9%31.3%–0%–1.6%0%–42.8%–9.8%1.25×–2.01×–1.18×

Learned structure cuts MLP FLOPs and accelerates ConvNet layers without degrading—and sometimes improving—test error.

AlexNet on ILSVRC 2012

Table 4 — Sparsity and speedup
#MethodTop-1 err.Statistic conv1conv2conv3conv4conv5
1ℓ₁44.67% sparsity67.6%92.4%97.2%96.6%94.3%
CPU ×0.802.914.843.832.76
GPU ×0.250.521.381.041.36
2SSL44.66% column sparsity0.0%63.2%76.9%84.7%80.7%
row sparsity9.4%12.9%40.6%46.9%0.0%
CPU ×1.053.376.279.734.93
GPU ×1.002.374.944.033.05
3pruning [7]42.80% sparsity16.0%62.0%65.0%63.0%63.0%
4ℓ₁42.51% sparsity14.7%76.2%85.3%81.5%76.3%
CPU ×0.340.991.301.100.93
GPU ×0.080.170.420.300.32
5SSL42.53% column sparsity0.00%20.9%39.7%39.7%24.6%
CPU ×1.001.271.641.681.32
GPU ×1.001.251.631.721.36
Caption: pruning [7] refers to Iterative Magnitude Pruning (IMP) from Han et al. 2015.
L1 applies l1-norm based regularization during training to encourage sparsity.
SSL refers to Structured Sparsity Learning (this work).
row-wise and column-wise sparsity = filter-wise and shape-wise sparsity (w.r.t GEMM)