one example of is flash attention - in torch, flash attention is just implemented as one big forward pass and one big backward pass. it's not broken into intermediate operations. initially, tanh here was implemented as one forward pass and one backward pass.
any time you want to inspect the intermediate results is a good time to not chunk it. squishing operators together saves hbm reads and writes.