On a more boring note... I could imagine using this pattern for documentation screenshots. It would be a neat sanity check during PR reviews to show whichever screenshots were changed by that PR in context with their surrounding docs.
Maybe they're meaningful, intended, changes which could be useful context for a reviewer. Maybe they're broken, indicating not only that you now need a different derivation to get the desired screenshot, but also that your prose probably needs to change (since the path to get from start to the screenshot has now also changed).
The underlying principle of translating those implicit temporal dependencies into explicit data dependencies very well might have applications in optimizing high-accuracy computer/console emulators. The usual implementations iterate over some variable that represents the passage of time and compute all possible side effects, which is expensive. Many compilers already do a related static optimization by eliminating intermediate variable mutations that can be proven to have no side effects.
(also, if you're not familiar with the "tool-assisted speedrun/superplay" tradition of optimizing video game input sequences as a flavor of metaprogramming, see https://tasvideos.org/NewcomerCorner)
For someone packaging things with Nix, it's tempting to think of derivations as just a builder - you write your scripts, accept your inputs, and your result ends up in the store.
The article suggests, taking advantage of lazy evaluation, that you can put a whole state machine in Nix (i.e., lift the DSL for describing states into the Nix language). If you had a team of developers that need various customizations of something, they could describe it in Nix, and benefit from the caching of intermediate steps. If someone's just changing something at the end of the long string of customizations, they only need to build their changes.
A more concrete, if very niche, example - Back in the days when we were all training our own deep learning models, my company had assembled a list of facts for every product we could find on the internet. Our data scientists wanted to be able to test datasets with different aggregates of the lists - we needed to both build the training set, and compute it at runtime for new products. There was some compositionality to the aggregates - one run might use the mean, another would subtract the mean from every element, etc. It was time consuming and annoying to manage. One of our developers made a DSL in Haskell to describe the set of aggregates desired for a given dataset, and needed to build caching for each column. The technique the article suggests could have done this out of the box (at build time, even) - build a function and a vector for each aggregate, throw them in the nix store, and then wrap them up as a training set and a server.
I gave 3 talks that I will update on my site[2] once they released.
[1]: https://nix.vegas [2]: https://fzakaria.com/talks