A new preprint analyzes hidden-state dynamics in eight locally instrumented small open Transformer models to determine if internal representations follow a structured progression during inference. The study distinguishes between changes occurring across model depth and those evolving over autoregressive generation time.

  • Local ordering across model depth was significantly more structured than random permutations (p = 0.00019996) and held in all leave-one-model-out checks.
  • Cross-model depth profiles showed high coherence with a mean correlation of approximately r = 0.789.
  • Functionally labelled events were statistically associated with normalized layer depth (p = 0.0024).
  • A previously observed common temporal pattern failed to replicate in the expanded panel, surviving 0/8 leave-one-model-out checks.
  • Models with similar functional outcomes or from the same architecture family did not show significant structural similarity.

The results suggest that Transformer inference contains reproducible structure along depth while remaining highly conditional in time and behavior, challenging the idea of a universal temporal dynamic.