Predictive Representation Learning at the Right Granularity

My initial work showed that JEPA-style prediction can be used as a regularizer alongside a primary learning objective, removing the need for a moving-average target network. This led to a broader question: rather than predicting between fixed views or individual tokens, can models learn the appropriate granularity of abstraction? Subsequent work explores randomized fractional views, boundary bottlenecks, and continuous concepts as mechanisms for concentrating information into representations of meaningful structure.