From JEPA-style regularization to variable-granularity concepts in language models
My initial work showed that JEPA-style prediction can be used as a regularizer alongside a primary learning objective, removing the need for a moving-average target network. This led to a broader question: rather than predicting between fixed views or individual tokens, can models learn the appropriate granularity of abstraction? Subsequent work explores randomized fractional views, boundary bottlenecks, and continuous concepts as mechanisms for concentrating information into representations of meaningful structure.
Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA
First author · Co-authored with Yann LeCun
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
First author · Co-authored with Yann LeCun
Current Direction: Developing a unified account of perception and generation as prediction at different levels of abstraction.
Read more