Resources

JEPA Hub

JEPA is an approach to training an AI model — similar to the way generative or diffusion models are approaches to training. JEPA isn't a product or even a specific technology implementation. There will be many, many products and companies built that train their models as JEPA models. While we are the first to bring a product built on a JEPA-style model to market commercially, we won't be the last. This hub is one place where anyone can learn about JEPA more broadly and why it represents the best path to what comes after LLMs.

Joint Embedding Predictive Architecture is the research lineage Primate Intelligence builds on. Papers, companies, analysis, and talks on world models that predict in representation space instead of pixel space.

Details on Primate's implementation of a JEPA model can be found on our technology, product and blog pages.

Papers 17

The JEPA family, in order of publication.

A Path Towards Autonomous Machine Intelligence2022

Yann LeCun

The founding position paper. Introduces the Joint Embedding Predictive Architecture (JEPA) as the centerpiece of a broader argument for world-model-based AI over pure autoregressive generation.

Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA)2023

Assran, Duval, Misra, Bojanowski, Vincent, Rabbat, LeCun, Ballas

The first working JEPA model. Learns image representations by predicting masked-patch embeddings in latent space instead of reconstructing pixels — no hand-crafted augmentations.

MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features2023

Bardes, Ponce, LeCun

Unifies optical flow estimation and content-feature learning in a single shared encoder trained with the JEPA objective.

V-JEPA: The Next Step Toward Advanced Machine Intelligence2024

Bardes, Garrido, Ponce, Chen, Rabbat, LeCun, Assran, Ballas

Extends JEPA to video: learns by predicting masked spatio-temporal regions of a clip in representation space, producing strong off-the-shelf video features with no labels or fine-tuning.

A-JEPA: Joint-Embedding Predictive Architecture Can Listen2023

Fei, Fan, Yu, Zhu, Sun, Jiang

Carries the JEPA masked-latent-prediction principle from images into the audio spectrogram domain.

Time-Series JEPA for Predictive Remote Control under Capacity-Limited Networks2024

Bariah, Zeydan, et al.

Applies JEPA-style latent prediction to sensor time-series, compressing high-dimensional signals into low-dimensional semantic embeddings for bandwidth-constrained remote control.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning2025

Assran, Bardes, Fan, Garrido, et al. (Meta FAIR)

Scales V-JEPA to 1M+ hours of internet video and adds an action-conditioned variant (V-JEPA 2-AC) that plans robot manipulation zero-shot in new environments from a small amount of unlabeled robot footage.

Joint Embeddings Go Temporal (TS-JEPA)2025

various

A JEPA architecture adapted specifically for general time-series representation learning, extending the masked-latent-prediction recipe beyond vision and audio.

Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving2026

various

Integrates V-JEPA representations with multimodal trajectory distillation to address mode collapse in end-to-end autonomous driving policies.

ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning2025

Vujinovic, et al.

Unifies imitation learning and self-supervised learning by jointly predicting action sequences and latent observation sequences — up to 40% improvement in world-model understanding and 10% higher task success vs. the strongest baseline. Published in IEEE Access.

LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures2025

Huang, Balestriero, et al.

Asks whether the JEPA recipe that works so well for vision can improve language model training. LLM-JEPA outperforms standard LLM training objectives across Llama3, OpenELM, Gemma2, and Olmo, while being robust to overfitting — direct evidence embedding-space objectives beat input-space reconstruction outside vision too.

VL-JEPA: Joint Embedding Predictive Architecture for Vision-Language2025

Chen, Shukor, et al. (Meta)

Vision-language model that predicts continuous text embeddings instead of autoregressively generating tokens. Beats CLIP, SigLIP2, and Perception Encoder on 16 video classification/retrieval benchmarks with 50% fewer trainable parameters, and matches InstructBLIP/QwenVL on VQA at just 1.6B params — the clearest evidence yet that JEPA-style training generalizes beyond pure vision into multimodal understanding.

Value-Guided Action Planning with JEPA World Models2026

Destrade, et al.

Shapes a JEPA world model's representation space so the goal-conditioned value function approximates a distance between state embeddings, significantly improving planning performance over standard JEPA on control tasks. Presented at the World Modeling Workshop 2026 (Mila).

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?2026

Terver, et al. (Meta FAIR)

Systematic ablation of architecture, training objective, and planning algorithm for JEPA-based world models (JEPA-WMs) — combines findings into a model beating both DINO-WM and V-JEPA-2-AC baselines on navigation and manipulation. Accepted at TMLR.

Causal-JEPA (C-JEPA): Learning World Models through Object-Level Latent Masking2026

Nam, et al.

Extends masked joint-embedding prediction from image patches to object-centric representations, forcing interaction-dependent prediction instead of shortcut solutions. ~20% absolute gain on counterfactual VQA reasoning and 100x fewer latent features needed for comparable planning performance. Accepted at ICML 2026.

LeWorldModel (LeWM): Stable End-to-End Joint-Embedding Predictive Architecture from Pixels2026

Maes, Balestriero, et al.

First JEPA that trains stably end-to-end from raw pixels with just two loss terms (down from six), using only ~15M params trainable on a single GPU. Plans up to 48x faster than foundation-model-based world models while staying competitive on 2D/3D control — and its latent space reliably detects physically implausible events.

A Generalization Theory for JEPA-Based World Models2026

Cui, et al.

First formal generalization theory for JEPA world models — frames JEPA pretraining as conditional spectral graph learning, connects pretraining error to downstream planning regret, and proves a finite-sample generalization bound. Theoretical grounding for why latent-space prediction beats input-level prediction.

World Model Landscape: JEPA vs. Other Approaches 3

The definitive long-form pieces mapping where JEPA sits against video diffusion, symbolic, and spatial-3D approaches.

Blogs & Thought Leadership 5

Essential reading on JEPA and world models.

Videos & Talks 9

Yann LeCun and others making the case for world models over LLMs, on camera.

Yann LeCun's $1B Bet Against LLMs [Part 1]2026

Welch Labs

The clearest public explainer of why LeCun raised ~$1B for AMI Labs to bet against autoregressive LLMs — walks through what a JEPA world model is, how it differs from next-token prediction, and why he thinks pixel-level generation is a dead end. Start here.

Yann LeCun's $1B Bet Against LLMs [Part 2]2026

Welch Labs

Follow-on to Part 1, going deeper on the JEPA architecture itself — energy-based self-supervised learning, representation-space prediction, and the hierarchical planning story LLMs can't do.

Joint-Embedding Predictive Architecture (JEPA): Self-Supervised World Modeling2025

AI, Career Growth and Life Hacks

Focused walkthrough of the JEPA architecture as self-supervised world modeling — the core mechanics of predicting in representation space instead of pixel space.

Why Meta's VL-JEPA Destroys All LLMs2026

Better Stack

Covers VL-JEPA, Meta's vision-language JEPA variant, and how it diverges from standard LLM architectures under LeCun's guidance.

Beyond LLMs: JEPA and the Road to AGI — the main milestones so far2026

Turing Post TV

Traces the JEPA roadmap from self-supervised learning through to LeCun's AGI milestones — a useful timeline view of how the architecture evolved.

V-JEPA: Revisiting Feature Prediction for Learning Visual Representations from Video (Explained)2024

Yannic Kilcher

Kilcher's paper-explainer breakdown of the original V-JEPA paper — unsupervised representation learning from video via feature prediction alone, no pixel reconstruction.

Yann LeCun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI2024

Lex Fridman Podcast #416

2.5-hour deep technical conversation. Timestamped sections on the limits of autoregressive LLMs, video prediction, JEPA, JEPA vs. LLMs, and hierarchical planning — the most complete public explainer of LeCun's reasoning on video.

Yann LeCun on What Comes After LLMs2026

Redpoint Ventures (with Jacob Effron)

LeCun's clearest post-Meta statement: LLMs are "a dead-end on the path to human-level intelligence" despite being useful products, because they don't build a world model. Covers why he left Meta and what AMI Labs is betting on instead.

The Next Phase of Artificial Intelligence2026

Bloomberg (Yann LeCun & JP Vert)

LeCun and JP Vert discuss how LLMs translate — or fail to translate — into physical-world understanding, and the infrastructure world models will need.