Daniel Platnick, Marjan Alirezaie and Hossein Rahnama
Work for a Member organization and need a Member Portal account? Register here with your official email address.
May 7, 2026
Daniel Platnick, Marjan Alirezaie and Hossein Rahnama
Modern AI agents can plan, reflect, reason, and act over long-horizon digital tasks. Even so, they cannot plan to steer a social interaction in real time. Doing so requires anticipating how latent factors like trust and resistance will evolve under the agent’s actions, fast enough to plan within a conversational turn. Generative world models approach this by narrating possible futures, but autoregressive text generation is both too slow for real-time planning and fundamentally lossy. Representation-predictive methods can be superior, but are underexplored for interaction. We build LID-Bench, the first controlled testbed with oracle latent states for interaction dynamics, enabling systematic comparison of generative and representation-predictive world models against known dynamics. A bottleneck decomposition on three generative models spanning 117M–350M parameters reveals that all learn interaction dynamics internally, with text-rendering signal losses of 89–100% observed by horizon k=5. Our representation-predictive model (Social-JEPA, 125M params, 500K trainable) bypasses text entirely, forecasting latent dynamics 2.8–3.9×more accurately while performing rollouts 1,059–2,314×faster—making real-time planning within conversational turn-taking latencies feasible.