Pub-AI: AI in Science

Video World Models · 2026-09-30

Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as, Chirag Raman

arXiv:2609.37708PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

BRAID (Bilevel Representations for Agent Interaction Dynamics) is a hierarchical sequential latent-variable model for generative multi-person social motion. It uses a group latent for shared interaction dynamics and person latents for individual behavior conditioned on the evolving group context, enabling coherent generation and compact social-state vectors for full, sparse, or partial observations.

Why it matters

The paper targets a gap in social motion modeling: interaction state is often implicit, limiting transfer across groups, tasks, and partial-observation regimes. BRAID makes group interaction explicit via a reusable group/person latent hierarchy trained as a meta-transfer learning problem across datasets, and evaluates forecasting, tracking/in-filling, and response generation with interaction-structure metrics (synchrony, temporal alignment, interpersonal coordination) rather than reconstruction alone.

Method

  • Hierarchical latent sequential model: a group latent z_t^g captures shared interaction dynamics (phase, tempo, proxemics) and person latents {z_t^p} encode individual behavior conditioned on the group latent.
  • Meta-transfer learning: shared interaction priors learned across conversational, dance, and boxing data, then adapted through arbitrary context sets of observed people and joints.
  • Inference via amortised variational posterior with precision-weighted (product-of-experts) fusion of bottom-up evidence (full target data) and top-down prior (context/latent history), enabling generation under sparse or absent context.
Abstract (from arXiv)

Human social behaviour is not a collection of independent motions, but a jointly organised process in which group dynamics and individual variation continuously shape one another. Yet existing social motion models often prioritise plausible trajectories while leaving interaction state implicit, limiting their ability to transfer across groups, tasks, and partial-observation regimes. To address this gap, we introduce Bilevel Representations for Agent Interaction Dynamics (BRAID), a hierarchical sequential latent-variable model for generative multi-person interaction. BRAID explicitly formulates social motion generation as a meta-transfer learning problem: shared interaction priors are learned across datasets and adapted through arbitrary context sets of observed people and joints. The model represents each scene through a group-level latent state that captures shared interaction dynamics and person-level latent states that capture individual behaviour conditioned on the evolving group context. This modelling choice enables coherent generation under full, sparse, or partial observations while exposing compact social-state vectors that can serve as an interface for downstream embodied-agent systems. We evaluate BRAID under a unified SMPL-based representation on social forecasting, tracking and in-filling, and response generation, using metrics that assess not only reconstruction accuracy but also realism, diversity, temporal alignment, and interpersonal coordination. We further analyse the hierarchical latent space, showing that it captures separable group- and individual-level structure.

Related papers