Technical Assessment Platforms

Academia is racing to build multimodal interview datasets — and says they're scarce

Source: ACM Multimedia / arXiv · Sep 23, 2026

Academic research is converging on exactly the kind of data an interview-focused business is built to produce. The ACM Multimedia AVI Challenge 2026 released a labeled dataset of 644 asynchronous video interviews, used to predict HEXACO personality traits and cognitive ability directly from candidates' recorded video responses. A separate effort, the Multi-TPC dataset, frame-aligns speech, motion capture, and gaze across three people interacting at once — a dense, synchronized multimodal record of real human interaction.

The most telling detail isn't any single result — it's the researchers' own framing. They repeatedly describe labeled, multimodal interview data as scarce and difficult to obtain. That scarcity is the whole point. Real recordings of people reasoning, explaining and reacting under interview conditions, captured across video, audio and behavioral channels and cleared for use, are genuinely hard to assemble — which is precisely what makes them valuable.

When independent academic groups are standing up their own multimodal-interview corpora and flagging how hard they are to source, it's strong outside validation that a consented, structured interview corpus sits on a real and growing demand curve — spanning both commercial model builders and the research community.

Key Points

  • The ACM Multimedia AVI Challenge 2026 released a labeled async-video-interview dataset of 644 subjects to predict HEXACO personality traits and cognitive ability from video responses
  • The Multi-TPC dataset frame-aligns speech, motion capture, and gaze across three simultaneous participants
  • Researchers explicitly call labeled multimodal-interview data scarce and hard to obtain
  • Direct external validation of the multimodal-interview-corpus thesis