Selected Publications

I'm interested in Computer Vision and Generative AI. My current research spans controllable image and video generation and editing, Vision Language Models, video understanding, reinforcement learning and model evaluation. During my Ph.D., I did research in Open-Set Recognition and Human Cognitive Science.

ACaM: Natural Language Camera Movement Understanding

Yuwen Tan, Joey Huang, Jin Huang, Haoxiang Li, Boqing Gong

ECCV, 2026  ·  project page / arXiv

We show that existing vision-language models frequently confuse camera movement with object movement, translation with rotation, and left with right. We introduce a two-level cinematographic taxonomy and an atomic benchmark of real and synthetic videos, and curate a large-scale training set with targeted camera-movement augmentation. Our fine-tuned 8B model outperforms Gemini 3.1 Pro by 10% and 11% on real and synthetic videos respectively, though a substantial gap to human performance remains.

Extreme Value Theory for Modeling Category Decision Boundaries in Visual Recognition

Jin Huang, Deeksha Arun, Terrance Boult, Walter Scheirer

bioRxiv

Category learning models vary in how they place decision boundaries relative to human behavior. We find evidence that Extreme Value Theory, which preferentially weights the extremes of a distribution, better predicts human category judgments than central-tendency models across two experiments with line stimuli and face morph sequences, offering new insight into how discriminative information is encoded for decision making.

Analysis of Human Perception in Distinguishing Real and AI-Generated Faces: An Eye-Tracking Based Study

Jin Huang, Subhadra Gopalakrishnan, Trisha Mittal, Jake Zuena, Jaclyn Pytlarz

FG, 2025  ·  paper

We investigate how humans perceive and distinguish real faces from AI-generated ones through a perceptual experiment using eye-tracking technology. Analyzing StyleGAN-3 generated images, we find that participants distinguish real from fake faces with an average accuracy of 76.80%, and that they scrutinize images more closely when they suspect a fake.

Human Activity Recognition in an Open World

Derek Prijatelj, Sam Grieggs, Jin Huang, Walter Scheirer, et al.

JAIR, 81 (2024) 85-122  ·  paper / arXiv

Managing novelty in perception-based human activity recognition (HAR) is critical for realistic, real-world settings. We formalize novelty for HAR, propose an incremental open world learning (OWL) protocol applied to the Kinetics datasets to build a new benchmark (KOWL-718), analyze how current state-of-the-art HAR models perform as novelty is introduced over time, and release a containerized pipeline for reproducing and extending the protocol.

Measuring Human Perception to Improve Open-Set Recognition

Jin Huang, Derek Prijatelj, Justin Dulay, Walter Scheirer

T-PAMI, 2023  ·  paper / arXiv

Human perception, as measured through visual psychophysics, currently outperforms machine models at recognizing novelty in visual recognition tasks. We run a large-scale experiment collecting over 900,000 human reaction time measurements and use them to design a novel regularization term for the Open-Set Recognition loss function.