← Back to Projects

SocialPulse: Real-World Social Interaction Sensing

NIMH R01MH132138 3Cavaliers Seed Grant Commonwealth Cyber Initiative TYDE Seed Grant

Leveraging privacy-preserving wearable sensing and multimodal AI to detect and characterize social interactions as they unfold in everyday life.

Illustration of a smartwatch showing a speech waveform and an interaction indicator, with sensing pulses reaching two people who are having a conversation with speech bubbles.
SocialPulse runs directly on a smartwatch, detecting in-person and virtual social interactions in everyday life while processing audio on the device to protect privacy.

Description

Social interactions are fundamental to health and well-being, yet they remain difficult to measure as they naturally occur in daily life. SocialPulse develops privacy-preserving wearable sensing and multimodal AI methods to detect and characterize social interactions in real-world settings, including in-person, virtual, and hybrid interactions. Our smartwatch-based system processes audio directly on the device rather than transmitting raw recordings to a phone or the cloud, helping protect the privacy of users and the people around them. Through naturalistic studies, we develop and evaluate models that identify when social interactions occur and characterize their context and dynamics. This work provides new ways to measure social behavior at scale and creates a foundation for personalized, context-aware systems that can better understand and respond to an individual's social environment.

We present an on-watch interaction detection system designed to capture diverse interactions in naturalistic settings. A core component is a foreground speech detector trained on a public dataset. Evaluated on over 100,000 labeled foreground speech and background sound instances, the detector achieves a balanced accuracy of 85.51%, outperforming prior work by 5.11%.

We evaluated the system in a real-world deployment (N=38), with over 900 hours of total smartwatch wear time. The system detected 1,691 interactions, 77.28% were confirmed via participant self-report, with durations ranging from under one minute to over one hour. Among correct detections, 81.45% were in-person, 15.7% virtual, and 1.85% hybrid. We further developed a 15-second window-level audio-only model that enables faster interaction prediction, achieving a balanced accuracy of 90.39% and a sensitivity of 91.01% on 33,698 labeled windows. These results demonstrate the feasibility of real-world interaction sensing and open the door to adaptive, context-aware systems responding to users' dynamic social environments.

Publications

Team