New webcam-based AI system reads and responds to human facial expressions

A webcam-based research system connects human facial movements with AI responses, revealing both technical progress and limitations.

Shy Cohen
Edited By: Shy Cohen/
ERC Writer: Merilin Reede
Add as a preferred source in Google
A Tallinn University project uses webcam input to generate AI facial responses, while testing accuracy, context and demographic bias.

A Tallinn University project uses webcam input to generate AI facial responses, while testing accuracy, context and demographic bias. (CREDIT: The Brighter Side of News)

  • A Tallinn University doctoral project connects webcam-based expression analysis with facial responses from a human-like AI agent.
  • Tracking individual facial movements gave the system finer, more stable control than broad emotion labels.
  • Experiments exposed disagreements between analysis tools and demographic bias, limiting claims about reliably interpreting emotions.

A digital face can look convincing until its expression fails to match the conversation. Realistic skin and detailed animation do not guarantee an appropriate reaction to another person’s smile, frown or changing expression.

A doctoral project at Tallinn University addresses that gap by connecting facial-expression analysis with the generation of an artificial character’s response. Rather than assigning a label to one photograph, the system treats facial behavior as an exchange unfolding over time.

Abdallah Hussein Sham developed the Enactive Facial Expression Pipeline, or EFEP. The work combines experimental datasets and previously published research, including studies in Sensors and Signal, Image and Video Processing.

The resulting system uses a standard webcam to analyze a person’s expression, estimate a response and animate the agent’s face. Its findings concern recognition and synthesis of facial behavior. They do not establish that an AI character understands someone’s private feelings.

Abdallah Hussein Sham. (CREDIT: Tallinn University)

Building a response, rather than reading a snapshot

The thesis starts from an interaction problem: an expression acquires meaning within an ongoing encounter. A character needs to produce a response, rather than simply identify what the other face appears to display.

EFEP is modular, linking separate stages for analyzing facial input, estimating a reaction and generating visible facial movement. This arrangement connects the observation of a person with the behavior of an artificial agent.

The research found that individual facial muscle movements provided finer and more stable control than categories such as “happy” or “sad.” Those broad labels compress many visible details into a single description.

A supporting synthesis study used action units from the Facial Action Coding System to control an artificial character. These units describe particular facial movements, providing a way to manipulate its facial points.

The researchers used the OpenFace software interface and Unreal Engine 4 for that work. They first programmed the character to mimic expressions, then trained a model to generate reactions using sequences of earlier expressions.

Sixty participants supplied reactions in different contexts

Training an agent to respond requires examples of reactions, not only pictures labeled with emotions. The research combined existing expression datasets with a new collection involving 60 participants and five imagined social situations.

Screenshot of avatar trying to mimic the user’s facial expression using OpenFace from a recorded video. (CREDIT: Abdallah Hussein Sham et al, Etera 2026)

In the dataset study, published in Sensors in 2023, participants watched prerecorded videos of people displaying changing expressions. A webcam recorded the viewers’ facial behavior while instructions supplied different social contexts.

The people on screen were introduced as a best friend, colleague, stranger, someone the participant hated or someone they feared. Participants were asked to respond as they would in the specified situation.

These were simulated encounters with videos, rather than unrestricted conversations between two people. That distinction defines what the collected reactions can demonstrate and how closely they represent everyday interaction.

The researchers also gathered written feedback about the procedure and its perceived naturalness. Most participants described their reactions as natural, although responses varied. Some considered actual interactions easier or more expressive than responding to a recording.

The collection used ordinary equipment, including a 1080p webcam and a laptop. Its accessible recording method was intended to support research into context-dependent facial reactions for virtual or robotic characters.

Analysis tools disagreed about the same faces

The reaction dataset revealed a significant methodological problem. Three facial-expression analysis systems did not agree on a single emotion prediction for the recorded reactions, despite participants reporting that their responses felt genuine.

Dynamics of emotional facial expression and speech prosody in a sentence. (CREDIT: Abdallah Hussein Sham et al, Etera 2026)

Those systems were iMotions, a Mini-Xception model and the Py-Feat toolkit. Their disagreement led the authors to call for a more robust approach to predicting reactions.

That finding separates visible behavior from an interpretation assigned by software. Recording a face does not provide a definitive emotional label, and different models can classify the same material differently.

A separate 2024 synthesis paper evaluated generated reactions against test data using average root-mean-square error. Sixteen test participants also reported on the apparent naturalness of the character’s responses to changing human expressions.

These are different kinds of evaluation. Comparing generated output with a dataset tests correspondence, while participant reports address how the animation appears. Neither alone establishes successful communication across all people or situations.

The small participant assessment also leaves broader testing important. Evidence from a prototype should not be treated as proof that an agent will respond appropriately throughout a long, unpredictable conversation.

Diverse training data mattered to performance

Fairness was another central strand of the research. Models trained on images representing only limited demographic groups showed accuracy differences when applied to other groups.

Robots with animated facial expressions. (CREDIT: Abdallah Hussein Sham et al, Etera 2026)

The university’s account reports that the observed gap disappeared with diverse training data. A supporting study provides a more specific qualification: adding variation can reduce bias, but does not necessarily remove it completely.

That study examined several expression-recognition methods, including Deep Emotion, Self-Cure Network, ResNet50, InceptionV3 and DenseNet121. It compared performance using images of people from different racial groups.

The models tended to favor groups represented in their training data. The authors found that improving overall performance could also increase bias when the training collection remained imbalanced.

Their findings supported equally including previously missing groups in the training data. This is evidence about the tested models and datasets, rather than a guarantee of fairness in every deployment.

The result matters for expression generation as well as recognition. Within the pipeline, the interpretation of a person’s face informs the response produced by the agent.

A research prototype for limited settings

The project targets interactive media, creative applications and user experience research. These settings offer opportunities to investigate nonverbal interaction without presenting the system as a medical assessment or an employment tool.

Grad-CAM visualisation for DenseNet121. The model attends to the mouth region for happy expressions, consistent with AU12 (lip corner puller). (CREDIT: Abdallah Hussein Sham et al, Etera 2026)

Its limited scope also reflects the sensitivity of emotion-related AI. The EU AI Act prohibits systems used to infer emotions in workplaces and educational institutions, subject to medical or safety exceptions.

That restriction does not establish that every entertainment application is harmless or automatically compliant. It underscores the importance of separating experimental facial interaction from judgments about people in consequential settings.

Voice and body language are proposed additions, rather than demonstrated features of the current pipeline. Incorporating them would expand the research beyond camera-based facial behavior and require further evaluation.

EFEP offers a tested approach to connecting an observed expression with an animated response. Its strongest lesson is also a boundary: a more responsive digital face can be technically useful without becoming a reliable window into human emotion.

Dig deeper into facial expressions and human–AI interaction

These resources examine facial analysis tools, emotional interpretation and the ethical limits of expression-based AI.

Py-Feat: Python Facial Expression Analysis Toolbox: Introduces an open-source toolkit for detecting, analyzing and visualizing facial-expression data. (Affective Science, 2023)

Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements: Reviews the evidence and limitations behind assigning emotional meaning to facial movements. (Psychological Science in the Public Interest, 2019)

OpenFace 2.0: Facial Behavior Analysis Toolkit: Describes software for analyzing facial landmarks, action units, head pose and gaze. (IEEE International Conference on Automatic Face & Gesture Recognition, 2018)

Not in My Face: Challenges and Ethical Considerations in Automatic Face Emotion Recognition Technology: Examines scientific and ethical challenges in automated facial-emotion recognition. (Machine Learning and Knowledge Extraction, 2024)

Article 5: Prohibited AI practices: Presents the AI Act’s restrictions, including specified uses of emotion inference in employment and education. (European Commission, 2026)

Research findings are available online in the journal Etera.

The original story "New webcam-based AI system reads and responds to human facial expressions" is published in The Brighter Side of News.



Like these kind of feel good stories? Get The Brighter Side of News' newsletter.


Shy Cohen
Shy CohenScience and Technology Writer

Shy Cohen
Writer

Shy Cohen is a Washington-based science and technology writer covering advances in artificial intelligence, machine learning, and computer science. Having published articles on MSN, AOL News, and Yahoo News, Shy reports news and writes clear, plain-language explainers that examine how emerging technologies shape society. Drawing on decades of experience, including long tenures at Microsoft and work as an independent consultant, he brings an engineering-informed perspective to his reporting. His work focuses on translating complex research and fast-moving developments into accurate, engaging stories, with a methodical, reader-first approach to research, interviews, and verification.