New webcam-based AI system reads and responds to human facial expressions
A webcam-based research system connects human facial movements with AI responses, revealing both technical progress and limitations.
ERC Writer: Merilin Reede
A Tallinn University project uses webcam input to generate AI facial responses, while testing accuracy, context and demographic bias. (CREDIT: The Brighter Side of News)
- A Tallinn University doctoral project connects webcam-based expression analysis with facial responses from a human-like AI agent.
- Tracking individual facial movements gave the system finer, more stable control than broad emotion labels.
- Experiments exposed disagreements between analysis tools and demographic bias, limiting claims about reliably interpreting emotions.
A digital face can look convincing until its expression fails to match the conversation. Realistic skin and detailed animation do not guarantee an appropriate reaction to another person’s smile, frown or changing expression.
A doctoral project at Tallinn University addresses that gap by connecting facial-expression analysis with the generation of an artificial character’s response. Rather than assigning a label to one photograph, the system treats facial behavior as an exchange unfolding over time.
Abdallah Hussein Sham developed the Enactive Facial Expression Pipeline, or EFEP. The work combines experimental datasets and previously published research, including studies in Sensors and Signal, Image and Video Processing.
The resulting system uses a standard webcam to analyze a person’s expression, estimate a response and animate the agent’s face. Its findings concern recognition and synthesis of facial behavior. They do not establish that an AI character understands someone’s private feelings.
Building a response, rather than reading a snapshot
The thesis starts from an interaction problem: an expression acquires meaning within an ongoing encounter. A character needs to produce a response, rather than simply identify what the other face appears to display.
EFEP is modular, linking separate stages for analyzing facial input, estimating a reaction and generating visible facial movement. This arrangement connects the observation of a person with the behavior of an artificial agent.
The research found that individual facial muscle movements provided finer and more stable control than categories such as “happy” or “sad.” Those broad labels compress many visible details into a single description.
A supporting synthesis study used action units from the Facial Action Coding System to control an artificial character. These units describe particular facial movements, providing a way to manipulate its facial points.
The researchers used the OpenFace software interface and Unreal Engine 4 for that work. They first programmed the character to mimic expressions, then trained a model to generate reactions using sequences of earlier expressions.
Sixty participants supplied reactions in different contexts
Training an agent to respond requires examples of reactions, not only pictures labeled with emotions. The research combined existing expression datasets with a new collection involving 60 participants and five imagined social situations.
In the dataset study, published in Sensors in 2023, participants watched prerecorded videos of people displaying changing expressions. A webcam recorded the viewers’ facial behavior while instructions supplied different social contexts.
The people on screen were introduced as a best friend, colleague, stranger, someone the participant hated or someone they feared. Participants were asked to respond as they would in the specified situation.
These were simulated encounters with videos, rather than unrestricted conversations between two people. That distinction defines what the collected reactions can demonstrate and how closely they represent everyday interaction.
The researchers also gathered written feedback about the procedure and its perceived naturalness. Most participants described their reactions as natural, although responses varied. Some considered actual interactions easier or more expressive than responding to a recording.
The collection used ordinary equipment, including a 1080p webcam and a laptop. Its accessible recording method was intended to support research into context-dependent facial reactions for virtual or robotic characters.
Analysis tools disagreed about the same faces
The reaction dataset revealed a significant methodological problem. Three facial-expression analysis systems did not agree on a single emotion prediction for the recorded reactions, despite participants reporting that their responses felt genuine.
Those systems were iMotions, a Mini-Xception model and the Py-Feat toolkit. Their disagreement led the authors to call for a more robust approach to predicting reactions.
That finding separates visible behavior from an interpretation assigned by software. Recording a face does not provide a definitive emotional label, and different models can classify the same material differently.
A separate 2024 synthesis paper evaluated generated reactions against test data using average root-mean-square error. Sixteen test participants also reported on the apparent naturalness of the character’s responses to changing human expressions.
These are different kinds of evaluation. Comparing generated output with a dataset tests correspondence, while participant reports address how the animation appears. Neither alone establishes successful communication across all people or situations.
The small participant assessment also leaves broader testing important. Evidence from a prototype should not be treated as proof that an agent will respond appropriately throughout a long, unpredictable conversation.
Diverse training data mattered to performance
Fairness was another central strand of the research. Models trained on images representing only limited demographic groups showed accuracy differences when applied to other groups.
The university’s account reports that the observed gap disappeared with diverse training data. A supporting study provides a more specific qualification: adding variation can reduce bias, but does not necessarily remove it completely.
That study examined several expression-recognition methods, including Deep Emotion, Self-Cure Network, ResNet50, InceptionV3 and DenseNet121. It compared performance using images of people from different racial groups.
The models tended to favor groups represented in their training data. The authors found that improving overall performance could also increase bias when the training collection remained imbalanced.
Their findings supported equally including previously missing groups in the training data. This is evidence about the tested models and datasets, rather than a guarantee of fairness in every deployment.
The result matters for expression generation as well as recognition. Within the pipeline, the interpretation of a person’s face informs the response produced by the agent.
A research prototype for limited settings
The project targets interactive media, creative applications and user experience research. These settings offer opportunities to investigate nonverbal interaction without presenting the system as a medical assessment or an employment tool.
Its limited scope also reflects the sensitivity of emotion-related AI. The EU AI Act prohibits systems used to infer emotions in workplaces and educational institutions, subject to medical or safety exceptions.
That restriction does not establish that every entertainment application is harmless or automatically compliant. It underscores the importance of separating experimental facial interaction from judgments about people in consequential settings.
Voice and body language are proposed additions, rather than demonstrated features of the current pipeline. Incorporating them would expand the research beyond camera-based facial behavior and require further evaluation.
EFEP offers a tested approach to connecting an observed expression with an animated response. Its strongest lesson is also a boundary: a more responsive digital face can be technically useful without becoming a reliable window into human emotion.
Dig deeper into facial expressions and human–AI interaction
These resources examine facial analysis tools, emotional interpretation and the ethical limits of expression-based AI.
Py-Feat: Python Facial Expression Analysis Toolbox: Introduces an open-source toolkit for detecting, analyzing and visualizing facial-expression data. (Affective Science, 2023)
Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements: Reviews the evidence and limitations behind assigning emotional meaning to facial movements. (Psychological Science in the Public Interest, 2019)
OpenFace 2.0: Facial Behavior Analysis Toolkit: Describes software for analyzing facial landmarks, action units, head pose and gaze. (IEEE International Conference on Automatic Face & Gesture Recognition, 2018)
Not in My Face: Challenges and Ethical Considerations in Automatic Face Emotion Recognition Technology: Examines scientific and ethical challenges in automated facial-emotion recognition. (Machine Learning and Knowledge Extraction, 2024)
Article 5: Prohibited AI practices: Presents the AI Act’s restrictions, including specified uses of emotion inference in employment and education. (European Commission, 2026)
Research findings are available online in the journal Etera.
The original story "New webcam-based AI system reads and responds to human facial expressions" is published in The Brighter Side of News.
Related Stories
- AI tool tracks facial aging, and helps doctors gauge cancer risk
- AI can detect hypertension and diabetes by looking at your face and eyes
- New AI software can tell you where to apply makeup to fool facial recognition
Like these kind of feel good stories? Get The Brighter Side of News' newsletter.
Shy Cohen
Writer



