The thymus trains T cells like an AI model to avoid attacking the body
Immune system generalization may help T cells avoid attacking the body after seeing only a small sample of self-peptides.
CSHL Writer: Samuel Diamond

A CSHL model shows how T cells can learn from limited thymus training and still avoid most self-attacks. (CREDIT: Shutterstock)
- The immune system may learn not to attack the body by using a small but representative sample of the body’s own protein fragments.
- T cells can recognize groups of similar fragments, so they do not need to see every possible target during training.
- The model may help explain why failures in this training process can lead to autoimmune disease.
The thymus gives young T cells a test they cannot possibly study for in full. Each cell sees only a fraction of the body’s own protein fragments, yet it must learn not to attack the rest.
That puzzle sits at the center of a study from Cold Spring Harbor Laboratory (CHSL). The work borrows ideas from machine learning to explain how the immune system trains T cells to avoid “friendly fire.”
The process is called negative selection. In the thymus, developing T cells meet fragments of the body’s own proteins, called self-peptides. If a T cell binds too strongly to one of them, it gets deleted. That helps prevent the immune system from attacking healthy tissue.
But the body contains a huge number of self-peptides. Each T cell encounters only a small share before leaving the thymus.
“This has long been an open question in immunology,” explains CSHL Assistant Professor Hannah Meyer. “Negative selection is a crucial process, but if T cells had to test against every single one of the body’s peptides, it would take forever. So, how do they learn to avoid friendly fire? We think it’s through a process called generalization.”
A machine-learning idea inside immune training
Generalization is familiar in artificial intelligence. A machine-learning model can learn from limited examples, then apply that learning to new cases. A dog-detection model, for example, does not need to see every dog image online.
“Machine learning deals with the generalization problem all the time,” CSHL Associate Professor Saket Navlakha explains.
Navlakha, Meyer and their team asked whether the immune system might solve a similar problem. The thymus acts like a training environment. The rest of the body becomes the test environment.
“Machine learning tells us that generalization is possible, but only under certain conditions. It’s not black magic. So, let’s try to look at this immunology question from that perspective.”
For generalization to work, the training data must resemble the test data. In immune terms, the self-peptides in the thymus must resemble the self-peptides in tissues throughout the body.
The team found that they do. Self-peptide abundance in the thymus closely mirrored abundance in tissues elsewhere. That means common self-peptides in the body tend to be common in the thymus too.
Why one T cell can learn from many targets
The second condition involves similarity. T-cell receptors are cross-reactive, meaning one receptor can recognize several similar peptides. That gives the immune system a shortcut.
A T cell does not need to see every peptide it might later attack. If it recognizes one similar self-peptide during training, the thymus can remove it before it becomes dangerous.
The researchers modeled this process using a computational approach. They focused on T cells that recognize protein fragments presented by MHC class I molecules, which are important for CD8 T cells.
The analysis examined 9-amino-acid peptides that could bind to a well-studied human immune molecule, HLA-A02:01. Out of 11,136,576 reference proteome peptides, 434,276 were predicted to bind to HLA-A02:01.
The researchers then reduced those peptides to shorter 6-amino-acid versions. That made the peptide space easier to study while preserving positions important for T cell binding. The final analysis included 426,316 self-peptides.
The thymus covered almost all of them. Only 0.6%, or 2,494 of 426,316 peripheral self-peptides, had zero abundance in the thymus.
That close match matters because it gives developing T cells a training sample that reflects what they may meet later.
Sparse sampling, strong protection
The team then asked how much each T cell actually needs to see. Their model estimated that a developing T cell interacts with about 240 medullary thymic epithelial cells during a two-day window when deletion can occur.
Based on those interactions, each T cell may encounter about 20,000 to 200,000 unique self-peptides. That is still only a portion of the full set.
Yet the model showed that limited sampling could offer strong protection. When T cells sampled just 5% of unique self-peptides in the thymus, negative selection reduced peripheral self-reactivity by more than 80%. When sampling reached 23%, protection exceeded 95%.
The study also reported that 90% of self-reactive T cells could be correctly deleted despite each one encountering only 10% of the body’s self-peptides.
That result depends heavily on cross-reactivity. Without it, a self-reactive T cell would need to encounter its exact target peptide in the thymus. The chance of that happening was nearly zero in the model.
The researchers also found that human self-peptides are not scattered randomly. They cluster in peptide space, meaning similar self-peptides often sit near each other. Around a given self-peptide target, there were an average of 44 other self-peptides nearby. A random distribution would produce about 12.
This clustering helps the thymus teach more with less. One encounter can provide information about a neighborhood of related peptides.
When generalization breaks down
The team also used the model to study autoimmune disease. Autoimmunity happens when the immune system wrongly attacks healthy tissue.
The model reproduced features of autoimmune polyendocrine syndrome type 1, a rare autoimmune disorder linked to mutations in the AIRE gene. AIRE helps the thymus display many self-peptides during T cell training.
When the researchers modeled the loss of AIRE-dependent genes, the thymus lost part of its training set. The body’s tissues still contained those targets, but the thymus no longer taught T cells about all of them.
That change created weak spots. T cells with high survival probability became linked to self-peptides found in APS1-affected tissues, including the adrenal gland, liver, ovary, pancreas, skin, small intestine, thyroid and testis.
The model also matched known immune targets in APS1. Out of 21 high-confidence target proteins from a proteome-wide survey, 19 were predicted targets of T cells whose survival probability increased in the AIRE-deleted thymus. The listed exceptions were IFNA1 and GIF.
The limits of thymus training
The findings do not mean negative selection catches every dangerous T cell. Some self-reactive T cells escape the thymus and enter the body.
That is why other safeguards still matter. The immune system also uses peripheral tolerance mechanisms, including regulatory T cells and anergy, to control escaped self-reactive cells.
The model suggests negative selection makes that later job easier. Before negative selection, each self-peptide could bind to a median of 1,632 T cell receptors. After negative selection, that number dropped to 107 in one model condition.
That is a 15-fold reduction. It means the thymus may not need perfection to reduce autoimmune risk. It may only need to shrink the problem enough for other systems to manage.
The authors also note limits. Real T cell encounters in the thymus may not be fully random. T cells interact with other antigen-presenting cells, including B cells and dendritic cells. Better single-cell data and improved measurements of peptide presentation could sharpen future models.
The cross-reactivity model also has limits. It used a simplified peptide space, not the full 9-amino-acid space. A fuller analysis would be much larger and harder to explore.
Practical implications of the research
The study gives immunology a new way to frame an old question. It suggests the thymus may train T cells through a biological form of generalization, not by showing every possible self-peptide.
That idea could help researchers identify where immune tolerance is most likely to fail. It may also guide better models of autoimmune disease, especially disorders tied to missing or poorly displayed self-peptides in the thymus.
“We’re calling this direction ImmunoAI,” Navlakha says. “We’re not trying to create AI inspired by the immune system, but we’re studying how the immune system solves fundamental machine learning problems. When we start to look at the immune system as if it’s another kind of AI, we may find out some surprising things about human health and disease.”
Research findings are available online in the journal Science Advances.
The original story "The thymus trains T cells like an AI model to avoid attacking the body" is published in The Brighter Side of News.
Related Stories
- Neandertal ancestry significantly affects our immune system today
- Common asthma drug slows tumor growth and helps the immune system recover
- Stanford researchers cure type 1 diabetes in mice by resetting the body's immune system
Like these kind of feel good stories? Get The Brighter Side of News' newsletter.
Shy Cohen
Writer



