Low Immunogenicity in Human Panels? Examining Latent-X2’s Evidence
Analyzing the immunogenicity assay quality and binding validation in the first AI antibody study to test human immune responses
Low Immunogenicity in Human Panels? Examining Latent-X2’s Evidence
Latent Labs released preprint data on Latent-X2, their generative model for antibody and macrocycle design. This is the first study to report immunogenicity data for AI-generated antibodies, a critical milestone, since immunogenicity is a leading cause of clinical failure. The paper’s lead claim is “drug-like antibodies with low immunogenicity in human panels.”
But the immunogenicity evidence doesn’t support that claim. They tested only 4 VHH binders from a single target (TNFL9) in Fc-fusion format, not the naked VHH domains they designed. The 10-donor panel is heavily biased: 60% of donors share the same HLA B44 supertype and 40% share B08, meaning you’re testing ~7–8 immune backgrounds, not 10 independent ones. The assay lacks appropriate positive controls (no bococizumab, no KLH), uses indirect detection methods (bulk cytokines and ATP instead of T-cell activation markers), and has no predefined statistical framework to define what counts as an immunogenic response.
Beyond immunogenicity, the binding validation has gaps. The paper reports picomolar affinities for two HDAC8 binders (26–41 pM), but both sensorgrams show baseline drift during dissociation, a technical artifact that prevents accurate KD determination. There’s no epitope mapping for any construct, no competition binding data, no functional validation beyond binding.
In this post, I analyze both the immunogenicity assay quality and the overall reporting of tested binders.

Binding validation for computationally designed VHHs across 7 targets. Blue structures are VHH designs, red structures are scFv designs. Surface plasmon resonance (SPR) and biolayer interferometry (BLI) sensorgrams show concentration-dependent binding with fitted KD values. Black dashed lines show kinetic fits. Reported affinities span picomolar (HDAC8 binders at 26–41 pM) to micromolar (ONCM binder at 3.26 μM), with most VHH designs in the 2.75–797 nM range. Several sensorgrams show technical issues: TNFL9_10 displays biphasic binding; ONCM_3 shows potential biphasic behavior; multiple targets (MMP2, AHSP, LEP) exhibit incomplete dissociation; both HDAC8 constructs show baseline drift during dissociation; 1433B_1 and 1433B_6 report identical KD values (4.27 nM) despite different kinetic profiles.
How The Model Works
Latent-X2 is an all-atom generative model that jointly designs sequence and structure. The paper doesn’t disclose architectural details.
The model takes 3 inputs:
- Target protein structure (only the backbone atoms)
- Epitope specification
- Optional antibody framework with desired CDR lengths
And it outputs three things:
- Complete binder sequence
- All-atom 3D structure of the binder
- The bound complex with the target
Multi-modality: The model can generate different types of binders from the same architecture:
- VHH (nanobodies): Single-domain antibodies
- scFv: Two-domain antibodies with VH and VL chains connected by a linker
- Macrocycles: Cyclic peptides, 6–18 residues
- Mini-binders
You specify which modality you want during generation. For antibodies, you provide the framework and desired CDR lengths. For macrocycles, you specify the peptide length. Maximum input length is 512 residues, target + binder together.
Results
Against 20 targets, they tested 4–24 designs per modality (i.e., VHH, scFv, macrocycle) per target.
- ~55% target-level reported success rate (11/20 targets; 9 antibody targets + 2 macrocycle targets).
- No success rate reported per modality per target (i.e., we don’t know how many designs were generated & tested).
- 47% of binders met all four developability criteria without optimization (monomericity, hydrophobicity, thermostability, polyreactivity)
Methods used:
- SPR or BLI for binding kinetics and affinity
- Baculavirus particle (BVP) ELISA for polyreactivity
- NanoDSF for thermal stability (Tm)
- HIC-HPLC for hydrophobicity (used lirilumab and tremelimumab as positive and negative controls, respectively per Jain et al. 2017)
- SEC-HPLC for monomericity
- Cell-based assay with human PBMCs for ex vivo immunogenicity
Binders:
- HDAC8 (26.2 pM, 41.5 pM, both scFvs)
- 1433B (2.75 nM VHH; 4.27 nM VHH; 4.27 nM scFv)
- LEP (44.8 nM VHH)**
- BHRF1 (27.0 nM VHH)**
- TNFL9 (31.1 nM, 45 nM, 49.7 nM, all VHHs. 5/24 reported success rate.)
- MMP2 (77 nM, 555 nM, both VHHs)
- SAE1 (165 nM scFv)
- AHSP (797 nM)
- ONCM (3.26 µM)
- K-Ras(G12D) (5.53 µM, 23.90 µM, both macrocycles. 8/10 reported success rate.)
- PHD2 (1.54 nM, 4.70 nM, both macrocycles. 9/10 reported success rate.)
Targets failed: 1433E, CD98, IL-20, IL-3, PD-L1, PRL, SC2RBD, SOMA, UBE2B
* There might be an error in Figure 2 for the KD values calculated for LX2_VHH_1433B_1 and LX2_scFv_1433B_6. They both have the same KD value with different kinetics profiles.
** There is a discrepancy between the KD values reported in the paper and the values on their website (link and link). Additionally, there are 2 more designs listed per target on the website that don’t appear in the paper.
*** These two designs are not listed on the website. Only the first binder (31.1 nM) is listed.
They report antibody binders for 9 targets but tested immunogenicity for only 4 binders from 1 target (TNFL9). They confirmed 5 TNFL9 binders total but tested 4 for immunogenicity. Why exclude 1? What about VHHs to MMP2, LEP, 1433B, BHRF1, AHSP, ONCM? What about the scFvs including their best binder?
For TNFL9 immunogenicity: Four VHH binders showed no detectable immunogenic response across 10 healthy human donors in T-cell proliferation and cytokine release assays. But the paper reports binding data for only 3 of these. We don’t know how and why they picked those specific 4 VHHs for immunogenicity testing.
For macrocycle success rates: K-Ras(G12D) and PHD2 showed 80–90% hit rates from 10 designs per target, matching or exceeding trillion-scale mRNA display screens. But they only report KD values for 2 binders per target. That’s 13 confirmed binders with missing affinity data.
Reporting, Transparency, and Validation Gaps
1. Picomolar affinity claim based on unoptimized BLI experiment
Authors report picomolar affinities for two HDAC8-binding scFvs (26.2 pM and 41.5 pM) based on BLI measurements. However, both sensorgrams show baseline drift during the dissociation phase, which prevents accurate determination of the dissociation rate constant (koff). Since KD = koff/kon, compromised dissociation kinetics directly undermine the affinity calculation.
Baseline drift during dissociation typically indicates:
- Non-specific aggregation on the sensor surface
- Avidity effects (especially problematic for scFv constructs, which can form dimers)
- Rebinding due to insufficient washing or mass transport limitations
- Assay conditions not optimized for the interaction
Picomolar binding claims require rigorous validation. When SPR/BLI data show artifacts, standard practice is to:
- Optimize assay conditions (buffer composition, flow rate, surface density, regeneration conditions)
- Use orthogonal methods (e.g., KinExA)
- Demonstrate functional consequences (sub-picomolar binders should show potent inhibition at low nM concentrations)
The preprint provides none of these. The HDAC8 constructs likely bind tightly (the sensorgrams show clear concentration-dependent responses) but the actual KD values cannot be determined from these data. Without optimized dissociation kinetics or orthogonal validation, the picomolar claims are not supported by the data shown. These affinities should be reported as “low nanomolar or better” until proper validation confirms sub-nanomolar binding.
2. No success rates reported per modality per target
The paper states they tested “4–24 designs per modality per target” but only reports exact counts for 3 of 11 successful targets: TNFL9 (5/24 VHH = 21%), PHD2 (9/10 macrocycles = 90%), and K-Ras(G12D) (8/10 macrocycles = 80%). For the other 8 targets, we don’t know how many designs were tested per modality. I expected to see “X designs tested, Y binders confirmed” per target per modality.
3. Missing affinity data for confirmed binders
PHD2 had 9 confirmed binders but only 2 KD values are reported. K-Ras(G12D) had 8 confirmed binders but only 2 KD values are reported. That’s 13 missing KD values. They ran SPR or BLI to confirm binding, so the data exists. Why report only the best 2 per target? Are the other binders weak? Did they fail developability?
4. Paper-website discrepancies
Multiple targets show inconsistencies between the preprint and their website data:
LEP: Different KD values reported in paper versus website. The website also lists 2 additional designs that don’t appear in the paper.
BHRF1: Different KD values reported in paper versus website. The website also lists 2 additional designs that don’t appear in the paper.
TNFL9: This one is especially problematic. The paper reports 5 confirmed binders (31.1 nM, 45.0 nM, 49.7 nM, plus two others). They tested 4 of these for immunogenicity. But the website only lists 1 binder: the 31.1 nM design.
Why do some binders appear in the paper but not the website? Why do some appear on the website but not in the paper?
5. Unexplained selective reporting for immunogenicity testing
TNFL9 had 5 confirmed VHH binders. They tested 4 for immunogenicity and excluded 1 without explanation. Were the 4 selected randomly? Best binders? Most developable? Lowest predicted immunogenicity? If you’re claiming “no immunogenicity detected,” excluding 20% of your binders without explanation undermines that claim.
6. No failure mode analysis
9 targets failed (1433E, CD98, IL-20, IL-3, PD-L1, PRL, SC2RBD, SOMA, UBE2B). Zero information on which modalities were tested, how many designs per modality, or why they failed. You cannot assess model limitations without understanding failures.
7. No Epitope Validation
They claim the model generates binders that “engage the specified epitope” and that “binding specificity was assessed by alanine mutagenesis of key CDRH3 residues.” They mutated 3 positions in CDRH3, saw loss of binding in all 4 VHHs, and concluded this confirms “target recognition being driven by the designed epitope interactions.”
This doesn’t validate epitope specificity. Mutating CDRH3 and losing binding just shows those residues are important for binding. It doesn’t tell you where on the target you’re binding. You need actual epitope mapping (hydrogen-deuterium exchange, cross-blocking experiments, or co-crystal structures) to figure out if you are hitting the intended epitope or binding somewhere else that still requires those CDRH3 residues.
8. Undisclosed expression system selection criteria
The paper uses two different expression systems (mammalian, CHO cells, and cell-free) but provides no rationale for why specific targets were assigned to each system. From the methods (Section D.1):
- CHO expression (14 targets): 1433B, 1433E, AHSP, CD98, HDAC8, IL-20, IL-3, LEP, MMP2, ONCM, PRL, SOMA, TNFL9, UBE2B
- Cell-free expression (4 targets): BHRF1, PD-L1, SAE1, SC2RBD
Expression system affects protein folding, post-translational modifications, yield, and ultimately whether you can even test a design. Cell-free expression is sometimes used as a rescue strategy when designs fail in mammalian cells due to toxicity, poor folding, or low expression. If that happened here, it means some designs failed very early with critical developability/manufacturability criteria. A single sentence explaining the decision criteria would resolve this issue.
Immunogenicity Assay Quality Gaps
9. Single-target, single-format assessment
Only TNFL9 was tested for immunogenicity. That’s 1 of 9 successful antibody targets. Only the VHH format was tested, not scFv (which includes their best binder at 26.2 pM). No explanation for why TNFL9 was selected. We cannot assess whether this is typical performance or a cherry-picked low-immunogenicity example.
10. Wrong format tested
All VHHs were expressed as Fc-fusion constructs in pcDNA3.4 vectors with a human IgG1 Fc tag (hinge-CH2-CH3). This means they tested Fc-VHH chimeras, not naked VHH domains. The Fc region introduces additional biology through FcRn binding, Fcγ receptor–mediated APC uptake, immune complex formation, and innate effector signaling. These processes alter antigen exposure, processing, and presentation compared to a naked VHH. As a result, an Fc-fused VHH represents a different molecular format with a distinct immunogenic risk profile. Testing Fc-fusion constructs does not directly support claims about intrinsic VHH immunogenicity, only about that specific Fc-linked format.
Control format ambiquity: The paper doesn’t specify whether they reformatted caplacizumab to match their test constructs. Caplacizumab is approved as a bivalent VHH (two VHH domains linked together), not an Fc-fusion. For a valid comparison, they should have either: (1) cloned and expressed caplacizumab in-house with the same Fc tag used for their designs, or (2) tested their VHHs in bivalent format to match caplacizumab’s approved structure. We don’t know which (if either) they did.
11. HLA supertype imbalance
The 10-donor panel has some imbalances. HLA alleles cluster into supertypes based on shared peptide-binding specificities. Donors within the same supertype present similar epitope repertoires and can show correlated immunogenic responses.
Analyzing the donor panel by HLA supertype shows both overrepresentation of certain supertypes and underrepresentation of others:
- B44 supertype dominates (60%): B18:01 (donors 1, 3), B40:01 (donors 2, 10), B41:02 (donor 8), B50:01 (donor 7). B44 alleles prefer acidic residues at P2 (Glu, Asp) and aromatic/hydrophobic C-termini. Having 60% of donors share this specificity creates redundancy in peptide presentation. While B44 coverage is important (representing ~20–30% of most populations), this level of overrepresentation reduces the effective diversity of the panel.
- B08 supertype overrepresented (40%) Four of 10 donors carry B*08:01 (donors 1, 2, 5, 6). B08 supertype alleles have distinct peptide-binding preferences that create a separate peptide-presentation context. This oversampling reduces effective diversity.
- B62 supertype severely underrepresented (10%) Only 1 donor carries a B62 allele (B*15:02 in donor 4). B62 alleles prefer aliphatic residues at P2 and cover a significant portion of many populations. Single-donor representation provides insufficient coverage for this supertype.
- **Other B supertypes with reasonable coverage:
- *B07 supertype: 30% (3/10 donors: B07:02, B42:01, B51:01)
- B27 supertype: 20% (2/10 donors: B27:05, B39:06)
- B58 supertype: 30% (3/10 donors: B57:01, B58:01, B*58:02)
- HLA-A coverage is better but still clustered:
- A02 supertype: 40% (4/10 donors)
- A03 supertype: 30% (3/10 donors)
- A01 supertype: 20% (2/10 donors)
- A24 supertype: 20% (2/10 donors)
The effective ’N’ of this study is not 10, because of the imbalance with overrepresentation of B44 (60%) and B08 (40%), and severe underrepresentation of B62 (10%). Functional diversity of the peptide-presentation space is closer to an N of 7 or 8. When 60% of donors share the B44 supertype and 40% share B08, negative results cannot reliably exclude immunogenicity across diverse populations. Oversampling of B44 and B08 supertypes combined with B62 underrepresentation increases Type II error risk and limits the panel’s ability to detect immunogenicity in underrepresented immune contexts.
Note: I mapped the HLA types using Sidney et al. 2008 paper, based on their supertype classifications. Full mappings can be found in the Appendix, along with my recommended HLA type coverages.
12. Missing high-immunogenicity positive controls
They included PHA-L, anti-CD3, and anti-CD28 as positive controls. These are polyclonal activators. They don’t test the HLA-II/TCR antigen-specific pathway that causes therapeutic antibody immunogenicity. You need protein antigens known to be immunogenic, like bococizumab (48% ADA in patients), HuA33 (73% ADA), or KLH (keyhole limpet hemocyanin, the universal positive control). Without these, you cannot distinguish between “low immunogenic response” and “assay cannot detect responses.”
Here’s the irony. They do use bococizumab in this paper as the thermostability benchmark. Bococizumab (Tm = 61°C) defines their “acceptable” threshold in the developability criteria section. But bococizumab failed Phase 3 trials due to 48% ADA rates. They cite this failure in reference 17. So they use a clinical failure with massive immunogenicity as a developability standard but don’t use it as an immunogenicity positive control.
Why not also include KLH as well? It’s highly immunogenic in all donors, tests the full HLA-II/TCR pathway, and is used in vaccine adjuvant studies. If your assay can’t detect KLH immunogenicity, your assay doesn’t work.
13. Low assay sensitivity
PBMC-based T-cell activation assays have a fundamental sensitivity problem. Antigen-specific CD4+ T cells are rare. Detection depends on antigen dose and sampling enough donors to cover HLA class II diversity. This study used 30 µg/mL protein (maximum) and 10 donors. Those are conservative choices for this assay type. With these parameters, you cannot tell the difference between “no immunogenic response” and “we didn’t sample enough rare T-cell clones or HLA combinations to find a response.” Some groups run these assays with higher protein concentrations and 40+ donors specifically to address this limitation (e.g., Cohen et al. 2021).
14. Indirect detection methods
They measured T-cell activation indirectly: ATP levels (via CellTiter-Glo) for proliferation and bulk cytokine secretion for activation. PBMCs contain multiple cell types (T-cells, B-cells, monocytes, NK cells) all contributing to these readouts. ATP measures metabolic activity from all cells, not just dividing T-cells. Cytokines like TNF-α and IFN-γ come from monocytes and NK cells, not just T-cells. IL-8 is constitutively produced. These indirect measures mix signals from different cell types and pathways, making it harder to detect antigen-specific T-cell responses. Direct methods use activation markers like CD134 and CD137 to count activated CD4+ T-cells specifically.
15. No statistical framework
They don’t define what counts as an immunogenic response. No stimulation index calculation, no predefined threshold, no statistical significance testing. Their conclusion of “no detectable immunogenic response” comes from looking at bar graphs in Figure 3. Without objective criteria, you cannot distinguish between “no response,” “response below detection threshold,” and “assay noise.” The negative result isn’t falsifiable. You need a predefined decision framework to make quantitative claims about immunogenicity risk.
16. No mechanism confirmation
No controls to confirm the assay measures HLA-II restricted CD4+ T-cell activation. No HLA blocking experiments. They tested Fc-fusion proteins in PBMC culture. You need to rule out Fc-mediated effects and innate responses. Without these controls, you cannot tell if you’re measuring antigen-specific T-cell responses or something else.
17. Technical execution problems
They cultured cells for 120 hours. That’s long enough for cell death, overgrowth, and nonspecific activation. They don’t report viability measurements. They froze and thawed supernatants before cytokine analysis. IL-10 and IL-6 are unstable under freeze-thaw. They didn’t deplete CD8+ T-cells, which is needed for CD4+ T-cell specific responses. IL-8 was “above range” for multiple conditions, meaning the assay wasn’t optimized for the concentration range. They used only 2 technical replicates per condition with no biological replicates.
Conclusion
Latent-X2 is the first AI antibody design study to publish immunogenicity data. Immunogenicity assays are not cheap to run and they can get finicky. I am very happy to see it is finally included in an AI antibody design paper. Considering how critical immunogenicity is downstream, I hope to see more teams willing to go the extra mile and include similar assays in their developability panels.
Here’s what would make this work stronger: Test immunogenicity across multiple targets and formats, not just 4 VHHs from TNFL9. Include bococizumab and KLH as positive controls. Use direct T-cell activation markers and predefined statistical thresholds. Report design counts per modality, all confirmed binder affinities, and failure modes. Make the paper and website datasets consistent.
The binding work demonstrates functional binders across multiple targets with reasonable developability profiles. The macrocycles match trillion-scale mRNA display screens in hit rate. But the picomolar affinity claims for HDAC8 aren’t supported because both sensorgrams show baseline drift during dissociation that prevents accurate KD determination. Without optimized assays or orthogonal validation, those should be reported as “low nanomolar or better,” not 26–41 pM.
Too many acronyms in one post? **Check out my biotech abbreviation cheat sheet** and feel free to suggest additions.
Follow me on **Substack, Medium, Bluesky, and LinkedIn **for more posts on drug discovery, assay development, and screening workflows.
Appendix
HLA Type Mappings
Full map can be found here: link.
HLA-A Supertypes:
- A01: 5/10 donors (50%): donors 1, 2, 3, 5, 10
- A01/A03: 1/10 donors (10%): donor 3 (A*30:01)
- A02: 4/10 donors (40%): donors 2, 5, 6, 7
- A03: 4/10 donors (40%): donors 4, 8, 9, 10
- A24: 2/10 donors (20%): donors 4, 7
- A01/A24: 0/10 donors (0%): ABSENT
HLA-B Supertypes:
- B07: 3/10 donors (30%): donors 3, 4, 9
- B08: 4/10 donors (40%): donors 1, 2, 5, 6
- B27: 2/10 donors (20%): donors 6, 7
- B44: 6/10 donors (60%): donors 1, 2, 3, 7, 8, 10
- B58: 3/10 donors (30%): donors 8, 9, 10
- B62: 1/10 donor (10%): donor 4
- Unclassified: 1/10 donor (10%): donor 5 (B*13:01)
Acceptable Supertype Coverage
For a 10-donor immunogenicity panel to provide adequate population coverage, supertype representation should approximate global HLA frequencies while ensuring all major supertypes have sufficient representation to detect potential epitopes.
Minimum Acceptable Coverage (10-donor panel):
HLA-A Supertypes:
- A01: 2 donors (20%)
- A02: 4–5 donors (40–50%)
- A03: 3 donors (30%)
- A24: 2 donors (20%)
Current panel A coverage is acceptable, though A02 is slightly underrepresented at 40% (expected ~50%).
HLA-B Supertypes:
- B07: 3–4 donors (30–40%) → Current: 3 (30%) ✓
- B08: 2 donors (20%) → Current: 4 (40%) – 2x overrepresented
- B27: 1–2 donors (10–20%) → Current: 2 (20%) ✓
- B44: 2–3 donors (20–30%) → Current: 6 (60%) – 2–3x overrepresented
- B58: 1–2 donors (10–20%) → Current: 3 (30%) ✓
- B62: 2 donors (20%) → Current: 1 (10%) – 2x underrepresented
What the panel should look like:
To correct these imbalances, the study should:
- Reduce B44 representation from 6 to 2–3 donors (removing 3–4 donors)
- Reduce B08 representation from 4 to 2 donors (removing 2 donors)
- Increase B62 representation from 1 to 2 donors (adding 1 donor)
This would require replacing 4–5 donors to achieve balanced coverage. Alternatively, expanding to 20–30 donors would allow balanced representation across all supertypes without requiring perfect matching to population frequencies, which is impractical with only 10 donors.
메타데이터
- post_id
- 9c540b8f49ba
- slug
- low-immunogenicity-in-human-panels-examining-latent-x2s-evidence-9c540b8f49ba
- url
- https://medium.com/@enginyapici/low-immunogenicity-in-human-panels-examining-latent-x2s-evidence-9c540b8f49ba
- canonical_url
- https://medium.com/@enginyapici/low-immunogenicity-in-human-panels-examining-latent-x2s-evidence-9c540b8f49ba
- author_url
- https://medium.com/@enginyapici
- status
- ok
- fetched_at
- 2026-06-09 15:37:30