Where the training data comes from, and what sequence each measurement is attached to
Named publications, real sequences, and the exact training rows they produced. Written to be checked, not believed.
Generated 2026-08-08 18:48 from data/provenance_audit.json, which queries KKB directly. Every count below is read from that file.
Why this report exists
A kinase inhibitor that is potent against a wild-type kinase can be useless against a point mutant of the same kinase, and that difference is often the entire subject of the paper reporting it. So if a mutant measurement is filed against the wild-type sequence, the model is shown two contradictory potencies for one protein and the correct thing for it to learn is to ignore the protein.
This is not hypothetical. It happened here.An earlier version of this pipeline attached 9,490 mutant measurements to wild-type sequences. That is the failure this report exists to rule out, and it is why the evidence below is sequences and identifiers rather than a summary statistic.
The publications below were selected by a search over the mutation-rich kinases and are then pinned, because re-running that search across the full panel is prohibitively slow. Each one is verified end to end here, so the selection affects which examples you see and not whether they hold.
The worked examples
Example 1. publication reporting BOTH wild type and mutantTilting the Scales toward EGFR Mutant Selectivity: Expanding the Scope of Bivalent Type V Kinase Inhibitors.
J Med Chem 2024 · target EGFR · wild-type sequence 1,210 residues, md5 99d03b567dbc
L858R: the sequence actually attached to these 36 measurements
residues 834 to 882, substitution L858R
wild typeVHRDLAARNVLVKTPQHVKITDFGLAKLLGAEEKEYHAEGGKVPIKWMA
mutantVHRDLAARNVLVKTPQHVKITDFGRAKLLGAEEKEYHAEGGKVPIKWMA
L858R;T790M: the sequence actually attached to these 35 measurements
residues 766 to 814, substitution T790M
wild typeMASVDNPHVCRLLGICLTSTVQLITQLMPFGCLLDYVREHKDNIGSQYL
mutantMASVDNPHVCRLLGICLTSTVQLIMQLMPFGCLLDYVREHKDNIGSQYL
residues 834 to 882, substitution L858R
wild typeVHRDLAARNVLVKTPQHVKITDFGLAKLLGAEEKEYHAEGGKVPIKWMA
mutantVHRDLAARNVLVKTPQHVKITDFGRAKLLGAEEKEYHAEGGKVPIKWMA
L858R;T790M;C797S: the sequence actually attached to these 23 measurements
residues 766 to 814, substitution T790M
wild typeMASVDNPHVCRLLGICLTSTVQLITQLMPFGCLLDYVREHKDNIGSQYL
mutantMASVDNPHVCRLLGICLTSTVQLIMQLMPFGCLLDYVREHKDNIGSQYL
residues 773 to 821, substitution C797S
wild typeHVCRLLGICLTSTVQLITQLMPFGCLLDYVREHKDNIGSQYLLNWCVQI
mutantHVCRLLGICLTSTVQLITQLMPFGSLLDYVREHKDNIGSQYLLNWCVQI
residues 834 to 882, substitution L858R
wild typeVHRDLAARNVLVKTPQHVKITDFGLAKLLGAEEKEYHAEGGKVPIKWMA
mutantVHRDLAARNVLVKTPQHVKITDFGRAKLLGAEEKEYHAEGGKVPIKWMA
What this publication contributed
| mutation as recorded in KKB | resolved substitution | rows in this paper | compounds | training rows for this form | sequence md5 |
|---|
| L858R | L858R | 36 | 32 | 1,113 | f02e295764e7 |
| L858R;T790M | T790M, L858R | 35 | 31 | 541 | 19ebead2a842 |
| L858R;T790M;C797S | T790M, C797S, L858R | 23 | 23 | 581 | 78442f2fae03 |
| WT | wild type | 34 | 30 | 13,121 | 99d03b567dbc |
Label corrected by the tracePinned as mutant_only by a protocol-level search; the row-level trace finds 3 mutant and 1 wild-type groups, so it is reported as both.
What this provesThis publication contributed 4 distinct sequences for one gene, and 128 of its rows reach training. The wild-type rows carry the wild-type sequence and each mutant carries its own, and deduplication did not collapse them into one another. If the historical bug were present, this number would be 1 and every row would show the wild-type md5.
Example 2. publication reporting ONLY mutant dataIdentification of M4205A Highly Selective Inhibitor of KIT Mutations for Treatment of Unresectable Metastatic or Recurrent Gastrointestinal Stromal Tumors.
J Med Chem 2023 · target KIT · wild-type sequence 976 residues, md5 f753f2b2d975
V654A: the sequence actually attached to these 26 measurements
residues 630 to 678, substitution V654A
wild typeHLTEREALMSELKVLSYLGNHMNIVNLLGACTIGGPTLVITEYCCYGDL
mutantHLTEREALMSELKVLSYLGNHMNIANLLGACTIGGPTLVITEYCCYGDL
What this publication contributed
| mutation as recorded in KKB | resolved substitution | rows in this paper | compounds | training rows for this form | sequence md5 |
|---|
| V654A | V654A | 26 | 26 | 304 | c7df54462e46 |
| exon 11/13;(544-976,V559D;V654A) | unparseable | 6 | 6 | 0 | |
| exon 11/14;(544-976,V559D;T670I) | unparseable | 6 | 6 | 0 | |
| exon 11/17;(544-976,V560G;D816V) | unparseable | 6 | 6 | 0 | |
| exon 11/17;(544-976,V560G;N822K) | unparseable | 6 | 6 | 0 | |
| exon 11;(544-976);wildtype | unparseable | 6 | 6 | 0 | |
| exon 11;(544-976,557-558) | unparseable | 6 | 6 | 0 | |
| exon 11;(544-976,V559A) | unparseable | 6 | 6 | 0 | |
| exon 11;(544-976,V559D) | unparseable | 6 | 6 | 0 | |
| exon 11;(544-976,V560G) | unparseable | 6 | 6 | 0 | |
| exon 13;(544-976;K642E) | unparseable | 6 | 6 | 0 | |
| exon 13;(544-976;V654A) | unparseable | 6 | 6 | 0 | |
| exon 14;(544-976;T670I) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;A829P) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;D816E) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;D816F) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;D816H) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;D816I) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;D816V) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;D816Y) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;D820E) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;D820Y) | unparseable | 6 | 6 | 0 | |
| exon 17;(544-976;Y823D) | unparseable | 6 | 6 | 0 | |
132 of this publication's rows were dropped, not mislabelledThey name variants that change the length of the protein, such as exon-19 deletions or internal tandem duplications, which cannot be expressed as a substitution on a fixed-length sequence. They are discarded rather than filed against the wild type. That is the safe failure, but it is still a loss and it is counted here rather than left silent.
What this provesEvery row this publication contributed carries a sequence whose md5 differs from the wild type, at exactly the substituted positions highlighted above. If the historical bug were still present, all of these rows would show the wild-type md5 and the model would be trained on two contradictory examples of the same protein.
Example 3. publication whose variant data is unrepresentable, so only its wild-type rows surviveDiscovery of potent and selective HER2 inhibitors with efficacy against HER2 exon 20 insertion-driven tumors, which preserve wild-type EGFR signaling.
Nat Cancer 2022 · target EGFR · wild-type sequence 1,210 residues, md5 99d03b567dbc
What this publication contributed
| mutation as recorded in KKB | resolved substitution | rows in this paper | compounds | training rows for this form | sequence md5 |
|---|
| (empty, wild type) | wild type | 2 | 2 | 13,121 | 99d03b567dbc |
| WT | wild type | 48 | 47 | 13,121 | 99d03b567dbc |
| del19 | unparseable | 48 | 47 | 0 | |
| del19,T790M | unparseable | 48 | 47 | 0 | |
Label corrected by the tracePinned as both by a protocol-level search; the row-level trace finds 0 mutant and 2 wild-type groups, so it is reported as wild_type_only.
96 of this publication's rows were dropped, not mislabelledThey name variants that change the length of the protein, such as exon-19 deletions or internal tandem duplications, which cannot be expressed as a substitution on a fixed-length sequence. They are discarded rather than filed against the wild type. That is the safe failure, but it is still a loss and it is counted here rather than left silent.
What this showsEvery variant this publication reports changes the length of the protein, so none of them can be represented. Its wild-type rows are trained on and its variant rows are dropped. Nothing is mislabelled, and nothing about the variants is learned.
The sharpest form: one compound, both proteins
Compound Cc1cc(Nc2ncnc3cnc(nc23)N4CCN(CC4)C(=O)C=C)ccc1Oc5ccc6c(c5)ncn6C was measured against both forms. Its training rows:
| form | pIC50 | class | sequence md5 |
|---|
| WT | 5.54 | Low | 99d03b567dbc |
The same compound appears against both forms, on separate rows, with different sequences. Deduplication did not collapse them.1 distinct sequences across these rows.
Example 4. publication whose variant data is unrepresentable, so only its wild-type rows surviveCellular Context Influences Kinase Inhibitor Selectivity.
J Med Chem 2026 · target ABL1 · wild-type sequence 1,130 residues, md5 d24f1ea01ac4
What this publication contributed
| mutation as recorded in KKB | resolved substitution | rows in this paper | compounds | training rows for this form | sequence md5 |
|---|
| (empty, wild type) | wild type | 3 | 3 | 6,055 | d24f1ea01ac4 |
| E255K-phosphorylated | unparseable | 1 | 1 | 0 | |
| F317I-nonphosphorylated | unparseable | 1 | 1 | 0 | |
| F317I-phosphorylated | unparseable | 1 | 1 | 0 | |
| F317L-nonphosphorylated | unparseable | 1 | 1 | 0 | |
| F317L-phosphorylated | unparseable | 1 | 1 | 0 | |
| H396P-nonphosphorylated | unparseable | 1 | 1 | 0 | |
| H396P-phosphorylated | unparseable | 1 | 1 | 0 | |
| M351T-phosphorylated | unparseable | 1 | 1 | 0 | |
| Q252H-nonphosphorylated | unparseable | 1 | 1 | 0 | |
| Q252H-phosphorylated | unparseable | 1 | 1 | 0 | |
| T315I-nonphosphorylated | unparseable | 1 | 1 | 0 | |
| T315I-phosphorylated | unparseable | 1 | 1 | 0 | |
| Y253F-phosphorylated | unparseable | 1 | 1 | 0 | |
| nonphosphorylated | unparseable | 1 | 1 | 0 | |
| phosphorylated | unparseable | 1 | 1 | 0 | |
Label corrected by the tracePinned as both by a protocol-level search; the row-level trace finds 0 mutant and 1 wild-type groups, so it is reported as wild_type_only.
15 of this publication's rows were dropped, not mislabelledThey name variants that change the length of the protein, such as exon-19 deletions or internal tandem duplications, which cannot be expressed as a substitution on a fixed-length sequence. They are discarded rather than filed against the wild type. That is the safe failure, but it is still a loss and it is counted here rather than left silent.
What this showsEvery variant this publication reports changes the length of the protein, so none of them can be represented. Its wild-type rows are trained on and its variant rows are dropped. Nothing is mislabelled, and nothing about the variants is learned.
Every mutant row that does not reach training, and why
The examples above show individual publications. This is the whole corpus. A record naming a mutation we cannot resolve is dropped rather than filed against the wild-type sequence, so the failure mode is lost data and never wrong data. That is the right trade, but it is still a loss, so here is its size and its composition.
| category | rows | why it cannot be represented |
| insertion or deletion | 2,383 | changes the length of the protein, so it cannot be applied to a fixed-length sequence as a substitution |
| other unparsed | 1,054 | does not match any substitution grammar we recognise |
| construct or domain description | 640 | names which part of the protein was expressed, not a variant |
| phosphorylation state | 495 | describes the enzyme preparation, not a sequence change |
| internal tandem duplication | 264 | adds residues, so the sequence length changes |
| total dropped | 4,836 | |
|---|
For scale, the training table holds 13,189 mutant rows across 237 distinct mutations, so 4,836 rows are lost against 13,189 retained.
The honest limitationMost of the loss is not a parser weakness that better code would fix. An EGFR exon-19 deletion removes residues; an internal tandem duplication adds them. This model represents a protein as a fixed-length amino-acid string with substitutions applied in place, so a variant that changes the length has nowhere to go. Those variants are clinically important, and the model has never seen them. Supporting them is an architecture change, not a bug fix.
One genuine bug, found by writing this report54 rows record the target as "Wild Type" rather than "WT". The check matched only the literal "WT", so these went down the mutation path, failed to parse, and were discarded under a reason that was false. They are wild type. Fixed; the count is small but the drop reason was actively misleading.
Single-shot inhibition data, traced to the optimiser
Percent-inhibition measurements at a stated concentration are not potencies, and an earlier version of this project ignored them. They are converted to a potency bound and enter training as censored records. The reason to trace them rather than assert them is that this project once lost 35,495 rows between the training table and the optimiser, 98.7 percent of the greater-than records, and it went unnoticed for two cycles.
Conversion ruleIC50 = C(100-I)/I with margins at 20 and 80 percent inhibition. Below 20 percent gives an upper bound on potency (class Low); above 80 percent a lower bound (class High); between them the record is uninformative and is dropped.
| stage | rows |
|---|
| KKB percent-inhibition records (all targets) | 449,460 |
| in the training table after conversion | 107,787 |
| surviving embedding coverage, i.e. reaching the optimiser | 107,784 |
| lost to embedding coverage | 3 |
100.0 percent of the converted single-shot rows reach the optimiser. By class: {'Low': 106685, 'High': 1099}.