A feature of VBASE2 that I like is its nomenclature, and nomenclature may need to be an issue addressed by AIRR. The IMGT nomenclature has the approval of the Human Genome organization, and IMGT has found fairly simple solutions to the challenges to the underlying logic of the nomenclature that arose from recent studies. The situation with the mouse is more complicated.
Our recent mouse study identified such divergence between the IGHV genes of the BALB/c mouse and the C57BL/6 strain that I think it will be difficult for the IMGT mouse nomenclature to survive. This is because it is no longer possible to be certain of the relationship between C57 and BALB/c genes that are presently considered allelic variants of one another. If there are 50% more IGHV genes in the BALB/c strain than in the C57 strain, their sequences may never be paired, and the apparent logic of the IMGT nomenclature breaks down.
VBASE2 has a nomenclature that does not attempt to describe relationships between sequences. Instead, each sequence is given a unique number, along with identifiers showing the species and gene type. eg musIGHV233. The three classes of sequences is also a useful idea. VBASE explains it this way: “Class 1 sequences are supported by a genomic sequence and a rearrangement. Class 2 contains sequences with genomic evidence only and class 3 holds sequences which have been found in rearrangements only.” A strategy of this kind would allow the incorporation of seqeunces that are inferred from VDJ rearrangements. I consider it to be essential for such sequences to be included for two reasons. Firstly, the evidence for inferences from hundreds and even thousands of VDJ rearrangements can be very convincing. And secondly, such inferences are likely to be most if not all we have to consider, until technologies change again, or research interests evolve. Genomic studies of antibody genes will for the time being probably be few and far between.