How to align V9 region - out of bounds?

Hi all (Pat)

I am aligning 18S V9 (using 1510R) and the alignment step is cutting the last 19 bases. I assume this is because is beyong the SILVA or seed alignment region? I am puzzled since I have analyzed V9 in the past and this did not happen before (keep my 133 bases). Am also puzzled since, in some cases, it does retain the 133 bases. The fasta before the alignment is mostly the 130-133 bp of the Euk V9 region, but after the alignment, it is chopped away A LOT. I tried both the mothur SEED downloaded recently and the silva nr128.

Is this something other people has seen? Using mothur v.1.48.0 Last updated: 5/20/22

Thanks,

Leo

I created a small subset of sequences, and aligned it against an old silva128V9 region alignment I have (which includes also the primer region, so I am sure it will not cut the end). Not sure why, still some sequences are getting trimmed on the 3’ end. Does this trimming depend on the exact sequence they align on the template?

The alignment algorithm picks the closest sequence in the database and then aligns to that sequence. So, if that reference sequence is missing the 3’ end, it gets removed in the alignment. Unfortunately, because the 18S database is pretty limited, I think I’ve had to use a more relaxed criteria to screen 18S sequences which brought in more partial sequences. I’d encourage you to use filter.seqs on the reference to only include sequences that start before the V9 region and end at the after the V9 region. Even though you’ll likely lose a lot of sequences, I think there will still be enough to get good alignments from the database.

Pat

thanks - will try that!! No worried for losing some sequences, I agree - ID will be done on PR2