# Too much OTUs

**URL:** <https://forum.mothur.org/t/too-much-otus/2564>\
**Category:** Commands in mothur\
**Created:** [October 7, 2015, 8:51pm UTC](https://forum.mothur.org/t/too-much-otus/2564 "2015-10-07T20:51:27Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![sebasdiazz](https://avatars.discourse-cdn.com/v4/letter/s/c5a1d2/32.png) [@sebasdiazz](https://forum.mothur.org/u/sebasdiazz)\
**Post date:** [October 7, 2015, 8:51pm UTC](https://forum.mothur.org/t/too-much-otus/2564/1 "2015-10-07T20:51:27Z")

</div>

Hi, Mothur community!

I am working with Miseq data from gut communities from 6 insect species. From my original dataset of 4’105.742 after trim short, nucleotide ambiguity, unaligned, non-bacterial and chimeric sequences, my total number of sequences are 2’062.377 and after unique.seqs and pre.cluster of 2 bp, 640.944 unique sequences. For every species I had something like 300.000 sequences and 90.000 unique sequences, and in the clustering at 97%, I found more than 200 OTUs per sample some even with 1000!!! (and you know for insect gut I must be something like 20-30 OTUs), with most of the OTUs as singletons and underrepresented OTUs (less than 10 seqs), so I guess I have a lot of spurious reads.

I want to back in my pre-processing steps, to try to identified the reads with sequencing errors. I changed in the pre.cluster the threshold from 2 to 4 (that will represent a error in the 1-2% of the sequence length of 427 bp), and now I have 332.577 unique sequences. Also, using split.abund with cutoff=1, I found 314.161 singletons (7.6% of my original dataset), what I still think that is a high number.

What do you recommend? Try with a higher pre.cluster or maybe eliminate the singletons or other alternative?

Thank you,  
Sebastián

---

<div class="post-metadata">

**Author:** ![pschloss](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.mothur.org/pschloss/32/4_2.png) [@pschloss](https://forum.mothur.org/u/pschloss)\
**Post date:** [October 12, 2015, 3:09pm UTC](https://forum.mothur.org/t/too-much-otus/2564/2 "2015-10-12T15:09:24Z")

</div>

I suspect you aren’t sequencing the V4 region with paired 250 nt reads, right? See this…

[http://blog.mothur.org/2014/09/11/Why-such-a-large-distance-matrix%3F/](http://blog.mothur.org/2014/09/11/Why-such-a-large-distance-matrix%3F/)
