# screen.seqs: removal of high percent of sequences

**URL:** <https://forum.mothur.org/t/screen-seqs-removal-of-high-percent-of-sequences/2804>\
**Category:** mothur bugs\
**Created:** [April 11, 2016, 8:17am UTC](https://forum.mothur.org/t/screen-seqs-removal-of-high-percent-of-sequences/2804 "2016-04-11T08:17:58Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![JuanjoR](https://avatars.discourse-cdn.com/v4/letter/j/94ad74/32.png) [@JuanjoR](https://forum.mothur.org/u/JuanjoR)\
**Post date:** [April 11, 2016, 8:17am UTC](https://forum.mothur.org/t/screen-seqs-removal-of-high-percent-of-sequences/2804/1 "2016-04-11T08:17:58Z")

</div>

Hi all,

I’m suspicious about the following step in MiSeq SOP:

screen.seqs(fasta=stability.trim.contigs.good.unique.align, count=stability.trim.contigs.good.count\_table, summary=stability.trim.contigs.good.unique.summary, start=2, end=7448, maxhomop=8)

After this step, the number of total sequences was reduced from 1.376.092 to 454.912. It seems to me a great deal of sequences removed. Could it be due to some mistake during the filtering process or to possible bad quality from the dataset?

Many thanks,

Juanjo

---

<div class="post-metadata">

**Author:** ![pschloss](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.mothur.org/pschloss/32/4_2.png) [@pschloss](https://forum.mothur.org/u/pschloss)\
**Post date:** [April 12, 2016, 2:17pm UTC](https://forum.mothur.org/t/screen-seqs-removal-of-high-percent-of-sequences/2804/2 "2016-04-12T14:17:00Z")

</div>

You can look at the bad.accnos file that is generated and see what tags are appended to the end of the sequence names that were rejected.

---

<div class="post-metadata">

**Author:** ![Sebastien](https://avatars.discourse-cdn.com/v4/letter/s/f475e1/32.png) [@Sebastien](https://forum.mothur.org/u/Sebastien)\
**Post date:** [April 29, 2016, 1:12pm UTC](https://forum.mothur.org/t/screen-seqs-removal-of-high-percent-of-sequences/2804/3 "2016-04-29T13:12:38Z")

</div>

Hi,

What is a “normal” percentage of sequences remaining after trimming by screen.seqs (based on start and end)?

In addition, do you recommend the use of the “optimize” option? Maybe I am mistaken, but in my thought, I think we force the number of sequences to be retained… In our side, we use the 97.5 tile for the start and the 2.5 tile for the stop from the summary in screen.seqs.

Looking forward to hearign from you.

Seb

---

<div class="post-metadata">

**Author:** ![pschloss](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.mothur.org/pschloss/32/4_2.png) [@pschloss](https://forum.mothur.org/u/pschloss)\
**Post date:** [May 4, 2016, 12:39pm UTC](https://forum.mothur.org/t/screen-seqs-removal-of-high-percent-of-sequences/2804/4 "2016-05-04T12:39:43Z")

</div>

It’s hard to say what to expect. If you’re using MiSeq with paired reads, you shouldn’t lose a lot of sequences. You should also be able to be pretty specific about your start / end positions. Can you post the output of running summary.seqs on the input to screen.seqs?

Pat
