# Unifrac with groups vs count file

**URL:** <https://forum.mothur.org/t/unifrac-with-groups-vs-count-file/20831>\
**Category:** Commands in mothur\
**Created:** [January 17, 2021, 11:07pm UTC](https://forum.mothur.org/t/unifrac-with-groups-vs-count-file/20831 "2021-01-17T23:07:38Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![lsteinberg](https://avatars.discourse-cdn.com/v4/letter/l/919ad9/32.png) [@lsteinberg](https://forum.mothur.org/u/lsteinberg)\
**Post date:** [January 17, 2021, 11:07pm UTC](https://forum.mothur.org/t/unifrac-with-groups-vs-count-file/20831/1 "2021-01-17T23:07:39Z")

</div>

I ran unweighted Unifrac on my sequences with either a group file or a count file and got different values for pairwise comparisons between groups. I’m having trouble understanding why this would be? I assume you would need a names or count file for weighted Unifrac because abundances are accounted for, but I thought Unifrac unweighted did not consider abundances? Does anyone have insight on this? Is the recommendation to run unweighted Unifrac with a group file or a count file?  
(I double-checked my log files and there is no mismatch between name, group, and count files)

---

<div class="post-metadata">

**Author:** ![westcott](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.mothur.org/westcott/32/18_2.png) [@westcott](https://forum.mothur.org/u/westcott)\
**Post date:** [January 21, 2021, 5:52pm UTC](https://forum.mothur.org/t/unifrac-with-groups-vs-count-file/20831/2 "2021-01-21T17:52:51Z")

</div>

Mothur does not use the abundances in the unweighted calculation, but the samples the unique sequences represent are needed. Without the name file mothur is unable to map the unique reads to all the samples they represents.

Consider the following:

seq1 group1  
seq2 group2  
seq3 group3  
seq4 group2  
seq5 group1  
seq6 group3

seq1 seq1,seq2  
seq3 seq3,seq4  
seq5 seq5  
seq6 seq6

In the tree seq1 represents reads from group1 and group2, but without the name file mothur would only see group1.

To find the UW value mothur finds the total branch length and the unique branch length for each pairing. Since a leaf node in the tree may represent sequences from several groups, leaving out the name file causes the unique branch total to be artificially inflated.

---

<div class="post-metadata">

**Author:** ![lsteinberg](https://avatars.discourse-cdn.com/v4/letter/l/919ad9/32.png) [@lsteinberg](https://forum.mothur.org/u/lsteinberg)\
**Post date:** [January 21, 2021, 7:11pm UTC](https://forum.mothur.org/t/unifrac-with-groups-vs-count-file/20831/3 "2021-01-21T19:11:13Z")

</div>

OMG, that is so simple and makes so much sense. Thank you!

---

<div class="post-metadata">

**Author:** ![system](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.mothur.org/system/32/2_2.png) [@system](https://forum.mothur.org/u/system)\
**Post date:** [January 31, 2021, 7:11pm UTC](https://forum.mothur.org/t/unifrac-with-groups-vs-count-file/20831/4 "2021-01-31T19:11:13Z")

</div>

This topic was automatically closed 10 days after the last reply. New replies are no longer allowed.
