# Interpreting output of multiple comparisons using loo

**URL:** <https://discourse.mc-stan.org/t/interpreting-output-of-multiple-comparisons-using-loo/5846>\
**Category:** Modeling\
**Tags:** loo, interpret-results\
**Created:** [October 2, 2018, 3:08pm UTC](https://discourse.mc-stan.org/t/interpreting-output-of-multiple-comparisons-using-loo/5846 "2018-10-02T15:08:04Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![simon.dp](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/simon.dp/32/2474_2.png) [@simon.dp](https://discourse.mc-stan.org/u/simon.dp)\
**Post date:** [October 2, 2018, 3:08pm UTC](https://discourse.mc-stan.org/t/interpreting-output-of-multiple-comparisons-using-loo/5846/1 "2018-10-02T15:08:04Z")

</div>

If I use the loo function compare() to compare two models, I get how much the second model is better and standard errors of the measurement.

So if I run and get:

```
compare(loom0V,loom1V)

elpd_diff se 
     31.3 8.1 

```

I would say that the second model seems to be roughly 4 SE’s better.

But exactly how do I interpret a comparison of multiple models at the same time - specifically:

```
compare(loom0V,loom1V,loom1Vb,loom2Va,loom2Vai)

         elpd_diff elpd_loo se_elpd_loo p_loo se_p_loo looic se_looic
loom2Va 0.0 -5146.2 41.5 32.1 0.4 10292.4 83.0 
loom2Vai -4.0 -5150.2 41.7 39.1 0.5 10300.4 83.3 
loom1Vb -35.8 -5182.0 40.8 31.0 0.4 10364.1 81.5 
loom1V -171.6 -5317.8 39.4 25.1 0.3 10635.6 78.8 
loom0V -202.9 -5349.1 38.7 23.9 0.3 10698.3 77.3

```

I am all of a sudden getting a lot more info and I am not so sure… What can I conclude in this case?

This may be a question answered elsewhere, but I have not found a simple answer so far so your input is valued.

---

<div class="post-metadata">

**Author:** ![avehtari](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/avehtari/32/5935_2.png) [@avehtari](https://discourse.mc-stan.org/u/avehtari)\
**Post date:** [October 3, 2018, 7:15am UTC](https://discourse.mc-stan.org/t/interpreting-output-of-multiple-comparisons-using-loo/5846/2 "2018-10-03T07:15:20Z")

</div>

> [@simon.dp](#):
>
> I would say that the second model seems to be roughly 4 SE’s better.

Yes.

> [@simon.dp](#):
>
> ```no-highlight
> compare(loom0V,loom1V,loom1Vb,loom2Va,loom2Vai)
> 
> elpd_diff elpd_loo se_elpd_loo p_loo se_p_loo looic se_looic
> 
> ```

Unfortunately this is missing diff\_se, which would then give you the same information as when comparing two models. The difference is computed to the model with highest log predictive density (elpd\_loo). You can still use this see the order, and then you can compare two models at time to check diff\_se’s. Before seeing other diff\_se’s my guess is that there is uncertainty about the difference between loom2Va and loom2Vai, but these models provide clearly better predictive performance than others.

Changing this output to be more clear is on our (with @jonah) todo list.

---

<div class="post-metadata">

**Author:** ![simon.dp](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/simon.dp/32/2474_2.png) [@simon.dp](https://discourse.mc-stan.org/u/simon.dp)\
**Post date:** [October 3, 2018, 10:39am UTC](https://discourse.mc-stan.org/t/interpreting-output-of-multiple-comparisons-using-loo/5846/3 "2018-10-03T10:39:43Z")

</div>

That makes sense - thank you for your response, [avehtari](https://discourse.mc-stan.org/u/avehtari).

I have just analysed them against each other and it seems like you have some good intuition regarding the two bigger models being better than the rest, but similar to each other. May I ask you which way you would prefer this reported?

One way could be to plot all the models against the null model (loom0V) with 2 SEs as error bars:  
 ![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/e/e0bb81b960a840254181e065bbfd6af898d0ea8e.png)

In this case, however, it seems that there is barely any difference between 3 of the models.

When I compare them pairwise, one can see that the 2 complex models are at least 2 SE’s better than the best model with one predictor, but this may look more confusing.  
 ![](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/1/1c9d4e8f88d4908a28719ec736814c3bd77d4c2e.png)

Do you have any preference or would you choose a completely different way of reporting this?

Edit: While writing this, I was thinking that one could also do as the table in my original post and compare all models against the best model. I think I would prefer that myself:  
 ![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/7/732a9203ad3aaf2e58117e060c741bb5951dd3f6.png)

---

<div class="post-metadata">

**Author:** ![avehtari](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/avehtari/32/5935_2.png) [@avehtari](https://discourse.mc-stan.org/u/avehtari)\
**Post date:** [October 3, 2018, 12:56pm UTC](https://discourse.mc-stan.org/t/interpreting-output-of-multiple-comparisons-using-loo/5846/4 "2018-10-03T12:56:08Z")

</div>

> [@simon.dp](#):
>
> Edit: While writing this, I was thinking that one could also do as the table in my original post and compare all models against the best model. I think I would prefer that myself:

Yes, this is what I recommend, too. And presenting the result as a plot like you did is much better than as a table. It’s easy to quickly see the differences!

You may further consider whether you would have some application specific measure to give more interpretable calibration of the model differences.
