# Multivariate formula with different number of observations

**URL:** <https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401>\
**Category:** brms\
**Tags:** cognitive-science\
**Created:** [July 5, 2020, 6:02pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401 "2020-07-05T18:02:28Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![GianniRewer](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/giannirewer/32/2835_2.png) [@GianniRewer](https://discourse.mc-stan.org/u/GianniRewer)\
**Post date:** [July 5, 2020, 6:02pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/1 "2020-07-05T18:02:28Z")

</div>

Hi,

I would like to model two response variables (`y1`, `y2`) using `brms` multivariate syntax to (1) model each response with a different distribution family, and (2) model the random effects `g` as correlated. **The number of obs for `y1` and `y2` differ.**

The idea is to have something like:

```
bf_1 <- bf(y1 ~ 1 + (1|ID|g)) + normal()
bf_2 <- bf(y2 ~ 1 + (1|ID|g)) + lognormal()

```

According to the [brms multivariate vignette](https://paul-buerkner.github.io/brms/articles/brms_multivariate.html), it seems that we need the same number of observations for each component in the formula (in the vignette `tarsus` and `back` are two columns in the dataset). Can this assumption be relaxed? A previous user raised the same question [here](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864), but the suggested solution does not make specifying different families possible.

Would you be able to help?

---

<div class="post-metadata">

**Author:** ![martinmodrak](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/martinmodrak/32/133_2.png) [@martinmodrak](https://discourse.mc-stan.org/u/martinmodrak)\
**Post date:** [July 9, 2020, 10:54am UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/2 "2020-07-09T10:54:53Z")

</div>

Sorry, short on time, so just a quick notice - If I understand the problem correctly the `subset` term should let you use only some rows for each outcome (see the help for `brmsformula` for details).

---

<div class="post-metadata">

**Author:** ![GianniRewer](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/giannirewer/32/2835_2.png) [@GianniRewer](https://discourse.mc-stan.org/u/GianniRewer)\
**Post date:** [July 9, 2020, 10:56am UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/3 "2020-07-09T10:56:43Z")

</div>

No worries. Thank you for taking the time! I’ll try to use `subset` and I’ll report back.

---

<div class="post-metadata">

**Author:** ![GianniRewer](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/giannirewer/32/2835_2.png) [@GianniRewer](https://discourse.mc-stan.org/u/GianniRewer)\
**Post date:** [July 10, 2020, 11:19am UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/4 "2020-07-10T11:19:57Z")

</div>

It worked! Using `subset` was the trick. Thank you 🙏🏻

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [August 4, 2020, 9:54am UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/5 "2020-08-04T09:54:54Z")

</div>

I’m curious to see the solution using subset. Can you post the brms formula you used?

---

<div class="post-metadata">

**Author:** ![jroon](https://avatars.discourse-cdn.com/v4/letter/j/fbc32d/32.png) [@jroon](https://discourse.mc-stan.org/u/jroon)\
**Post date:** [August 4, 2020, 2:33pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/6 "2020-08-04T14:33:51Z")

</div>

Found an example here in the last post:

> <https://github.com/paul-buerkner/brms/issues/360>
>
> Currently brms allows us to specify a multivariate model, but only allows a single unified data.frame. This seems to limit cases...

Nice!

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [August 4, 2020, 2:37pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/7 "2020-08-04T14:37:35Z")

</div>

Thanks for posting!

I was hoping that I could have a different number of observations for two different outcomes and still fit the models in brms using the multivariate normal. But, rescor still requires the outcomes to have the same number of observations :( I need rescor = TRUE so I can model the covariances between predictors and residuals.

For now, I’ll keep trying to do this in Stan… but it’s still beyond my skill level… I might be ready to give it another few-month break…

---

<div class="post-metadata">

**Author:** ![jroon](https://avatars.discourse-cdn.com/v4/letter/j/fbc32d/32.png) [@jroon](https://discourse.mc-stan.org/u/jroon)\
**Post date:** [August 4, 2020, 2:40pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/8 "2020-08-04T14:40:19Z")

</div>

Oh I didn’t spot `set_rescor(FALSE) ` there. Also not useful for me so since I too need to model rescor = TRUE 😭

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [August 4, 2020, 2:41pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/9 "2020-08-04T14:41:34Z")

</div>

Haha, sorry to be bearer the bad news! Well, if you figure out a solution for rStan, be sure to let me know! :) (or find a vignette… but I’ve searched high and low for one…)

---

<div class="post-metadata">

**Author:** ![jroon](https://avatars.discourse-cdn.com/v4/letter/j/fbc32d/32.png) [@jroon](https://discourse.mc-stan.org/u/jroon)\
**Post date:** [August 4, 2020, 2:44pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/10 "2020-08-04T14:44:56Z")

</div>

haha no worries - better you told me than I did a bunch of work before noticing.

Hmm you could possibly treat the outcomes with ‘absent rows’ as missing values, but I have not tried this myself as yet.

See the vignette here: [https://cran.r-project.org/web/packages/brms/vignettes/brms\_missings.html#imputation-during-model-fitting](https://cran.r-project.org/web/packages/brms/vignettes/brms_missings.html#imputation-during-model-fitting)

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [August 4, 2020, 2:47pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/11 "2020-08-04T14:47:11Z")

</div>

Ya, I’ve thought of that, but the data aren’t truly “missing”. I do behavioral experiments where the unbalanced nature is occurs as part of the experiments. e.g., if I’m typically only interested in correct trial data (response times). The number of errors is never constant across participants…

---

<div class="post-metadata">

**Author:** ![jroon](https://avatars.discourse-cdn.com/v4/letter/j/fbc32d/32.png) [@jroon](https://discourse.mc-stan.org/u/jroon)\
**Post date:** [August 4, 2020, 2:55pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/12 "2020-08-04T14:55:11Z")

</div>

But does that matter if it is a dependent variable? When you extend the model as in the example there, my understanding is that the ‘missing values’ are treated as parameters - so for missing outcome variables it is similar to a doing a prediction for new data. If you want you could make a model with this parameters included in order to include your uneven outcomes, and then ignore these parameters when interpreting your results… I don’t think it can effect your coefficients other than by allowing you to include all the data

---

<div class="post-metadata">

**Author:** ![paul.buerkner](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/paul.buerkner/32/3303_2.png) [@paul.buerkner](https://discourse.mc-stan.org/u/paul.buerkner)\
**Post date:** [August 4, 2020, 2:57pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/13 "2020-08-04T14:57:19Z")

</div>

Why is it that you have different number of observations?

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [August 4, 2020, 3:06pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/14 "2020-08-04T15:06:52Z")

</div>

I have unequal observations, because I look at response time data based on response accuracy (and only trials where a participant actually responds.

For example, one person could have 300 correct trials and 20 error trials and another could have 290 correct trials and 25 error trials (with 5 discarded non-response trials).

@paul.buerkner you actually helped me with a similar problem before ([Modeling unbalanced multivariate outcomes in brms](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864)), which was referenced above.

My hope is to get an identical output to using  
bf(congruent ~ 0 + (1+trial|p|id)) +  
bf(incongruent ~ 0 + (1+trial|q|id) + rescor(TRUE)

In this case it would be something like  
bf(correct ~ 0 + (1+trial|p|id)) +  
bf(error ~ 0 + (1+trial|q|id) + rescor(TRUE)

That output of this brms fit has all of the (co)variance components that I need, and none that I don’t :) I don’t know if that helps to clarify what I’m after… I can only get it to work if I have the same number of correct and error trials across participants.

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [August 4, 2020, 3:08pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/15 "2020-08-04T15:08:23Z")

</div>

@jroon That is an interesting idea… maybe the issue is that I don’t understand what’s happening when mi() is used, as in… bf(congruent | mi() ~ 0 + (1+trial|p|id))

---

<div class="post-metadata">

**Author:** ![paul.buerkner](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/paul.buerkner/32/3303_2.png) [@paul.buerkner](https://discourse.mc-stan.org/u/paul.buerkner)\
**Post date:** [August 4, 2020, 3:09pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/16 "2020-08-04T15:09:32Z")

</div>

What keeps you from stacking correct and error trials into the same variable (long format) and then modeling that?

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [August 4, 2020, 3:11pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/17 "2020-08-04T15:11:09Z")

</div>

I couldn’t figure out how to get the covariance between residual variances for the two model types… :)

I want the sd(correct\_Intercept), sd(error\_Intercept), sd(correct\_trial), sd(error\_trial) and the covariances among them. I also need the sigma\_correct, sigma\_error, and rescor(correct, error).

---

<div class="post-metadata">

**Author:** ![jroon](https://avatars.discourse-cdn.com/v4/letter/j/fbc32d/32.png) [@jroon](https://discourse.mc-stan.org/u/jroon)\
**Post date:** [August 4, 2020, 3:14pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/18 "2020-08-04T15:14:28Z")

</div>

> [@paul.buerkner](#):
>
> What keeps you from stacking correct and error trials into the same variable (long format) and then modeling that?

Is this the same model however?

If I understand you correctly you mean make one long format output variable and an indicator variable to say which outcome in the wide format that each row corresponds to. Then when you fit that model all outcomes are considered to come from the same univariate distribution (e.g. normal) with the indicator having a moderating effect, whereas in the multivariate version the outcomes are considered to come from a multivariate normal ?

I’ve seen this done in e.g. jags models in the past - but I’ve never understood are these things truly equivalent.

---

<div class="post-metadata">

**Author:** ![paul.buerkner](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/paul.buerkner/32/3303_2.png) [@paul.buerkner](https://discourse.mc-stan.org/u/paul.buerkner)\
**Post date:** [August 4, 2020, 3:41pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/19 "2020-08-04T15:41:25Z")

</div>

You can only define a correlation between two variables when there is a one to one correspondence between values. This is not a brms limitation but how correlations work. So i am unsure how you would define this residual correlation in the first place.

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [August 4, 2020, 3:50pm UTC](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401/20 "2020-08-04T15:50:51Z")

</div>

Hm, all participants have observations for each event type, so it’s not like any participants don’t have a person-specific estimate of correct or error trial. I’m interested in the group-level effects, so I figured as long as everyone had at least a few good observations of each trial type it should be doable…

[Next page](https://discourse.mc-stan.org/t/multivariate-formula-with-different-number-of-observations/16401.md?page=2)
