# Modeling unbalanced multivariate outcomes in brms

**URL:** <https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864>\
**Category:** brms\
**Created:** [September 6, 2019, 1:59pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864 "2019-09-06T13:59:55Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [September 6, 2019, 1:59pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864/1 "2019-09-06T13:59:56Z")

</div>

Hi, I’m trying to fit a basic unbalanced multivariate outcomes model in brms that I can then build upon. These data are from a behavioral paradigm (Stroop). The paradigm has two trial events/conditions: congruent and incongruent. However, I only examine responses to correct trials, which means that the data are unbalanced because people commit errors.

Here is an example of how the data might look. There are rows with measurements for one event (e.g., congruent) but not the other event (incongruent). (making up data for example…)

id: participant ID  
trial: trial number… 1 = first trial, 2 = second trial, etc.  
congruent: response for a congruent trial  
incongruent: response for an incongruent trial

NA = empty

| id | trial | congruent | incongruent |
| --- | --- | --- | --- |
| 100 | 1 | 300 | 350 |
| 100 | 2 | 310 | 345 |
| 100 | 3 | 305 | _NA_ |
| 100 | 4 | 310 | _NA_ |
| 101 | 1 | 300 | 350 |
| 101 | 2 | 310 | 345 |
| 101 | 3 | _NA_ | 340 |
| 102 | 1 | 300 | 350 |
| 102 | 2 | 310 | 345 |
| 102 | 3 | 305 | 350 |
| 102 | 4 | 306 | 346 |

I have data on about a hundred participants and a few hundred trials of each condition for everyone.

If I had the same number of observations for congruent and incongruent trials, I could model the data using the following…

```
brm_fit <- brm(
  bf(congruent ~ 0 + (1+trial|p|id)) + 
    bf(incongruent ~ 0 + (1+trial|q|id)),
  data = df_wide,
  chains = 4,
  cores = 4,
  iter = 1000)

```

But, when I try to model the unbalanced data, brms throws out the rows with NA. How can I go about running this model with unbalanced observations?

Thank you very much for any insight!

- Operating System: OS X 10.14.16
- brms Version: 2.9.0

---

<div class="post-metadata">

**Author:** ![jackbailey](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jackbailey/32/14978_2.png) [@jackbailey](https://discourse.mc-stan.org/u/jackbailey)\
**Post date:** [September 7, 2019, 10:31pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864/2 "2019-09-07T22:31:06Z")

</div>

I’m not sure I understand how essential it is that the data are unbalanced, but you could always impute the missing values when you fit your model.

```
brm_fit <- brm(
  bf(congruent | mi() ~ 0 + (1+trial|p|id)) + 
    bf(incongruent | mi() ~ 0 + (1+trial|q|id)),
  data = df_wide,
  chains = 4,
  cores = 4,
  iter = 1000)
```

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [September 7, 2019, 11:44pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864/3 "2019-09-07T23:44:00Z")

</div>

Thanks for your response. I thought of modeling the data as missing, but I’m not sure that is theoretically defensible. I’m not actually “missing” data, because I am using all correct trials. The responses definitely are not missing at random. It’s not like there is a random hardware malfunction for some trials or something similar.

At least that’s how I reasoned that I can’t model the responses as missing. I could be wrong though.

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [September 16, 2019, 3:42pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864/4 "2019-09-16T15:42:16Z")

</div>

@paul.buerkner, I was wondering whether you had any thoughts? I apologize for the direct tag. I’m still banging my head against the wall, and I know this is probably a straightforward reply for you :) Can I even do this in brms?

---

<div class="post-metadata">

**Author:** ![Guido\_Biele](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/guido_biele/32/3414_2.png) [@Guido\_Biele](https://discourse.mc-stan.org/u/Guido_Biele)\
**Post date:** [September 16, 2019, 4:53pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864/5 "2019-09-16T16:53:24Z")

</div>

Does anything speak against converting the data to long-format, which would allow to simply omit the trials with incorrect responses?

something like `rt ~ 1 + (1 + trialtype/trial ) + (1 | id)`

I realize this isn’t exactly what your original formula does, but I am not sure what you were grouping with `|p|` and `|q|`.

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [September 16, 2019, 5:22pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864/6 "2019-09-16T17:22:00Z")

</div>

Thank you for your thoughts.

I am not opposed toward using a long-format data. But, I need to model the varying effects of trial type (congruent, incongruent) as correlated. I’m interested in the between-trialtype covariances. This is why I have been trying to use the multivariate model. I am not sure how to write a multivariate outcomes model using long format… I’ll keep looking!

---

<div class="post-metadata">

**Author:** ![paul.buerkner](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/paul.buerkner/32/3303_2.png) [@paul.buerkner](https://discourse.mc-stan.org/u/paul.buerkner)\
**Post date:** [September 16, 2019, 7:19pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864/7 "2019-09-16T19:19:30Z")

</div>

I would recommend the following model, which uses the long format as @Guido_Biele suggested,  
that is all RTs under each other and a new variable trialtype. With that, we could use

```
rt ~ 0 + trialtype + (1 | trial) + (0 + trialtype | id) 

```

that way you estimate both trial types in the same model while accounting for the dependencies  
across the two observations per trial via a varying intercept.

---

<div class="post-metadata">

**Author:** ![Peter\_Clayson](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/peter_clayson/32/251_2.png) [@Peter\_Clayson](https://discourse.mc-stan.org/u/Peter_Clayson)\
**Post date:** [September 16, 2019, 8:04pm UTC](https://discourse.mc-stan.org/t/modeling-unbalanced-multivariate-outcomes-in-brms/10864/8 "2019-09-16T20:04:10Z")

</div>

Excellent. This worked.

Thank you everyone for your help! I appreciate it.
