# Is Ragged Array allowed in Stan?

**URL:** <https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752>\
**Category:** General\
**Created:** [September 25, 2018, 4:13am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752 "2018-09-25T04:13:48Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![serenejiang](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/serenejiang/32/3252_2.png) [@serenejiang](https://discourse.mc-stan.org/u/serenejiang)\
**Post date:** [September 25, 2018, 4:13am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/1 "2018-09-25T04:13:48Z")

</div>

The Stan manual says that ragged arrays are not allowed, but I also found this github tutorial on ragged arrays spec ([https://github.com/stan-dev/stan/wiki/Ragged-array-spec](https://github.com/stan-dev/stan/wiki/Ragged-array-spec)). The github tutorial seems to be more latest than the manual; does this mean that ragged array is now allowed in Stan?

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [September 25, 2018, 4:25am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/2 "2018-09-25T04:25:01Z")

</div>

No

---

<div class="post-metadata">

**Author:** ![ermeel](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/ermeel/32/4180_2.png) [@ermeel](https://discourse.mc-stan.org/u/ermeel)\
**Post date:** [September 25, 2018, 4:33am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/3 "2018-09-25T04:33:27Z")

</div>

The Stan 2.18. user guide states in section 9:

> Stan does not directly support either sparse or ragged data structures, though both can be accommodated with some programming effort.

Check under assets here:

> **[Releases · stan-dev/stan](https://github.com/stan-dev/stan/releases)**
>
> Stan development repository. The master branch contains the current release. The develop branch contains the latest stable development. See the Developer Process Wiki for details. - stan-dev/stan

---

<div class="post-metadata">

**Author:** ![jjramsey](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jjramsey/32/3319_2.png) [@jjramsey](https://discourse.mc-stan.org/u/jjramsey)\
**Post date:** [September 25, 2018, 11:20am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/4 "2018-09-25T11:20:16Z")

</div>

> [@serenejiang](#):
>
> The github tutorial

That’s not a tutorial. That’s a proposal for a possible _future_ feature of Stan.

---

<div class="post-metadata">

**Author:** ![serenejiang](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/serenejiang/32/3252_2.png) [@serenejiang](https://discourse.mc-stan.org/u/serenejiang)\
**Post date:** [September 25, 2018, 4:38pm UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/5 "2018-09-25T16:38:48Z")

</div>

I see. No wonder I ran into a lot of syntax errors when trying to implementing those codes.

---

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [September 25, 2018, 4:54pm UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/6 "2018-09-25T16:54:14Z")

</div>

OTOH there are very good ways to implement ragged arrays in Stan. For example:

```
int n_total_entries;
vector[n_total_entries] flat_array;

int n_l1_entries;
int starts[n_l1_entries];
int stops[n_l1_entries];
int lengths[n_l1_entries];

```

Then to get the `k`'th vector:

```
vector[lengths[k]] v = flat_array[starts[k]:stops[k]];

```

---

<div class="post-metadata">

**Author:** ![serenejiang](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/serenejiang/32/3252_2.png) [@serenejiang](https://discourse.mc-stan.org/u/serenejiang)\
**Post date:** [September 25, 2018, 5:40pm UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/7 "2018-09-25T17:40:02Z")

</div>

@sakrejda Thank you! Do you have suggestions for a ragged array of matrices, where each matrix has different size? I am very new to Stan programming; any suggestion or reference would be appreciated!

---

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [September 25, 2018, 5:46pm UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/8 "2018-09-25T17:46:57Z")

</div>

What you want to do with them in Stan and what are the constraints on their sizes? There’s no general answer but there are really good specific answers.

---

<div class="post-metadata">

**Author:** ![serenejiang](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/serenejiang/32/3252_2.png) [@serenejiang](https://discourse.mc-stan.org/u/serenejiang)\
**Post date:** [September 25, 2018, 6:45pm UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/9 "2018-09-25T18:45:53Z")

</div>

@sakrejda I want to input into Stan an array of matrices of size [num\_subjects, num\_basis, num\_visits]. In other words, there is a corresponding matrix for each subject, and each matrix has number of rows = number of basis,  
and number of columns = number of visits. Note that the number of basis is the same for all subjects, but the number of visits is different. I record the number of visits for each subject in a vector called V. I am wondering whether I could extract the value from the vector V when loading the array of matrices as well. After loading such data, then I would perform MCMC for each subject based on their matrices of basis information. Is this clear enough?

---

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [September 25, 2018, 7:39pm UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/10 "2018-09-25T19:39:29Z")

</div>

Yes that’s clear. That’s only ragged in one dimension so in Stan I would stack subjects and visits:

```
int num_subjects;
int num_visits[num_subjects];
int subject_starts[n_subjects]; // row of x where each subject starts
int subject_stops[n_subjects]; // row of x where each subject ends

int num_basis;
matrix[sum(num_visits), num_basis] X;

```

The rest depends on how you use X. If it’s a model matrix then you have a pretty  
standard (G)LM setup with X as the model matrix.

---

<div class="post-metadata">

**Author:** ![wds15](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/wds15/32/908_2.png) [@wds15](https://discourse.mc-stan.org/u/wds15)\
**Post date:** [September 25, 2018, 7:42pm UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/11 "2018-09-25T19:42:55Z")

</div>

My preference by now is to not have a starts and stops vector, but rather have a single one which has `n+1` elements whenever you have `n` items. In each entry `i` you have the starting position of the `i`th element and the `i+1`th element is one past the end of the `i`th one (which happens to be the start of the `i+1`th one). This way you keep things tidy with a single index vector instead of two… which makes it less cluttered for me.

---

<div class="post-metadata">

**Author:** ![rsmall](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/rsmall/32/4395_2.png) [@rsmall](https://discourse.mc-stan.org/u/rsmall)\
**Post date:** [March 22, 2019, 2:39am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/12 "2019-03-22T02:39:06Z")

</div>

I have a similar problem in which I need to access a distance matrix \mathbf{D}\_j for each family to define a correlation structure via Gaussian Process Regression: \mathbf{K}\_j=\exp\left[-\rho\mathbf{D}\_j\right] (element-wise). Any ideas as to how to go about passing/building each \mathbf{D}\_j in Stan using clever indexing? I’m sorry, but I don’t follow the examples given above.

I can build \mathbf{K}\_j “on demand” within the model or transformed parameters block, but I don’t want to have to continuously rebuild each \mathbf{D}\_j - as I have ~10,000 families this will likely be very slow.

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [March 22, 2019, 2:45am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/13 "2019-03-22T02:45:09Z")

</div>

Need more info here. Are the distance matrices fixed and so \rho is the only unknown? Do you mean that you have an array of J distance matrices, like `matrix[T, T] D[J];`?

---

<div class="post-metadata">

**Author:** ![rsmall](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/rsmall/32/4395_2.png) [@rsmall](https://discourse.mc-stan.org/u/rsmall)\
**Post date:** [March 22, 2019, 2:54am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/14 "2019-03-22T02:54:48Z")

</div>

Yes, my distance matrices are fixed - I am trying to learn the unknown tuning parameter \rho. If I understand your question correctly, then my T would essentially an array of ints giving the dimensionality of \mathbf{D}\_j, e.g. T=c(1, 1, 2, 5, 7). To put it another way, I have (in R) a list D in which D[[j]] accesses \mathbf{D}\_j with size n\_j \times n\_j.

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [March 22, 2019, 3:00am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/15 "2019-03-22T03:00:14Z")

</div>

OK, I understand now. That is going to be annoying currently. I think the least annoying option would be to compute the maximum value of `T` and declare the thing in data like `matrix[max(T), max(T)] D[J];` and then in R put all the distances for the j-th family in the top left of the j-th matrix, leaving any unused cells as \infty.

Then in your model block, do something like

```
model {
  for (j in 1:J) {
    matrix[T[j], T[j]] K = exp(-rho * block(D[j], 1, 1, T[j], T[j]));
    // likelihood contribution of the j-th family
  }
}

```

---

<div class="post-metadata">

**Author:** ![rsmall](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/rsmall/32/4395_2.png) [@rsmall](https://discourse.mc-stan.org/u/rsmall)\
**Post date:** [March 22, 2019, 3:08am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/16 "2019-03-22T03:08:03Z")

</div>

Makes sense - thanks! Since the `block` command just selects the appropriate entries of the “padded” `D[j]` I could just as easily define the unused cells as zero, say?

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [March 22, 2019, 3:10am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/17 "2019-03-22T03:10:24Z")

</div>

You could, but if you make a mistake, it will show up more destructively if you make the unused cells infinity.

Actually negative infinity would be better due to the `exp(-rho * ...)` thing.

---

<div class="post-metadata">

**Author:** ![rsmall](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/rsmall/32/4395_2.png) [@rsmall](https://discourse.mc-stan.org/u/rsmall)\
**Post date:** [March 22, 2019, 3:25am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/18 "2019-03-22T03:25:02Z")

</div>

Great idea, thanks. I assume that `block` also works on arrays of integers as well?

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [March 22, 2019, 4:15am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/19 "2019-03-22T04:15:26Z")

</div>

No, but you can use `1:T[j]`.

---

<div class="post-metadata">

**Author:** ![rsmall](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/rsmall/32/4395_2.png) [@rsmall](https://discourse.mc-stan.org/u/rsmall)\
**Post date:** [March 27, 2019, 2:59am UTC](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752/20 "2019-03-27T02:59:16Z")

</div>

Thanks for your help on this, Ben. Just wanted to let you know that the model compiles and runs correctly.

[Next page](https://discourse.mc-stan.org/t/is-ragged-array-allowed-in-stan/5752.md?page=2)
