# How to vectorize for multi-variate regression

**URL:** <https://discourse.mc-stan.org/t/how-to-vectorize-for-multi-variate-regression/1569>\
**Category:** General\
**Tags:** stanc\
**Created:** [August 15, 2017, 5:48am UTC](https://discourse.mc-stan.org/t/how-to-vectorize-for-multi-variate-regression/1569 "2017-08-15T05:48:31Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![bhomass](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bhomass/32/407_2.png) [@bhomass](https://discourse.mc-stan.org/u/bhomass)\
**Post date:** [August 15, 2017, 5:48am UTC](https://discourse.mc-stan.org/t/how-to-vectorize-for-multi-variate-regression/1569/1 "2017-08-15T05:48:31Z")

</div>

In section 9.1, the manual shows an example of vectorization for a single variable linear regression

```
data {
  int<lower=0> N;
  vector[N] x;
  vector[N] y;
}
parameters {
  real alpha;
  real beta;
  real<lower=0> sigma;
} model {
  y ~ normal(alpha + beta * x, sigma);
}

```

where y can also be written in unvectorized form

```
for (n in 1:N)
        y[n] ~ normal(alpha + beta * x[n], sigma);

```

The advantage is the vectorized form is apparently faster, which is what I am trying to do.

for my case, its multi-variate and I use a custom log likelihood

```
data {
	int N; // number of users
	int T; // number of days in campaign
	int K; // number of predictors
	
	int y[N, T]; // outcome
	row_vector[K] z[N, T]; 

	real alpha_tilda;
}
parameters {
	vector<lower=0>[K] gamma;
}

model {
	gamma ~ normal(0, 0.1);
	for(n in 1:N) {
		y[n] ~ iLogLikelihood(T, alpha_tilda, gamma, z[n]);
	}
}

```

Is this also vectorizable?

---

<div class="post-metadata">

**Author:** ![bbbales2](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bbbales2/32/77_2.png) [@bbbales2](https://discourse.mc-stan.org/u/bbbales2)\
**Post date:** [August 15, 2017, 1:29pm UTC](https://discourse.mc-stan.org/t/how-to-vectorize-for-multi-variate-regression/1569/2 "2017-08-15T13:29:11Z")

</div>

I don’t think it’s automatic for custom likelihoods.

You’ll have to provide a custom implementation for vector types.

By the way loops in Stan are fast (they convert straight over to C++ loops). Where vectorization is handy is for autodiff. If your likelihood has repeated expensive operations that can be shared between likelihood evaluations, then you’ll probably benefit from writing a big custom likelihood. If not, then the benefit is mostly cosmetic.

---

<div class="post-metadata">

**Author:** ![bhomass](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bhomass/32/407_2.png) [@bhomass](https://discourse.mc-stan.org/u/bhomass)\
**Post date:** [August 15, 2017, 5:28pm UTC](https://discourse.mc-stan.org/t/how-to-vectorize-for-multi-variate-regression/1569/3 "2017-08-15T17:28:08Z")

</div>

I thought it would make a significant difference in speed due to the manual text

“In addition to being more concise, the vectorized form is much faster”

---

<div class="post-metadata">

**Author:** ![bbbales2](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bbbales2/32/77_2.png) [@bbbales2](https://discourse.mc-stan.org/u/bbbales2)\
**Post date:** [August 15, 2017, 6:48pm UTC](https://discourse.mc-stan.org/t/how-to-vectorize-for-multi-variate-regression/1569/4 "2017-08-15T18:48:33Z")

</div>

There are a couple things.

If there’s shared computation, vector versions of functions can be faster (this happens with multivariate normals and how the matrix solves are handled).

If you can replace loops and such with matrix/matrix or matrix/vector ops you’ll be better off because the autodiff has been optimized for those (less autodiff variables are created, and the necessary derivatives are hard coded in – check [https://arxiv.org/abs/1509.07164](https://arxiv.org/abs/1509.07164) for more details).

Otherwise it’s probly cosmetic. Go ahead and try it if it’s not too hard, just saying that basic Stan loops are pretty fast as is.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [August 17, 2017, 1:38am UTC](https://discourse.mc-stan.org/t/how-to-vectorize-for-multi-variate-regression/1569/5 "2017-08-17T01:38:47Z")

</div>

> [@bhomass](#):
>
> “In addition to being more concise, the vectorized form is much faster”

That is correct as written. The chapter on efficiency tries to explain why, but I can summarize:

1. vectorized built-in operations cut down on the expression graph size and at the same time replace virtual function calls (pointer chasing, require branch prediction) with regular function calls

2. you can avoid recomputing shared operations

These won’t hold for your own code unless you’re careful to compute and reuse shared expressions. And even so, you can’t write custom derivatives to really unfold everything here.
