# Prior on R2

**URL:** <https://discourse.mc-stan.org/t/prior-on-r2/8291>\
**Category:** rstanarm\
**Created:** [March 30, 2019, 9:29am UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291 "2019-03-30T09:29:08Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![ermeel](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/ermeel/32/4180_2.png) [@ermeel](https://discourse.mc-stan.org/u/ermeel)\
**Post date:** [March 30, 2019, 9:29am UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291/1 "2019-03-30T09:29:08Z")

</div>

Hi Stanimals,

I am trying to follow the logic and derivation behind the idea to put a prior on R^2 for linear regression models. I am trying to write a step by step derivation to fully comprehend it.

I have to admit that I struggle a bit with the current description/derivation in the prior section of [this](https://cran.r-project.org/web/packages/rstanarm/vignettes/lm.html#priors) vignette for rstanarm. In particular the identity of \theta relating it to \rho\_k. I see how this works for the single variable case (best linear predictor interpretation of OLS) but not for the multiple regression case…

Is there an extended version of the vignette available somewhere, or any relevant literature pointers or derivations?

Otherwise I would ask my questions here…

@jonah @bgoodri?

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [March 30, 2019, 1:40pm UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291/2 "2019-03-30T13:40:23Z")

</div>

\boldsymbol{\theta} is a vector of coefficients between the outcome and the columns of \mathbf{Q}, which are orthogonal to each other. So, the coefficient for each is just a correlation times a ratio of the standard deviation of the outcome to the standard deviation of that column of \mathbf{Q} (which is \frac{1}{\sqrt{N - 1}}).

---

<div class="post-metadata">

**Author:** ![hhau](https://avatars.discourse-cdn.com/v4/letter/h/2bfe46/32.png) [@hhau](https://discourse.mc-stan.org/u/hhau)\
**Post date:** [March 30, 2019, 4:56pm UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291/3 "2019-03-30T16:56:38Z")

</div>

The first part of section 2.1 in this paper might be helpful?

> **[High Dimensional Linear Regression via the R2-D2 Shrinkage Prior](https://arxiv.org/abs/1609.00046)**
>
> We propose a new class of priors for linear regression, the R-square induced
> Dirichlet Decomposition (R2-D2) prior. The prior is induced by a Beta prior on
> the coefficient of determination, and then the total prior variance of the
> regression...

---

<div class="post-metadata">

**Author:** ![ermeel](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/ermeel/32/4180_2.png) [@ermeel](https://discourse.mc-stan.org/u/ermeel)\
**Post date:** [March 31, 2019, 5:54pm UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291/4 "2019-03-31T17:54:14Z")

</div>

> [@bgoodri](#):
>
> θ\boldsymbol{\theta} is a vector of coefficients between the outcome and the columns of Q\mathbf{Q}, which are orthogonal to each other. So, the coefficient for each is just a correlation times a ratio of the standard deviation of the outcome to the standard deviation of that column of Q\mathbf{Q} (which is 1√N−1\frac{1}{\sqrt{N - 1}}).

Thanks. Just another related question: Does centering in \mathbf{X} in imply that \mathbf{Q} is centered (in addition to having uncorrelated columns)? If so how can you show that easily? This is the only missing step I need for the derivation to be completed.

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [March 31, 2019, 7:30pm UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291/5 "2019-03-31T19:30:27Z")

</div>

If \mathbf{X} without an intercept has centered columns, the Q factor in its QR decomposition is the same as a Q factor — after its first column has been removed — in the QR decomposition of a design matrix with an intercept in the first column and no centering. And the remaining columns will have mean zero in order to be orthogonal to the removed column.

---

<div class="post-metadata">

**Author:** ![ermeel](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/ermeel/32/4180_2.png) [@ermeel](https://discourse.mc-stan.org/u/ermeel)\
**Post date:** [June 20, 2019, 5:49pm UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291/6 "2019-06-20T17:49:06Z")

</div>

Thanks @bgoodri.

Is it possible to get a self-contained example (Stan and R code) to see what’s actually implemented? I gave it a try, but I am sure this it not yet what you describe in the vignette and above.

```
data {
    int<lower=1> N;
    int<lower=1> K;
    matrix[N,K] X;
    vector[N] y;
    real<lower=0> eta;
    real<lower=0> s_y;
}
transformed data {
    matrix[N,K] Q = qr_thin_Q(X);
    matrix[K,K] R = qr_thin_R(X);
    matrix[K,K] R_inv = inverse(R);
}
parameters {
    cholesky_factor_corr[K+1] L;
    real<lower=0> omega;
}
transformed parameters {
    matrix[K+1,K+1] corr = L*L';
    vector[K] rho = corr[2:,1];
    real<lower=0> sigma_y = omega*s_y;
    vector[K] theta = sqrt(N-1)*sigma_y * rho;
    real R2 = dot_product(rho,rho);
    real<lower=0> sigma=omega*s_y*sqrt(1-R2);
}
model {
    L ~ lkj_corr_cholesky(eta);
    target += -log(square(omega));
    y ~ normal(Q*theta, sigma);
}
generated quantities {
    vector[K] beta = R_inv * theta;
}

```

and here some R-code:

```
library(rstan)
data("clouds", package = "HSAUR3")

ols <- lm(rainfall ~ seeding * (sne + cloudcover + prewetness + echomotion) +
            time, data = clouds)
X <- model.matrix(ols) 
X <- X[,2:ncol(X)]
X_c <- scale(X, scale=F)
y<- clouds$rainfall
stan_data <- list(
  
  N= nrow(X_c),
  K= ncol(X_c),
  y = y,
  X=X_c,
  eta=ncol(X_c)/2,
  s_y=sd(y)
)
fit <- stan("~/Desktop/prior_r2.stan", data=stan_data, chains=4, cores=1, seed=42)

```

Any help is appreciated.

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [June 20, 2019, 6:34pm UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291/7 "2019-06-20T18:34:53Z")

</div>

The Stan code for that is fairly self-contained

> <https://github.com/stan-dev/rstanarm/blob/master/src/stan_files/lm.stan#L89>

  
The R code less so but most of the action is actually in `stan_biglm.fit`  

> <https://github.com/stan-dev/rstanarm/blob/master/R/stan_biglm.fit.R>

---

<div class="post-metadata">

**Author:** ![arya](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/arya/32/1626_2.png) [@arya](https://discourse.mc-stan.org/u/arya)\
**Post date:** [September 18, 2019, 3:12am UTC](https://discourse.mc-stan.org/t/prior-on-r2/8291/8 "2019-09-18T03:12:53Z")

</div>

Thanks for linking that paper. R2-D2 looks pretty cool.

I guess the possible advantage over a Regularized Horseshoe prior would be that you get to specify a prior on R^2 instead of on the proportion of non-zero coefficients. They also mention in that paper that the prior density on the coefficients is unbounded at zero which leads to tighter shrinkage around zero.

Has anyone tried R2-D2 in Stan?
