# QR Regression Questions

**URL:** <https://discourse.mc-stan.org/t/qr-regression-questions/1415>\
**Category:** Modeling\
**Created:** [July 28, 2017, 3:10am UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415 "2017-07-28T03:10:00Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![betanalpha](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/betanalpha/32/4_2.png) [@betanalpha](https://discourse.mc-stan.org/u/betanalpha)\
**Post date:** [July 28, 2017, 3:10am UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415/1 "2017-07-28T03:10:00Z")

</div>

I’m putting together a case study on the QR regression, [https://github.com/betanalpha/knitr\_case\_studies/tree/master/qr\_regression](https://github.com/betanalpha/knitr_case_studies/tree/master/qr_regression), and I had a few questions that I was hoping some people could answer.

- Who wrote the QR section in the manual? I’m guessing either Ben or Jonah?

- The manual suggests scaling the Q and R matrices by sqrt(N - 1). This approximately makes Q orthonormal, but for unit scaling don’t we want to scale by the full N? This also seems to be the case empirically as demonstrated in the case study.

- Any thoughts on what’s causing the correlations in the transformed slopes? I thought it was the weakly informative prior on the nominal slopes, but I can’t seem to recover an isotropic posterior even with a uniform prior on the slopes. This problem should be simple enough that an isotropic posterior is achievable, no?

Thanks.

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [July 28, 2017, 3:35am UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415/2 "2017-07-28T03:35:01Z")

</div>

(1) Me  
(2) My thinking was that if `Q* = Q * sqrt(N- 1)` then the correlation matrix of `Q*` is the identity matrix. So, the units of the coefficients on `Q*` would be in standard deviations. I don’t think that helps all that much in terms of formulating a prior on the coefficients with respect to `Q*` or `X` though. Scaling by `N` is another option. Not scaling at all seems to be a bad idea for large `N`.  
(3) Did you center both columns of `X` before decomposing it?

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [July 28, 2017, 3:39am UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415/3 "2017-07-28T03:39:19Z")

</div>

One thing that needs to be said, which I didn’t say in the manual is that under the QR decomposition, the _last_ coefficient on `Q` or `Q*` is proportional to the last coefficient on `X`. So, if you only care / have informative prior information about one coefficient, it should be put last in `X` and then rescale your prior accordingly.

Also, I have since come to having the hunch that a polar decomposition would be better than a QR decomposition.

---

<div class="post-metadata">

**Author:** ![betanalpha](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/betanalpha/32/4_2.png) [@betanalpha](https://discourse.mc-stan.org/u/betanalpha)\
**Post date:** [July 28, 2017, 3:48am UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415/4 "2017-07-28T03:48:25Z")

</div>

Thanks.

Not scaling is definitely a bad idea.

I did not center the column of X before decomposing – I guess the QR needs to be done around the centered columns to completely decouple the model?

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [July 28, 2017, 4:11am UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415/5 "2017-07-28T04:11:35Z")

</div>

If you put the intercept into `X` first and then QR it, that is equivalent to doing QR on the centered `X` without the intercept.

---

<div class="post-metadata">

**Author:** ![betanalpha](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/betanalpha/32/4_2.png) [@betanalpha](https://discourse.mc-stan.org/u/betanalpha)\
**Post date:** [July 28, 2017, 8:09pm UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415/6 "2017-07-28T20:09:25Z")

</div>

Yup, centering did the trick.

---

<div class="post-metadata">

**Author:** ![noah](https://avatars.discourse-cdn.com/v4/letter/n/a87d85/32.png) [@noah](https://discourse.mc-stan.org/u/noah)\
**Post date:** [July 29, 2017, 3:48am UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415/7 "2017-07-29T03:48:30Z")

</div>

Michael,

I noticed a small discrpancy between your writeup and the Stan manual

- In the Stan manual, Q is scaled by sqrt(N-1). This creates a matrix with a standard deviation of 1 for all columns

- In your writeup, Q is scaled by N. This has a standard deviation very far from 1

Note: Data was centered, but not scaled, before doing the QR decomposition.

---

<div class="post-metadata">

**Author:** ![betanalpha](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/betanalpha/32/4_2.png) [@betanalpha](https://discourse.mc-stan.org/u/betanalpha)\
**Post date:** [July 29, 2017, 4:47pm UTC](https://discourse.mc-stan.org/t/qr-regression-questions/1415/8 "2017-07-29T16:47:01Z")

</div>

See the above discussion – there’s trade off being scaling the variance and the mean of the transformed distribution.
