# A better unit vector

**URL:** <https://discourse.mc-stan.org/t/a-better-unit-vector/26989>\
**Category:** Developers\
**Tags:** specification\
**Created:** [April 1, 2022, 10:50am UTC](https://discourse.mc-stan.org/t/a-better-unit-vector/26989 "2022-04-01T10:50:22Z")\
**Posts on this page:** 1\
**Showing post:** 30

<div class="post-metadata">

**Author:** ![Seth\_Axen](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/seth_axen/32/14059_2.png) [@Seth\_Axen](https://discourse.mc-stan.org/u/Seth_Axen)\
**Post date:** [May 23, 2022, 10:45pm UTC](https://discourse.mc-stan.org/t/a-better-unit-vector/26989/30 "2022-05-23T22:45:06Z")

</div>

Suppose we want to sample x \in \mathbb{S}^n \subset \mathbb{R}^{n+1} with the spherical embedding approach, i.e. sampling y \in \mathbb{R}^{n+1}, where x = \frac{y}{\lVert y \rVert}. As @betanalpha explained, the log-density “correction” of -\lVert y \rVert^2/2 that Stan uses puts an implicit density on r=\lVert y \rVert.

Specifically, we can think of the transformation from x to y as two steps:

1. augment the density \pi\_x(x) with a prior \pi\_r(r) for r\>0 to get a new joint density \pi\_x(x)\pi\_r( r).
2. bijectively map (x, r) \in \mathbb{S}^n \times R\_{\>0} to \mathbb{R}^{n+1} \backslash \{0\} with y = r x, and use the Jacobian correction -n \log r to get the log-density \log\pi\_y(y) = \log\pi\_x(y/\lVert y \rVert) + \log\pi\_r(\lVert y \rVert) - n\log \lVert y \rVert.

We are free to choose any continuous proper density for \pi\_r(r). Stan implicitly chooses the Chi distribution with n+1 degrees of freedom, which has the log-density \log\pi\_r(r) = n \log(r) -\frac{1}{2} r^2. The n \log r term, which attracts draws to 0 in one expression and repulses in the other, perfectly cancel, so that we are left with \log \pi\_y(y) = -\frac{1}{2} \lVert y\rVert^2.

We still have a singularity at y=0 that we need to avoid, and as noted above, when \pi\_x(x) is concentrated, then \pi\_y(y) has a wedge geometry that is challenging to sample. Due to concentration of measure, for large n, draws near y=0 should be rare, so this shouldn’t be as much of a problem. But for low n, this will manifest with divergences.

An alternative solution to those mentioned so far is to use \log\pi\_r(r) = \frac{1}{2} r^2 + (n + a) \log r for a \ge 0, i.e. a Chi distribution with n+a+1 degrees of freedom. This repels y from y=0, and the degree of repulsion can be tuned by the user by increasing a.

Here we see the log-density for a von Mises distribution with concentration of 100 transformed to the latent space for various values of a:

 ![tmp](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/5/55a581f11bebc567096474322b8c49a44e2e1a17.png)

Yet another alternative is to use \log \pi\_r(r) = -\frac{1}{2} (r-m)^2 + n \log r for m \ge 0:

 ![tmp2](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/d/dec8c6129f49584d8af245d3b553abb4b45f042b.png)

I’m not sure sure how this could be supported in Stan.

---

_[View the full topic](https://discourse.mc-stan.org/t/a-better-unit-vector/26989)._
