# Dirichlet Priors

**URL:** <https://discourse.mc-stan.org/t/dirichlet-priors/20544>\
**Category:** General\
**Created:** [January 31, 2021, 3:56pm UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544 "2021-01-31T15:56:20Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![CerulloE](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/cerulloe/32/15968_2.png) [@CerulloE](https://discourse.mc-stan.org/u/CerulloE)\
**Post date:** [January 31, 2021, 3:56pm UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544/1 "2021-01-31T15:56:20Z")

</div>

In [section 20.2](https://mc-stan.org/docs/2_26/stan-users-guide/reparameterizations.html) of the manual, some examples for priors for the beta distribution are given (from

```stan
parameters {
  real<lower=0,upper=1> phi;
  real<lower=0.1> lambda;
  ...
transformed parameters {
  real<lower=0> alpha = lambda * phi;
  real<lower=0> beta = lambda * (1 - phi);
  ...
model {
  phi ~ beta(1, 1); // uniform on phi, could drop
  lambda ~ pareto(0.1, 1.5);
  for (n in 1:N)
    theta[n] ~ beta(alpha, beta);
  ...

```

However, for the Dirichlet only the form is specificed but not example prior distributions,

```stan

parameters {
  simplex[K] phi;
  real<lower=0> kappa;
  simplex[K] theta[N];
  ...
transformed parameters {
  vector[K] alpha = kappa * phi;
  ...
}
model {
  phi ~ ...;
  kappa ~ ...;
  for (n in 1:N)
    theta[n] ~ dirichlet(alpha);

```

Obviously this will depend on the problem, but in general would be some sensible choices for `phi` and `kappa` ?

---

<div class="post-metadata">

**Author:** ![torkar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/torkar/32/1355_2.png) [@torkar](https://discourse.mc-stan.org/u/torkar)\
**Post date:** [February 10, 2021, 8:25am UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544/2 "2021-02-10T08:25:07Z")

</div>

Sorry for taking so long to reply. I’m a bit uncertain myself of suitable priors, but my hope is that @martinmodrak could help, or find the expertise to do it :)

---

<div class="post-metadata">

**Author:** ![CerulloE](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/cerulloe/32/15968_2.png) [@CerulloE](https://discourse.mc-stan.org/u/CerulloE)\
**Post date:** [February 10, 2021, 7:45pm UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544/3 "2021-02-10T19:45:38Z")

</div>

Thanks, I ended up doing:

```stan
parameters {
  simplex[K] phi;
  real kappa;
  simplex[K] theta[N];
.....
}
transformed parameters {
  vector[K] alpha = exp(kappa) * phi;
....
}
model {
  phi ~ beta(1, 1); // redundant
  kappa ~ normal(0, 5) ;// allows a lot of variation in the dirichlet hyperparameters 
  for (n in 1:N)
    theta[n] ~ dirichlet(alpha);
.....
}

```

---

<div class="post-metadata">

**Author:** ![torkar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/torkar/32/1355_2.png) [@torkar](https://discourse.mc-stan.org/u/torkar)\
**Post date:** [February 11, 2021, 6:23am UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544/4 "2021-02-11T06:23:23Z")

</div>

I guess the prior predictive checks looked sane?

---

<div class="post-metadata">

**Author:** ![martinmodrak](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/martinmodrak/32/133_2.png) [@martinmodrak](https://discourse.mc-stan.org/u/martinmodrak)\
**Post date:** [February 11, 2021, 3:41pm UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544/5 "2021-02-11T15:41:35Z")

</div>

So the combination of

```stan
vector[K] alpha = exp(kappa) * phi;

```

and

```stan
kappa ~ normal(0, 5) ;

```

is indeed extremely wide - but hey, as long as you have enough data to constrain it, it should not be a problem. To be safe, I would also consider running the model with a narrower (say `normal(0,1)`) prior on `kappa` and see if your inferences change. If not, you are probably good.

Here is a quick code I wrote to visualise bivariate marginal distributions from the prior to let me better understand it (you could work this out analytically, but I write code faster than I do math, so I chose simulation :-)

```r
N <- 1e5
K <- 10

res <- array(NA_real_, dim = c(K, N))
kappa <- rnorm(N, 0, 5)
phi <- MCMCpack::rdirichlet(N, rep(1, K)) #Uniform simplex
for(i in 1:N) {
  res[,i] <- MCMCpack::rdirichlet(1, exp(kappa[i]) * phi[i,])
}

print(paste0(sum(is.na(res)), " NAs"))

df_to_plot <- data.frame(x1 = res[1,], x2 = res[2,]) %>% filter(!is.na(x1), !is.na(x2))       
ggplot(df_to_plot, aes(x1, x2)) + geom_hex() + scale_fill_continuous(trans = "log10")

```

So for `K = 10` the bivariate implied prior is:

 ![obrazek](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/8/85fc9a264109749d12dd245f4da20c30e34b2d75.png)

which is pretty concentrated at “both almost zero” and “single almost 1, other almost zero” (not the log scale on fill)

Best of luck with your model.

---

<div class="post-metadata">

**Author:** ![CerulloE](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/cerulloe/32/15968_2.png) [@CerulloE](https://discourse.mc-stan.org/u/CerulloE)\
**Post date:** [February 11, 2021, 5:17pm UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544/6 "2021-02-11T17:17:12Z")

</div>

Thanks!

This model is part of a more complicated model, which involves a multivariate ordered probit, and i’m using a hierarchical induced-dirichlet model on the latent cutpoints. For prior predictive checks for this part of the model, I have been plotting the posteriors for the cumulative probabilities. I have incorporated domain expertise into other aspects on the model, and I don’t want to for this part of the model which uses ordered regression (as there isnt much available). In this case would I want the priors for the cumulative probabilities to be approximately uniform? Or, should I not look at the cumulative probabilities for prior predictive checks? I guess just because these are uniform it doesn’t mean it is less informative, since the priors might be “forcing” them to be more uniform…

I have 17 categories (K=17) and have plotted:

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/a/aec3212639a5e4dd021bed245c1d320d3465213b.png)

It looks like the N(0,5) would more easily allow 2 out of 17 of the categories to have most of the weight compared to the N(0, 1)?

---

<div class="post-metadata">

**Author:** ![CerulloE](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/cerulloe/32/15968_2.png) [@CerulloE](https://discourse.mc-stan.org/u/CerulloE)\
**Post date:** [February 12, 2021, 10:02am UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544/7 "2021-02-12T10:02:51Z")

</div>

> [@martinmodrak](#):
>
> is indeed extremely wide - but hey, as long as you have enough data to constrain it, it should not be a problem. To be safe, I would also consider running the model with a narrower (say `normal(0,1)` )

I actually get better mixing and less divergent transitions (none) with normal(0, 5) vs normal(0,1) on kappa. I guess N(0, 1) is conflicting with the likelihood too much as the posterior median for kappa is 4.81 with this prior.

Posterior draw summaries:

Normal(0,5)

![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/b/b25749c2987bf27fa9391413892429a4e6045551.png)

Normal(0, 1):

![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/c/c7a43a160b736f52113c64c73df75cb573ee66fe.png)

---

<div class="post-metadata">

**Author:** ![martinmodrak](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/martinmodrak/32/133_2.png) [@martinmodrak](https://discourse.mc-stan.org/u/martinmodrak)\
**Post date:** [February 12, 2021, 10:34am UTC](https://discourse.mc-stan.org/t/dirichlet-priors/20544/8 "2021-02-12T10:34:28Z")

</div>

> [@CerulloE](#):
>
> I actually get better mixing and less divergent transitions (none) with normal(0, 5) vs normal(0,1) on kappa.

Experiment beats my speculation :-). Yes this looks like indeed normal(0,1) would be a bad choice and normal(0,5) is better.

Best of luck with your model!
