# Beta-release Bayesian Posterior Database

**URL:** <https://discourse.mc-stan.org/t/beta-release-bayesian-posterior-database/12141>\
**Category:** Publicity\
**Created:** [December 2, 2019, 8:52am UTC](https://discourse.mc-stan.org/t/beta-release-bayesian-posterior-database/12141 "2019-12-02T08:52:39Z")\
**Posts on this page:** 1\
**Showing post:** 9

<div class="post-metadata">

**Author:** ![avehtari](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/avehtari/32/5935_2.png) [@avehtari](https://discourse.mc-stan.org/u/avehtari)\
**Post date:** [December 3, 2019, 5:04pm UTC](https://discourse.mc-stan.org/t/beta-release-bayesian-posterior-database/12141/9 "2019-12-03T17:04:55Z")

</div>

Thanks @betanalpha for the useful comments.

> [@betanalpha](#):
>
> I understand that it would be significantly more work, but what I would really like to see is functionality to give the “posterior” object a function and return a validated estimator (checking cube integrability empirically) with baseline error (estimated empirically) which the user could then use to make principled comparisons to other algorithms. Or alternatively a set of default functions those expectation value estimators have already been run and validated.

Can you make this suggestion more concrete? We’re still in progress of adding tools for comparing new results to “gold standard” and suggestions with as many details as possible are welcome. The idea of storing “gold standard” draws was that the user is not limited to pre-defined set of expectations, but we are happy to get recommendations for any default set.

> [@betanalpha](#):
>
> Tangentially I would also recommend more samples – I believe that I used at least 100,000 for the `stat_comp_benchmark` baseline expectation values. The standard error should really be beaten down to the sub percent level.

Preferably we would want have more than 100k, but the current choice is a compromise. We could sample for some posteriors even more than 100k draws and compute set of pre-defined expectations to save time and memory of the users of the database. Those using Stan could also anytime get more draws if there is doubt that 10k is not enough. As the database is meant for many things, I’m not sure if we would like to have 100k for all.

> [@betanalpha](#):
>
> The other issue is with the definition of “gold\_standard” – what exactly is it? Exact samples carry with them different guarantees than very long dynamic HMC chains, for example. In particular given a multimodal model like LDA I wouldn’t trust MCMC output at all.

Sure, and we don’t currently have “gold standard” for LDA and it’s set to NULL. We think it’s useful to have also models and posteriors which we know to be difficult. Using the keywords and whether “gold standard” exists, it is possible to run tests for only the cases you trust.

> [@betanalpha](#):
>
> One option would be to add tiers beyond gold. For example, gold for having exact PRNG samples and silver for running ensembles of _long_ dynamic HMC chains with verified diagnostics. Or gold for exact, silver for models with simulated data verified with both HMC diagnosis and with SBC, and bronze for just HMC.

The database has information how the “gold standard” has been obtained, but it might also useful to have this kind of tiers as shorthand notation for the most common choices.

> [@betanalpha](#):
>
> Relatedly `gold_standard_info` should also include the version of R, the makevars file, the systemInfo – anything that could influence the values created. It should also include any diagnostic checks that were run so that the user knows how they were validated (for example split Rhat for all coordinate functions below 1.1, no divergences, no E-FMI, etc).

I agree

> [@betanalpha](#):
>
> Finally a word of warning – the `stat_comp_benchmark` models were designed to be a very low bar and I recommend that you consider focusing on models more similar to actual applied use to fill out the database before considering its use to empirically validate methods and the like.

We know that about `stat_comp_benchmark` and that is one reason why we have been working on this.  
We hope this will be a community effort and people will submit pull requests for their favorite models and posteriors. And even if not making the effort for complete pull request, even just recommending models with best practices Stan code and description why the model/posterior is interesting.

---

_[View the full topic](https://discourse.mc-stan.org/t/beta-release-bayesian-posterior-database/12141)._
