# How to evaluate samplers for inclusion in Stan?

**URL:** <https://discourse.mc-stan.org/t/how-to-evaluate-samplers-for-inclusion-in-stan/38597>\
**Category:** Developers\
**Created:** [January 23, 2025, 7:56pm UTC](https://discourse.mc-stan.org/t/how-to-evaluate-samplers-for-inclusion-in-stan/38597 "2025-01-23T19:56:58Z")\
**Posts on this page:** 1\
**Showing post:** 11

<div class="post-metadata">

**Author:** ![martinmodrak](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/martinmodrak/32/133_2.png) [@martinmodrak](https://discourse.mc-stan.org/u/martinmodrak)\
**Post date:** [January 29, 2025, 6:41am UTC](https://discourse.mc-stan.org/t/how-to-evaluate-samplers-for-inclusion-in-stan/38597/11 "2025-01-29T06:41:49Z")

</div>

> [@Bob\_Carpenter](#):
>
> I’m not sure how we’d do this, but then I’m not the best person to talk to about C++ linking. Did you have something specific in mind? The way to get this kind of feature into Stan is to write a design doc. It’d have to somehow link to how you’d specify a sampler in the interfaces in addition to the very brittle way it’s done now in CmdStan now.
> 
> Right now, the easiest way to use the Stan language with a new sampler is through BridgeStan. Nutpie provides a Stan interface this way. This is also what we’re doing around GIST. You can still use `posterior` in R or `ArviZ` in Python.

I probably focused too narrowly on the technical part - since this is tangential, I made an outline at [Idea: A simple plugin system for CmdStan](https://discourse.mc-stan.org/t/idea-a-simple-plugin-system-for-cmdstan/38632) . But whatever the tech, I think a goal should be that a sampler that is in beta is implemented and documented in such a way that a substantial fraction of Stan users could (if they wanted) try the sampler for their problem with a reasonably short setup (say \<30 minutes). This could totally be achievable with a BridgeStan implementation.

> [@Bob\_Carpenter](#):
>
> I think we’d need the original form from Cook that had a hypothesis test for uniformity. I don’t see how we could do the Talts et al. version as it’s all subjective by-eye evaluations on a dimension-by-dimension basis.

That’s not true, therre are many known tests for discrete uniformity. The gamma statistic from [[2103.10522] Graphical Test for Discrete Uniformity and its Applications in Goodness of Fit Evaluation and Multiple Sample Comparison](https://arxiv.org/abs/2103.10522) is particularl useful IME (and the tail quantiles of the null distribution can be evaluated to use it in a test)

> [@Bob\_Carpenter](#):
>
> n some sense, I’m hoping we don’t need this if we have reference problems with reference answers. I am also interested in measuring samples that may have bias, which are going to fail SBC (e.g., various forms of VI or unadjusted MCMC methods).

True, but even for the approximate ones, it would IMHO be good to know the extent of the miscalibration. I don’t want to push SBC too strongly. It is definitely not a hard requirement IMHO.

---

_[View the full topic](https://discourse.mc-stan.org/t/how-to-evaluate-samplers-for-inclusion-in-stan/38597)._
