# Accessible explanation to the No-U-Turn Sampler

**URL:** <https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695>\
**Category:** Algorithms\
**Created:** [December 5, 2017, 10:00pm UTC](https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695 "2017-12-05T22:00:55Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![markhw](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/markhw/32/5701_2.png) [@markhw](https://discourse.mc-stan.org/u/markhw)\
**Post date:** [December 5, 2017, 10:00pm UTC](https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695/1 "2017-12-05T22:00:55Z")

</div>

In my Bayesian seminar today, we discussed at length how step-size and adapt-delta change the way we explore and sample from the posterior distribution. We were looking at the Hoffman & Gelman (2014) paper, but I’m wondering if there is a more intuitive or accessible explanation of what these hyperparameters do, how they affect how we explore the posterior, what the consequences of doing this is, and the thinking behind it was?

Does anyone know if a blog post or journal article or explanation elsewhere that explains the NUTS in a little more broader, conceptual terms?

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [December 6, 2017, 1:34am UTC](https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695/2 "2017-12-06T01:34:39Z")

</div>

[https://arxiv.org/abs/1701.02434](https://arxiv.org/abs/1701.02434)

---

<div class="post-metadata">

**Author:** ![monnahc](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/monnahc/32/1936_2.png) [@monnahc](https://discourse.mc-stan.org/u/monnahc)\
**Post date:** [December 12, 2017, 6:20pm UTC](https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695/3 "2017-12-12T18:20:55Z")

</div>

Michael’s conceptual intro is great. I also tried to explain it to non-mathy types in this paper:

[http://onlinelibrary.wiley.com/doi/10.1111/2041-210X.12681/abstract](http://onlinelibrary.wiley.com/doi/10.1111/2041-210X.12681/abstract)

I think it’s a nice complement to the other literature, but naturally that’s a biased opinion.

---

<div class="post-metadata">

**Author:** ![markhw](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/markhw/32/5701_2.png) [@markhw](https://discourse.mc-stan.org/u/markhw)\
**Post date:** [December 13, 2017, 12:51am UTC](https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695/4 "2017-12-13T00:51:29Z")

</div>

This is perfect, thanks!

---

<div class="post-metadata">

**Author:** ![tiagocc](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/tiagocc/32/85_2.png) [@tiagocc](https://discourse.mc-stan.org/u/tiagocc)\
**Post date:** [December 13, 2017, 8:47am UTC](https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695/5 "2017-12-13T08:47:39Z")

</div>

This blog post by Richard McElreath is also a very good way to start imo.

> **[Markov Chains: Why Walk When You Can Flow?](http://elevanth.org/blog/2017/11/28/build-a-better-markov-chain/)**
>
> Abstract: If you are still using a Gibbs sampler, you are working too hard for too little result. Newer, better algorithms trade random walks for frictionless flow. In 1989, Depeche Mode was popula…

---

<div class="post-metadata">

**Author:** ![markhw](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/markhw/32/5701_2.png) [@markhw](https://discourse.mc-stan.org/u/markhw)\
**Post date:** [December 13, 2017, 4:03pm UTC](https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695/6 "2017-12-13T16:03:29Z")

</div>

Ah, this is great! This is a perfect first introduction to the sampler, with nice interactives that one can use in a seminar.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [January 9, 2018, 5:25am UTC](https://discourse.mc-stan.org/t/accessible-explanation-to-the-no-u-turn-sampler/2695/7 "2018-01-09T05:25:58Z")

</div>

> [@markhw](#):
>
> step-size and adapt-delta change the way we explore and sample from the posterior distribution.

`adapt_delta` just sets the target “acceptance rate” for the sampler. A higher target acceptance rate means adaptation will find lower step sizes. Once warmup’s done, these are locked in.

How adaptation works has changed over versions. But that target acceptance is now complicated as we’re not using the basic NUTS algorithm.

The main issue you run into is conditioning—the usual bugbear of any kind of gradient-based algorithm. If you get into a location in the posterior where the step size is too large, you get divergences. We only use gradient-based approximations (i.e., first order) of the real posterior curvature, so sometimes we need small step sizes to do that accurately.
