# ADVI / Rats example / Adagrad

**URL:** <https://discourse.mc-stan.org/t/advi-rats-example-adagrad/1859>\
**Category:** Algorithms\
**Tags:** variational-bayes\
**Created:** [September 12, 2017, 3:11am UTC](https://discourse.mc-stan.org/t/advi-rats-example-adagrad/1859 "2017-09-12T03:11:55Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![JulianK](https://avatars.discourse-cdn.com/v4/letter/j/779978/32.png) [@JulianK](https://discourse.mc-stan.org/u/JulianK)\
**Post date:** [September 12, 2017, 3:11am UTC](https://discourse.mc-stan.org/t/advi-rats-example-adagrad/1859/1 "2017-09-12T03:11:55Z")

</div>

Curiosity drove me back to look at the in the context of ADVI that I posted [on the old Google groups discussion list](https://groups.google.com/forum/#!topic/stan-users/FaBvi8w7pc4).

I can get reasonable results from ADVI for RATS if I start with `init=0` and `eta=10` but not `eta=1`.

The higher the value of `eta` I use, the faster the convergence I get - if `vb()` runs (otherwise I get a message about the problem being too ill-conditioned).

This drove me back to look at the [Kucukelbir et al ADVI paper](https://arxiv.org/pdf/1506.03431.pdf).

**Limited memory Adagrad / length of memory**

In section E, “setting a step size for ADVI”, I note that the per-parameter step size for ADVI is derived based on a sliding window containing the sum of the squared gradients for the last 10 iterations.

How was this figure of “10” chosen? Some of the poor ADVI convergence seen in different models could be due to this quantity being too noisy due to the finite history. Did you ever consider trying a longer window?

Just something to consider…

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [September 17, 2017, 4:20pm UTC](https://discourse.mc-stan.org/t/advi-rats-example-adagrad/1859/2 "2017-09-17T16:20:21Z")

</div>

> [@JulianK](#):
>
> Did you ever consider trying a longer window?

Thanks for the input. I really don’t know the answer. Is the 10 a default that’s configurable? Is a different number more robust but maybe slower? If so, we’d go for more robustness in general.

Alp and Dustin aren’t really working on stan any more, so I don’t know what their original motiviations were. Andrew Gelman and Yuling Yao are looking at improving ADVI, but I don’t know where they’re at with it. I think their preliminary results were that scale mattered a lot.
