# Offset multiplier initialization

**URL:** <https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712>\
**Category:** General\
**Created:** [February 9, 2021, 7:38pm UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712 "2021-02-09T19:38:02Z")\
**Posts on this page:** 9\
**Page:** 3

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [April 24, 2025, 8:26pm UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/43 "2025-04-24T20:26:27Z")

</div>

> [@martinmodrak](#):
>
> you are adding work

Indeed, we are. But it’s a trivial amount of extra work in the scheme of a larger model, so I don’t think that should discourage use. If we had thought that was a dealbreaker, we wouldn’t have included it in the first place.

But the numerical behavior is really a problem, and that’s enough to make me not want to recommend using the offset/multiplier in its current form. Oh well.

---

<div class="post-metadata">

**Author:** ![spinkney](https://avatars.discourse-cdn.com/v4/letter/s/dec6dc/32.png) [@spinkney](https://discourse.mc-stan.org/u/spinkney)\
**Post date:** [April 25, 2025, 12:40am UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/44 "2025-04-25T00:40:41Z")

</div>

I always thought the offset/multiplier was backwards, like why can’t it be the same offset and multiplier in the parameters block but just declare the y ~ std\_normal()? I only have to write the mu and sigma once and this would fix the auto diff precision issue.

Like doesn’t offset and multiply just do what I want to this standard normal already?

---

<div class="post-metadata">

**Author:** ![avehtari](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/avehtari/32/5935_2.png) [@avehtari](https://discourse.mc-stan.org/u/avehtari)\
**Post date:** [April 25, 2025, 7:10am UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/45 "2025-04-25T07:10:49Z")

</div>

> [@Dalton](#):
>
> The general tendency of the NCP to dive towards very small values during warmup, particularly with high-dimensional random effect vectors.

This is over-generalization from one example. This depends on the initial values and where most of the posterior mass is. For illustration, the same example, but now init near 0 and at the same time the most of the posterior mass is far from 0, so we are not starting near the mode.

```no-highlight
set.seed(3)
n_grp <- 1000
n_per_grp <- 5
mu <- 10
sigma_re <- 2
sigma_residual <- 4

re <- rnorm(n_grp, mu, sigma_re)
y <- rnorm(n_grp * n_per_grp, re, sigma_residual)
grp_id <- rep(1:n_grp, n_per_grp)

re_data <- list(
  N = n_grp * n_per_grp,
  n_grp = n_grp,
  y = y, 
  grp_id = grp_id
)

fit <- re_mod$sample(data = re_data, chains = 60, parallel_chains = 8, 
                     save_warmup = T, 
                     iter_warmup = 1000, iter_sampling = 1, refresh=0,
                     init=.0001)

samps <- fit$draws(variables = c("mu","sigma","sigma_residual"),
                   inc_warmup = TRUE)
bayesplot::mcmc_trace(samps[1:50,,])

```

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/3X/d/2/d253e41dee0f08afdc5fae9a5e706b2217240cf8.jpeg)

No diving to problematically small sigma

---

<div class="post-metadata">

**Author:** ![avehtari](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/avehtari/32/5935_2.png) [@avehtari](https://discourse.mc-stan.org/u/avehtari)\
**Post date:** [April 25, 2025, 1:23pm UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/46 "2025-04-25T13:23:56Z")

</div>

Both @jsocolar’s and mine warmup traceplots show constant progress in the beginning. I also realised that the traceplots did not show the initial values. The following plots include the initial values and NUTS diagnostics. The first iteration is fine, but then there are several iterations which all have big step size, divergence, minimal treedepth and no progress. I repeated this with init=2 (bad inits), init=.1 (good init), and initializing with posterior draws (perfect init), and in all cases the sampler is stuck for several iterations which seems like a failure in the adaptation.

init=2

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/3X/f/0/f03dc2ba7137480049ebf056bfb39d15a02c9027.jpeg)

init=0.1

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/3X/d/8/d8abf115c8e43a5d7eacbc7ae1467f1c5f0f2ab8.jpeg)

init using the posterior draws

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/3X/c/f/cf65818b313d97bc652e0defeb636ca7b55487a9.jpeg)

---

<div class="post-metadata">

**Author:** ![jsocolar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jsocolar/32/2486_2.png) [@jsocolar](https://discourse.mc-stan.org/u/jsocolar)\
**Post date:** [April 25, 2025, 3:26pm UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/47 "2025-04-25T15:26:27Z")

</div>

> [@avehtari](#):
>
> The first iteration is fine, but then there are several iterations which all have big step size, divergence, minimal treedepth and no progress. I repeated this with init=2 (bad inits), init=.1 (good init), and initializing with posterior draws (perfect init), and in all cases the sampler is stuck for several iterations which seems like a failure in the adaptation.

This is the expected outcome of the dual averaging, right? When the adaptation sees a very high acceptance stat very early in the adaptation, it’ll aggressively explore a large step size that takes a few iterations to come back down. The computational cost is minimal because they all diverge immediately.

---

<div class="post-metadata">

**Author:** ![avehtari](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/avehtari/32/5935_2.png) [@avehtari](https://discourse.mc-stan.org/u/avehtari)\
**Post date:** [April 25, 2025, 3:46pm UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/48 "2025-04-25T15:46:43Z")

</div>

> [@jsocolar](#):
>
> This is the expected outcome of the dual averaging, right? When the adaptation sees a very high acceptance stat very early in the adaptation, it’ll aggressively explore a large step size that takes a few iterations to come back down. The computational cost is minimal because they all diverge immediately.

Even if the cost is small, it seems still silly to increase the step size 1000 times bigger after results from one iteration, even when

- initializing with posterior draws
- the step size during the first iteration is in the middle of later step size distribution
- and the accept\_stat of the first iteration is between .25 and 1 and not much different from the later iterations, so it seems your high accpetance\_stat claim does not hold

But maybe this silly behavior doesn’t matter, if Bob’s new sampling algorithm is better with step sizes anyway.

---

<div class="post-metadata">

**Author:** ![Dalton](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/dalton/32/4287_2.png) [@Dalton](https://discourse.mc-stan.org/u/Dalton)\
**Post date:** [April 25, 2025, 5:57pm UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/49 "2025-04-25T17:57:56Z")

</div>

> This is over-generalization from one example.

I think it’s fair to say that it’s _tendency_ under the default initialization of Uniform(-2, 2), no? Certainly it is not guaranteed to happen and it depends on the data and the model, but this is common enough occurrence that users encountering this problem in the wild may find this thread in their searching. That’s why this discussion is valuable. My comment was just an attempt to summarize the conversation for the non-developers who might stumble upon this thread and want a tl;dr.

For me, the other takeaway from this thread is that I should also be paying more attention to my initial values when I use this kind of structure, but that’s is something applies to both on offset/multiplier and the transformation without change of variables approaches.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [April 25, 2025, 5:58pm UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/50 "2025-04-25T17:58:22Z")

</div>

I’m writing up the C++ version as quickly as I can. Then I plan to release it Nutpie style at first with an interface through BridgeStan. Integrating directly in Stan is a multiple person-month job that touches half a dozen core and interface libraries (stan, CmdStan, cmdstanpy, cmdstanr, rstan, pystan, stan.jl).

---

<div class="post-metadata">

**Author:** ![jsocolar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jsocolar/32/2486_2.png) [@jsocolar](https://discourse.mc-stan.org/u/jsocolar)\
**Post date:** [April 25, 2025, 7:17pm UTC](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712/51 "2025-04-25T19:17:33Z")

</div>

> [@avehtari](#):
>
> and the accept\_stat of the first iteration is between .25 and 1 and not much different from the later iterations, so it seems your high accpetance\_stat claim does not hold

Ooof. Is the step-size accidentally getting boosted at the beginning of the init buffer in the same way that it does at the after all the metric adaptation windows?

[Previous page](https://discourse.mc-stan.org/t/offset-multiplier-initialization/20712.md?page=2)
