# Multicore Speedups are different between models

**URL:** <https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219>\
**Category:** Algorithms\
**Created:** [July 12, 2017, 5:44pm UTC](https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219 "2017-07-12T17:44:56Z")\
**Posts on this page:** 6\
**Page:** 2

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [July 13, 2017, 6:51pm UTC](https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219/21 "2017-07-13T18:51:30Z")

</div>

That’s only a couple percent variation, which is normal for any kind of speed comparison unless you take extreme measures to make sure no background processes ever run.

---

<div class="post-metadata">

**Author:** ![Emma926](https://avatars.discourse-cdn.com/v4/letter/e/58956e/32.png) [@Emma926](https://discourse.mc-stan.org/u/Emma926)\
**Post date:** [July 13, 2017, 6:55pm UTC](https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219/22 "2017-07-13T18:55:37Z")

</div>

I start to think the memory bandwidth bound is the real explanation. The 12.6GB/s is with single core, thus 4 cores should have higher bandwidth. I haven’t measure it but it’s may be several times higher.

---

<div class="post-metadata">

**Author:** ![Emma926](https://avatars.discourse-cdn.com/v4/letter/e/58956e/32.png) [@Emma926](https://discourse.mc-stan.org/u/Emma926)\
**Post date:** [July 13, 2017, 6:56pm UTC](https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219/23 "2017-07-13T18:56:19Z")

</div>

That’s right. The variation looks normal. There shouldn’t be any random scheduling problem.

---

<div class="post-metadata">

**Author:** ![Andre\_Pfeuffer](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/andre_pfeuffer/32/4939_2.png) [@Andre\_Pfeuffer](https://discourse.mc-stan.org/u/Andre_Pfeuffer)\
**Post date:** [July 14, 2017, 10:42am UTC](https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219/24 "2017-07-14T10:42:31Z")

</div>

And what, if you start 4 processes of ad\_advertisement with cmdstan, does it differ from RStan?

---

<div class="post-metadata">

**Author:** ![jerlich](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jerlich/32/2131_2.png) [@jerlich](https://discourse.mc-stan.org/u/jerlich)\
**Post date:** [September 11, 2017, 4:23am UTC](https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219/25 "2017-09-11T04:23:11Z")

</div>

I have a 64-core server, so I ran a fit with 60 chains. One chain behaved very badly (was off by 2 orders of magnitude - which may be a hint that we should have been fitting a log-scaled parameter), and my rhat’s were pretty bad. However, the other 59 chains converged.

When I compared the model predictions to the data, they actually did a pretty good job.

So my question is: is there ever a case where it is ok to ignore one bad chain? Or is the best practice to fix the problem (e.g. run longer warmup / rescale parameters) and re-fit?

---

<div class="post-metadata">

**Author:** ![bgoodri](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bgoodri/32/4451_2.png) [@bgoodri](https://discourse.mc-stan.org/u/bgoodri)\
**Post date:** [September 11, 2017, 4:59am UTC](https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219/26 "2017-09-11T04:59:41Z")

</div>

> [@jerlich](#):
>
> So my question is: is there ever a case where it is ok to ignore one bad chain? Or is the best practice to fix the problem (e.g. run longer warmup / rescale parameters) and re-fit?

In general, fix the problem. The number of situations where the one “bad” chain indicates there is a problematic part of the parameter space that the “good” chains never encountered is bigger than the number of situations where the one bad chain can be safely ignored. I would think about the scaling and also consider specifying a smaller value of `init_r`.

[Previous page](https://discourse.mc-stan.org/t/multicore-speedups-are-different-between-models/1219.md?page=1)
