# MPI Design Discussion

**URL:** <https://discourse.mc-stan.org/t/mpi-design-discussion/1103>\
**Category:** Developers\
**Created:** [June 30, 2017, 6:50am UTC](https://discourse.mc-stan.org/t/mpi-design-discussion/1103 "2017-06-30T06:50:05Z")\
**Posts on this page:** 1\
**Showing post:** 267

<div class="post-metadata">

**Author:** ![wds15](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/wds15/32/908_2.png) [@wds15](https://discourse.mc-stan.org/u/wds15)\
**Post date:** [January 23, 2018, 9:25am UTC](https://discourse.mc-stan.org/t/mpi-design-discussion/1103/267 "2018-01-23T09:25:20Z")

</div>

Ok, I have shown the good performance already a few times of MPI, but to this end mostly with synthetic examples. I just compiled a realistic example

- Hierarchical ODE based model
- 1300 subjects
- real world data set

The running time on a single core takes 2.6 days to finish. I have setup things such that the 10 & 15 core run were on a single machine while the 20, 40 & 80 core run were distributed in blocks of 10 (so 2, 4 and 8 machines) onto the cluster which is networked using infiniband. The key question to me was how well the performance scales and the result is stunning. The 64h on a single core go down to about 1h which is a 62x speedup, but look for yourself.

… and all results match exactly.

 ![mpi_wallclock](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/3/32c7c97867260e03b25ec673ef98f14a1a52f9be.png)

 ![mpi_speedup](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/c/c42fca8c34e0473015659c7c411e1f9326a6d8f7.png)

---

_[View the full topic](https://discourse.mc-stan.org/t/mpi-design-discussion/1103)._
