# Stan on GPU: looking for model+dataset examples for empirical evaluation of speedups

**URL:** <https://discourse.mc-stan.org/t/stan-on-gpu-looking-for-model-dataset-examples-for-empirical-evaluation-of-speedups/3002>\
**Category:** General\
**Created:** [January 16, 2018, 9:16am UTC](https://discourse.mc-stan.org/t/stan-on-gpu-looking-for-model-dataset-examples-for-empirical-evaluation-of-speedups/3002 "2018-01-16T09:16:11Z")\
**Posts on this page:** 1\
**Showing post:** 17

<div class="post-metadata">

**Author:** ![wds15](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/wds15/32/908_2.png) [@wds15](https://discourse.mc-stan.org/u/wds15)\
**Post date:** [January 25, 2018, 9:12pm UTC](https://discourse.mc-stan.org/t/stan-on-gpu-looking-for-model-dataset-examples-for-empirical-evaluation-of-speedups/3002/17 "2018-01-25T21:12:16Z")

</div>

MPI is coming, so the question is not silly at all. See here:

> [@MPI Design Discussion](http://discourse.mc-stan.org/t/mpi-design-discussion/1103/267):
>
> Ok, I have shown the good performance already a few times of MPI, but to this end mostly with synthetic examples. I just compiled a realistic example Hierarchical ODE based model 1300 subjects real world data set The running time on a single core takes 2.6 days to finish. I have setup things such that the 10 & 15 core run were on a single machine while the 20, 40 & 80 core run were distributed in blocks of 10 (so 2, 4 and 8 machines) onto the cluster which is networked using infiniband. The k…

However, the first implementation can speed up things dramatically, but it is not truly user friendly since you have to cram your data into the format we expect it for easy distribution via MPI.

---

_[View the full topic](https://discourse.mc-stan.org/t/stan-on-gpu-looking-for-model-dataset-examples-for-empirical-evaluation-of-speedups/3002)._
