# Proposal for consolidated output

**URL:** <https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263>\
**Category:** Developers\
**Tags:** features\
**Created:** [May 21, 2018, 12:41pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263 "2018-05-21T12:41:52Z")\
**Posts on this page:** 20\
**Page:** 2

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [June 4, 2018, 3:47pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/21 "2018-06-04T15:47:01Z")

</div>

Thanks for the feedback, good point about the GQ’s decoupling the number of parametres the algorithm deals with from the number of parameters in output. I’d like to see Bob/Daniel comment when they get to it but then I’ll move this to the design wiki with some clean-up.

I think the next step is for me to see if there are complications with consolidating output that would make it hard to (temporarily) change the innards of the services methods and algorithms w.r.t. output while still letting the interfaces produce the same output. This was a challenge with the original refactor that created the `services` method.

---

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [June 4, 2018, 3:47pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/22 "2018-06-04T15:47:27Z")

</div>

I checked it out, this is great!

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [June 5, 2018, 12:40am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/23 "2018-06-05T00:40:46Z")

</div>

> [@martinmodrak](#):
>
> **transformed data** : Similar to gen quants in that those have arbitrary dimension and may include ints. However those are output only once per-run. This is currently not piped through but should be useful for debugging

That would be useful. Especially with RNGs.

> [@martinmodrak](#):
>
> might be excusable to convert them to doubles ([all 32bit integers can be exactly represented as double](https://stackoverflow.com/questions/3793838/which-is-the-first-integer-that-an-ieee-754-float-is-incapable-of-representing-e)).

The plan is to upgrade to 64 bit ints as soon as possible.

Even then using actual int may be too difficult with this kind of writer structure. It is easier with just string output.

We also want to look downstream to serialization. We will need good int serialization for things like rng seeds.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [June 5, 2018, 12:57am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/24 "2018-06-05T00:57:32Z")

</div>

Thanks.

> [@sakrejda](#):
>
> `struct` holding the config.

not sure how we do this. each service gets its own flat config struct or is rhere looser typing?

> [@sakrejda](#):
>
> Model-file defined:

do the param names come with dimensions or types?

> [@sakrejda](#):
>
> **Model-file defined for sampling** :

There are different numbers of constrained and unconstrained params for simplex, cov matrix, etc.

> [@sakrejda](#):
>
> **Model-file defined for optimization** :

the hessian is an R thing, not in Stan. and is that constained or not?

> [@sakrejda](#):
>
> **Model-file defined for ADVI** :

the cov is on unconstrained scale. and it is dense or diagonal.

> [@sakrejda](#):
>
> **Model-file defined, for HMC** :

hmc only has one mass matrix after adaptation other than rhmc. it may be dense or diagonal.

> [@sakrejda](#):
>
> Model-file defined, for whatever algorithm:

gradients always unconstrained

> [@sakrejda](#):
>
> **Algorithm-defined for HMC** :

log density unnormalized on unconstrained scale with jacobians.

divergence is boolean, despite how it looks and is documented now.

> [@sakrejda](#):
>
> **Algorithm-defined for BFGS/LBFGS** :

The notes should definitely be broken out.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [June 5, 2018, 1:09am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/25 "2018-06-05T01:09:40Z")

</div>

> [@sakrejda](#):
>
> - With the three algorithms we have we could honestly just write some classes, one for each of sampling, optimization, and HMC. The specific sub-algorithms (e.g.-HMC with diagonal mass matrix and adaptation) would then just call their specific parts (e.g., `relay.send_mass_matrix(std::vector<double> x)` ).
> - As long as we build these classes up in a modular way we should be able to avoid boilerplate.

Thanks for the survey. It’s super helpful.

I am still not sure what is being proposed to dispatch events sent to the relays to handlers defined by the interfaces. Is the relay generic across interfaces? Do interfaces register handlers with the relay?

I was trying to understand the templating proposal as a way to streamline writing all these.

> [@sakrejda](#):
>
> The dev-side problem @martinmodrak ran into could be solved by deciding that the `logger` class could have more over-rides with a default conversion to text. If an interface didn’t implement a specific type, it could always get the output as text.

For text defaults, that can just be an implementation of a handler that an interface could plug in. I think the default implementations should be no ops to stop stray text getting through anywhere.

---

<div class="post-metadata">

**Author:** ![syclik](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/syclik/32/6_2.png) [@syclik](https://discourse.mc-stan.org/u/syclik)\
**Post date:** [June 7, 2018, 10:12am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/26 "2018-06-07T10:12:20Z")

</div>

@sakrejda, thanks for taking the huge effort of writing that out. I’ve tried multiple times to write it that clearly, but haven’t.

I think everything you’ve said about going from this current implementation to something workable is great.

I still don’t see the benefit of the templating. It’s not templating, per se. It’s the indirection in how the templates are used (structs living somewhere else).

I’d like to mention a future complication: in some cases, we won’t know how many iterations we will have. With that in mind, I’ve started thinking that this is an important distinction: things that are written once per iteration, more frequently than that, less frequently, and ad hoc. (@sakrejda, I think you laid all of that out nicely.)

---

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [June 7, 2018, 11:46am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/27 "2018-06-07T11:46:33Z")

</div>

Having written this out, I think we can do something that lowers the bar for understanding the relay code… there may still be room for templating to avoid code duplication but I think the confusing application is in the relay and given we have three kinds of algorithms it’s avoidable. I’ll try to write up a specific suggestion this weekend.

---

<div class="post-metadata">

**Author:** ![syclik](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/syclik/32/6_2.png) [@syclik](https://discourse.mc-stan.org/u/syclik)\
**Post date:** [June 7, 2018, 12:11pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/28 "2018-06-07T12:11:04Z")

</div>

3 is a manageable number. Brute force isn’t always so bad.

---

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [June 7, 2018, 12:51pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/29 "2018-06-07T12:51:27Z")

</div>

True, it’s the level below the relay where templating might make sense to take care of sending heterogeneous tables but those would bee helpers for the interface code rather than an all encompassing relay class.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [June 7, 2018, 11:32pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/30 "2018-06-07T23:32:24Z")

</div>

Thanks! I’m really looking forward to seeing this. I feel better having the initial survey in hand already.

I’d like to see simple APIs for the RStan, PyStan, and CmdStan interface clients if possible, even if they’re a bit redundant. The API for the algorithms and the back-end code in stan-dev/stan is much less of a concern for me because it’s not being used by external clients.

---

<div class="post-metadata">

**Author:** ![martinmodrak](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/martinmodrak/32/133_2.png) [@martinmodrak](https://discourse.mc-stan.org/u/martinmodrak)\
**Post date:** [June 12, 2018, 5:23am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/31 "2018-06-12T05:23:51Z")

</div>

Just a ping that I am unsure what the current status is - am I supposed to try to write some more code proposal or does someone else have the ball now?

---

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [June 12, 2018, 10:29am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/32 "2018-06-12T10:29:51Z")

</div>

Hey, sorry about that, I should’ve added more here. Next step is I need to write up what the calls will look like from the algorithm side and the interface side. There’s no point to writing more code until that’s clearly laid out. I was going to do it last weekend but I needed a faster .csv reader for Stan so I did that instead, I’ll be able to do it this week though.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [June 12, 2018, 9:51pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/33 "2018-06-12T21:51:32Z")

</div>

Thanks @sakrejda, that’s my understanding of where we’re at, too.

---

<div class="post-metadata">

**Author:** ![syclik](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/syclik/32/6_2.png) [@syclik](https://discourse.mc-stan.org/u/syclik)\
**Post date:** [June 20, 2018, 3:34am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/34 "2018-06-20T03:34:36Z")

</div>

Thanks! I’m looking at this again. @sakrejda, seriously: thanks! And @martinmodrak for digging too.

I’m starting to think we could have writing done at each iteration. By that, I mean once at the end of each iteration. I can’t think of any algorithm that doesn’t have iterations (and even if it did, we could say it’s 1 iteration).

For each algorithm, I think we can write these things at the end:

- parameter values
- generated quantities
- algorithm output (i.e. `treedepth__`)
- unconstrained values / latent algorithm parameters
- random seed

If we have the random seed(s), we can always generate the next iteration, _if_ there is no other action between the end of the iteration and the start of the next iteration.

I think we also need to report the same information right before the start of the first iteration. If we had that, we would be able to reconstruct each iteration exactly. Am I correct in thinking that’s one useful thing about the output? If we had that, we could generate the information within an iteration (leapfrog steps).

Anyway, looking forward to @sakrejda’s next suggestions.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [June 20, 2018, 8:34pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/35 "2018-06-20T20:34:41Z")

</div>

The other thing we were looking at is trajectory within iteration for HMC. So that’s a big block of information.

The RNG we use now isn’t restartable. The seeds are 8 bytes, but the states are something like 32 bytes. We could probably reseed from state if there’s a constructor.

> [@syclik](#):
>
> I’m starting to think we could have writing done at each iteration.

Can we generalize from the algorithms notions of iterations to an output process that somehow comes with a header and streaming. It’s not so much that it’s an iteration, but that it provides a set of parameter values. Could we use that same notion internally with a sequence of leapfrog steps?

---

<div class="post-metadata">

**Author:** ![syclik](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/syclik/32/6_2.png) [@syclik](https://discourse.mc-stan.org/u/syclik)\
**Post date:** [June 21, 2018, 12:25am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/36 "2018-06-21T00:25:01Z")

</div>

> [@Bob\_Carpenter](#):
>
> Can we generalize from the algorithms notions of iterations to an output process that somehow comes with a header and streaming. It’s not so much that it’s an iteration, but that it provides a set of parameter values. Could we use that same notion internally with a sequence of leapfrog steps?

Absolutely!

I was thinking there are iterations and things done within iterations, but I’m sure there’s a better abstraction if we think about it.

---

<div class="post-metadata">

**Author:** ![syclik](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/syclik/32/6_2.png) [@syclik](https://discourse.mc-stan.org/u/syclik)\
**Post date:** [June 28, 2018, 1:58pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/37 "2018-06-28T13:58:57Z")

</div>

> [@Bob\_Carpenter](#):
>
> The RNG we use now isn’t restartable. The seeds are 8 bytes, but the states are something like 32 bytes. We could probably reseed from state if there’s a constructor.

There’s a constructor that takes two seeds! It’s also easy to get the state out as two numbers.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [July 2, 2018, 5:36am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/38 "2018-07-02T05:36:01Z")

</div>

> [@syclik](#):
>
> There’s a constructor that takes two seeds! It’s also easy to get the state out as two numbers.

That’s good news. I never realized they had those. The PRNG we’re using is here:

[https://www.boost.org/doc/libs/1\_66\_0/doc/html/boost/random/ecuyer1988.html](https://www.boost.org/doc/libs/1_66_0/doc/html/boost/random/ecuyer1988.html)

I see the two argument constructor, but I don’t see how to get a two-componet seed out of the state. Do you have an example somewhere?

---

<div class="post-metadata">

**Author:** ![syclik](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/syclik/32/6_2.png) [@syclik](https://discourse.mc-stan.org/u/syclik)\
**Post date:** [July 3, 2018, 1:29am UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/39 "2018-07-03T01:29:56Z")

</div>

I got curious, so I just `std::cout << rng << std::endl` at some point and it put out two numbers. I used that as the seed and got back the exact same results. I don’t know if there’s a more direct way to get the two numbers out, but I’m sure we can manage even if it’s not documented.

---

<div class="post-metadata">

**Author:** ![sakrejda](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sakrejda/32/845_2.png) [@sakrejda](https://discourse.mc-stan.org/u/sakrejda)\
**Post date:** [July 8, 2018, 9:45pm UTC](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263/40 "2018-07-08T21:45:26Z")

</div>

Btw, I put together issues for this on github and went on vacation, another week before I pick it up.

[Previous page](https://discourse.mc-stan.org/t/proposal-for-consolidated-output/4263.md?page=1)
