# Using cmdstanr without returning the output to R?

**URL:** <https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700>\
**Category:** Interfaces\
**Tags:** cmdstanr\
**Created:** [December 3, 2025, 6:56am UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700 "2025-12-03T06:56:19Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![helske](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/helske/32/3376_2.png) [@helske](https://discourse.mc-stan.org/u/helske)\
**Post date:** [December 3, 2025, 6:56am UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/1 "2025-12-03T06:56:19Z")

</div>

I have a model which results in csv files with sizes of tens of GB, This is fine, expect creating the resulting object in R is slow and sometimes impossible due to the size. However, I don’t need to process all the variables at once, so I realized that I can use the following to read in only subset of samples:

```no-highlight
output <-read_cmdstan_csv(files = "samples.csv", variables = c("lp__", "x", "y"))
fit <- cmdstanr:::CmdStanMCMC_CSV$new(output, "samples.csv", FALSE)

```

(Not sure if there is cleaner way which doesn’t rely on the unexported functionality)

My question is that that when I run `model$sample(...)`, is there a way to tell `sample` that it should not even try to create and return the results to R (which will cause error due to lack of memory)?

---

<div class="post-metadata">

**Author:** ![jsocolar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jsocolar/32/2486_2.png) [@jsocolar](https://discourse.mc-stan.org/u/jsocolar)\
**Post date:** [December 3, 2025, 2:22pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/2 "2025-12-03T14:22:02Z")

</div>

I don’t think this functionality is anywhere exposed by `cmdstanr`, but I agree that you have a good use case for it and perhaps it would be a nice feature. If you want to use the package internals to do this, check out `run_cmdstan` here [cmdstanr/R/run.R at master · stan-dev/cmdstanr · GitHub](https://github.com/stan-dev/cmdstanr/blob/master/R/run.R)

cc @jonah

EDIT: This is wrong. See below.

---

<div class="post-metadata">

**Author:** ![huffyhenry](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/huffyhenry/32/781_2.png) [@huffyhenry](https://discourse.mc-stan.org/u/huffyhenry)\
**Post date:** [December 3, 2025, 4:19pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/3 "2025-12-03T16:19:53Z")

</div>

I don’t think that `$sample` reads in all the samples by default. You have to call `$draws` for that. Perhaps it is the calculation of diagnostics that runs out of RAM? If so then you can set `diagnostics = FALSE` if you know what you’re doing.

If you have auxiliary variables for which you don’t need to save draws at all, define them in the `model` block or in a local block (an extra pair of `{` `}`) in `transformed parameters`.

Working with such large models with `R` and `cmdstanr` can be frustrating. Perhaps you will find my package [Stanislaw](https://github.com/huffyhenry/Stanislaw) useful. It extracts subsets of draws directly from CmdStan CSVs and can also calculate posterior summaries much faster than `$summary`.

---

<div class="post-metadata">

**Author:** ![jonah](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jonah/32/200_2.png) [@jonah](https://discourse.mc-stan.org/u/jonah)\
**Post date:** [December 3, 2025, 4:55pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/4 "2025-12-03T16:55:21Z")

</div>

> [@huffyhenry](#):
>
> I don’t think that `$sample` reads in all the samples by default. You have to call `$draws` for that. Perhaps it is the calculation of diagnostics that runs out of RAM? If so then you can set `diagnostics = FALSE` if you know what you’re doing.

This is right, it shouldn’t read in all the draws until you ask it to do something that requires them (e.g. `$draws()`, `$summary()`, printing, etc.). For turning off reading in the diagnostics I would use `diagnostics=""` or `diagnostics=NULL`, although `FALSE` might also work, I haven’t tested it.

---

<div class="post-metadata">

**Author:** ![helske](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/helske/32/3376_2.png) [@helske](https://discourse.mc-stan.org/u/helske)\
**Post date:** [December 3, 2025, 7:26pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/5 "2025-12-03T19:26:23Z")

</div>

Thanks, indeed I was mixing my experiences with `rstan` fit object; the out of memory issue actually happens later in the batch jobs when using `fit$save_object().`

I prefer not to define these variables inside local `{}` or inside model block as I need them also later in generated quantities, although I could of course just recompute them as it probably doesn’t matter much in terms of the overall computing time.

---

<div class="post-metadata">

**Author:** ![jonah](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jonah/32/200_2.png) [@jonah](https://discourse.mc-stan.org/u/jonah)\
**Post date:** [December 3, 2025, 9:13pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/6 "2025-12-03T21:13:13Z")

</div>

> [@helske](#):
>
> when using `fit$save_object().`

Ah, yeah that makes sense. `save_object()` will read everything into memory in order for it to be available when the object is loaded.

---

<div class="post-metadata">

**Author:** ![huffyhenry](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/huffyhenry/32/781_2.png) [@huffyhenry](https://discourse.mc-stan.org/u/huffyhenry)\
**Post date:** [December 4, 2025, 9:08am UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/7 "2025-12-04T09:08:14Z")

</div>

> [@helske](#):
>
> I prefer not to define these variables inside local `{}` or inside model block as I need them also later in generated quantities, although I could of course just recompute them as it probably doesn’t matter much in terms of the overall computing time.

I have this dilemma often. Yes it does not matter much in terms of time, but such duplicated code means bugs and it can rarely be refactored into a function. I’d love it if the language had a decorator that you could apply to a variable to exclude its draws from the CSVs.

---

<div class="post-metadata">

**Author:** ![jonah](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jonah/32/200_2.png) [@jonah](https://discourse.mc-stan.org/u/jonah)\
**Post date:** [December 4, 2025, 3:53pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/8 "2025-12-04T15:53:25Z")

</div>

> [@huffyhenry](#):
>
> I’d love it if the language had a decorator that you could apply to a variable to exclude its draws from the CSVs.

If I remember correctly, I think this is something @WardBrian had also expressed interest in, so maybe we can make that happen at some point.

---

<div class="post-metadata">

**Author:** ![WardBrian](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/wardbrian/32/12078_2.png) [@WardBrian](https://discourse.mc-stan.org/u/WardBrian)\
**Post date:** [December 4, 2025, 3:58pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/9 "2025-12-04T15:58:43Z")

</div>

Yes, and I still am interested in that feature (the most recent idea was you would annotate the variable with `@silent`). If I remember correctly the primary concerns were that it interacts badly with things like standalone generated quantities - a model that had a silenced variable can’t have its results loaded back in for further processing by the same model. But that also seems like an obvious and “fair” tradeoff.

---

<div class="post-metadata">

**Author:** ![jonah](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jonah/32/200_2.png) [@jonah](https://discourse.mc-stan.org/u/jonah)\
**Post date:** [December 4, 2025, 5:08pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/10 "2025-12-04T17:08:00Z")

</div>

> [@WardBrian](#):
>
> But that also seems like an obvious and “fair” tradeoff.

I agree that’s a tradeoff worth accepting.

---

<div class="post-metadata">

**Author:** ![helske](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/helske/32/3376_2.png) [@helske](https://discourse.mc-stan.org/u/helske)\
**Post date:** [December 4, 2025, 5:54pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/11 "2025-12-04T17:54:03Z")

</div>

It would be great to have something like `@silent` in Stan code, I too find repeated the code in multiple places in order to avoid saving some auxiliary stuff annoying and “dangerous”.

However, disregarding the issue of extra variables in the output CSVs, I think there’s a simple solution for avoiding reading everything to R: Just include argument `variables` for `as_cmdstan_fit()` and pass it to `read_cmdstan_csv()` which already accepts `variables` argument?

Current definition of `as_cmdstan_fit`:

```no-highlight
as_cmdstan_fit <- function(files, check_diagnostics = TRUE, format = getOption("cmdstanr_draws_format")) {
  csv_contents <- read_cmdstan_csv(files, format = format)
  switch(
    csv_contents$metadata$method,
    "sample" = CmdStanMCMC_CSV$new(csv_contents, files, check_diagnostics),
    "optimize" = CmdStanMLE_CSV$new(csv_contents, files),
    "variational" = CmdStanVB_CSV$new(csv_contents, files),
    "pathfinder" = CmdStanPathfinder_CSV$new(csv_contents, files),
    "laplace" = CmdStanLaplace_CSV$new(csv_contents, files)
  )
}

```

edit: Except it seems that `model$sample()` does not call `as_cmdstan_fit()`. Well, would at least help when manually creating the fit object for CSVs.

---

<div class="post-metadata">

**Author:** ![jonah](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jonah/32/200_2.png) [@jonah](https://discourse.mc-stan.org/u/jonah)\
**Post date:** [December 5, 2025, 4:38pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/12 "2025-12-05T16:38:05Z")

</div>

I made a branch of the cmdstanr repo called `"as_cmdstan_fit-with-variables"` that adds the `variables` argument to `as_cmdstan_fit()`:

> <https://github.com/stan-dev/cmdstanr/pull/1121/>
>
> \#### Submission Checklist
> 
> \- \[x\] Run unit tests
> \- \[x\] Declare copyright holde…r and agree to license (see below)
> 
> \#### Summary
> 
> Adds \`variables\` argument to \`as\_cmdstan\_fit()\` to allow creating objects from a subset of variables in the CSV files. 
> 
> \#### Copyright and Licensing
> 
> Please list the copyright holder for the work you are submitting
> (this will be you or your assignee, such as a university or company):
> \*\*Columbia University\*\*
> 
> 
> By submitting this pull request, the copyright holder is agreeing to
> license the submitted work under the following licenses:
> 
> \- Code: BSD 3-clause (https://opensource.org/licenses/BSD-3-Clause)
> \- Documentation: CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/)

It would be great if you (or anyone else reading this) could test it out and let me know if it works for your use case.

---

<div class="post-metadata">

**Author:** ![robertgrant](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/robertgrant/32/21233_2.png) [@robertgrant](https://discourse.mc-stan.org/u/robertgrant)\
**Post date:** [December 12, 2025, 10:40am UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/13 "2025-12-12T10:40:09Z")

</div>

I’m going to try this out as soon as I get a chance. Busy pre-Christmas though. Do we have alternative outputs to CSV? It hasn’t been a limit for me in the past but it is such a storage-inefficient format.

---

<div class="post-metadata">

**Author:** ![jonah](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jonah/32/200_2.png) [@jonah](https://discourse.mc-stan.org/u/jonah)\
**Post date:** [December 12, 2025, 5:32pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/14 "2025-12-12T17:32:03Z")

</div>

Unfortunately we’re still only using CSV. There have been various proposals for other formats (which would need to be changed in CmdStan itself), but as far as I know we haven’t had a developer take on that project yet.

---

<div class="post-metadata">

**Author:** ![jonah](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jonah/32/200_2.png) [@jonah](https://discourse.mc-stan.org/u/jonah)\
**Post date:** [January 13, 2026, 9:09pm UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/15 "2026-01-13T21:09:27Z")

</div>

> [@jonah](#):
>
> I made a branch of the cmdstanr repo called `"as_cmdstan_fit-with-variables"` that adds the `variables` argument to `as_cmdstan_fit()`:

This has now been merged into master

---

<div class="post-metadata">

**Author:** ![helske](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/helske/32/3376_2.png) [@helske](https://discourse.mc-stan.org/u/helske)\
**Post date:** [January 14, 2026, 7:27am UTC](https://discourse.mc-stan.org/t/using-cmdstanr-without-returning-the-output-to-r/40700/16 "2026-01-14T07:27:38Z")

</div>

Thanks, this is great, I was planning to test this out earlier but got distracted by other things.
