# Including data in covariance matrix

**URL:** <https://discourse.mc-stan.org/t/including-data-in-covariance-matrix/35283>\
**Category:** Modeling\
**Tags:** specification\
**Created:** [May 28, 2024, 11:00am UTC](https://discourse.mc-stan.org/t/including-data-in-covariance-matrix/35283 "2024-05-28T11:00:46Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![levis](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/levis/32/8315_2.png) [@levis](https://discourse.mc-stan.org/u/levis)\
**Post date:** [May 28, 2024, 11:00am UTC](https://discourse.mc-stan.org/t/including-data-in-covariance-matrix/35283/1 "2024-05-28T11:00:46Z")

</div>

one of the objectives of my STAN model is to see how age covaries with the parameters of the model. However, in simulations, the model recovers a covariance matrix that is much lower than the actual covariance between parameters. Does anyone have any insight regarding why that could be with an approach like this?

TIA.

```no-highlight
data {
  int<lower=1> nS; // number of subjects
  vector[nS] age_zscored; // Z-scored age
}

parameters {
  vector[2] theta_mu; 
  matrix[2, nS] theta_pr_minus2;
  vector<lower=0>[2] sigma; // parameter variance  
  cholesky_factor_corr[3] L; // include age 
}

transformed parameters {
  matrix[2, nS] theta_pr;
  vector<lower=0>[3] sigma_pr;
  matrix[nS, 3] theta;
  //parameters that may covary with age 
  vector<lower=0,upper=1>[nS] alpha;
  vector[nS] beta_1;

  // Assign fixed standard deviation for age (first element)
  sigma_pr[1] = 1;
  sigma_pr[2:3] = sigma;

  // Assign age as the first row of theta_pr
  theta_pr[1,] = to_row_vector(age_zscored);
  theta_pr[2:3,] = theta_pr_minus2;

  // Transform theta_pr to get theta
  theta = transpose(diag_pre_multiply(sigma_pr, L) * theta_pr);
  beta_1 = (theta_mu[1] + theta[,2]) * 5;
  alpha = Phi_approx(theta_mu[2] + theta[,3]);
  
}

model {
  // Priors
  target += normal_lpdf(theta_mu | 0, 2);
  target += std_normal_lpdf(to_vector(theta_pr_minus2));
  target += student_t_lpdf(sigma | 3, 0, 1);
  target += lkj_corr_cholesky_lpdf(L | 1);
  
  // model specification goes here 
}

generated quantities {
  matrix[3,3] Omega = multiply_lower_tri_self_transpose(L);
} 

```

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [May 28, 2024, 7:12pm UTC](https://discourse.mc-stan.org/t/including-data-in-covariance-matrix/35283/2 "2024-05-28T19:12:47Z")

</div>

> [@levis](#):
>
> a covariance matrix that is much lower

Do you mean a covariance matrix with smaller absolute correlations than you used to simulate? Are you sure you’re simulating from the model’s priors and data-generating process? The reason I ask is that if you simulate data from a Bayesian model, then fit it using the same model, it should be well calibrated if the sampler is functioning correctly. So my question is really whether you’ve tried something like simulation-based calibration and seen it fail?

P.S. This statement is redundant because the 1 parameter makes the distribution over L uniform and thus a no-op.

```stan
  target += lkj_corr_cholesky_lpdf(L | 1);

```

P.P.S. “Stan” is not an acronym—it’s named after Stanislaw Ulam.
