# Repeat Sales Regression ported from BUGS

**URL:** <https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267>\
**Category:** Modeling\
**Created:** [May 21, 2018, 2:41pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267 "2018-05-21T14:41:40Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![sar2160](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sar2160/32/2370_2.png) [@sar2160](https://discourse.mc-stan.org/u/sar2160)\
**Post date:** [May 21, 2018, 2:41pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/1 "2018-05-21T14:41:40Z")

</div>

I’m trying to port a hierarchical repeat sales index model from BUGS , the original code from the paper is attached (I think the author is @avdminne, ) . My Stan model compiles and fits, but I need to turn max\_treedepth up to at least 14 in order to avoid warnings - usually all or almost all iterations exceed the treedepth at default settings. This is fine but taking max\_treedepth from 10 to 14 lengthens the runtime 6x+ and I’m already only using a subset of the data. Is there anything I could try to avoid a higher treedepth, maybe reparameterization tricks or alternative priors?

I also read on the “Guide to Errors” about looking for parameters correlated with energy\_ , but I have a lot of parameters and since all iterations typically exceed the max\_treedepth I wonder whether I’m defining something in a not ideal way.

```stan
data {
 int<lower=1> N; // N observations
 int<lower=1> Na; // N areas
 int<lower=1> Nt; //N timeperiods
 vector[N] y; // log price diff
 int<lower=1> month1[N]; //month of repeat trxn
 int<lower=1> month0[N]; //month of first trxn
 int<lower=1> area_id[N]; // area id
 real<lower=0> adf; // parameter for DoF
}

parameters {
  real<lower=0> sigma_y;
  real beta[Nt];
  real inc[Nt];
  real kappa[Nt];
  real<lower=0> tau_alpha;
  real<lower=0> tau_kappa;
  real<lower=0> tau_eta;
  matrix[Na,Nt] alpha; 
  real<lower=1> nu; // degrees of freedom for t
  
}

transformed parameters {
  
  real mu_beta[Nt];
  real mu[N];
    
  mu_beta[1] = 0;
  for (t in 2:Nt){
        mu_beta[t] = beta[t-1]+inc[t]*kappa[t]; // 'matt trick?'
   }      
  for (i in 1:N){
       mu[i] = (beta[month1[i]] - beta[month0[i]]) + 
     (alpha[area_id[i], month1[i]] - alpha[area_id[i], month0[i]]);
   }
  
}

model {

  
  // Priors
  sigma_y ~ inv_gamma(0.1,0.1);
  tau_alpha ~ lognormal(5,1);
  tau_kappa ~ inv_gamma(0.001, 0.001);
  tau_eta ~ inv_gamma(0.001, 0.001);
  nu ~ exponential(adf);
  beta ~ normal(mu_beta, tau_eta);
  inc ~ normal(0,1);

   for (t in 2:Nt) {
    kappa[t] ~ normal(kappa[t-1], tau_kappa);
   }
   
   for (a in 1:Na){
     for (t in 2:Nt){
       alpha[a,t] ~ normal(alpha[a,t-1], tau_alpha);
     }
   }
   
y ~ student_t(nu, mu, sigma_y);
}

```

I’m omitting the other hierarchical features for property types that is in the BUGS model. Also I tried to incorporate the “Matt trick” as described in this other post [Structural Time Series research using Stan](http://discourse.mc-stan.org/t/structural-time-series-research-using-stan/1548).

EDIT: Based on reading discussions elsewhere and the guide to priors it seems like inverse gamma priors are necessary for BUGS but not ideal for Stan, I’ve changed them to standard student\_t with DoF between 2-4.

[BUGScode-2017-hiearchichal-repeat-sales.pdf](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/6/68019690228cf188ed7817eed8566aa4b0261d78.pdf) (56.1 KB)

---

<div class="post-metadata">

**Author:** ![aaronjg](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/aaronjg/32/3768_2.png) [@aaronjg](https://discourse.mc-stan.org/u/aaronjg)\
**Post date:** [May 21, 2018, 10:02pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/2 "2018-05-21T22:02:54Z")

</div>

> [@sar2160](#):
>
> I need to turn max\_treedepth up to at least 14 in order to avoid warnings - usually all or almost all iterations exceed the treedepth at default settings

This suggests there is some correlation in your parameters. Looking at the model, this is where I would start. try, for a few, `i` plotting pairs of:

- beta[i], inc[i], and kappa[i]
- mu\_beta, tau\_eta, and beta[i]
- kappa[i] and kappa[i-1]
- kappa[i] and tau\_kappa
- nu and sigma

Also, I’m not totally sure I follow what you are trying to do, but this paper may be of interest, which models repeat sales in Stan.

> **[11-07-2017-Dew-Ryan-PAPER-Ansari\_BNP\_CBA-JMP.pdf](https://marketing.wharton.upenn.edu/wp-content/uploads/2017/08/11-07-2017-Dew-Ryan-PAPER-Ansari_BNP_CBA-JMP.pdf)**
>
> 1089.65 KB

---

<div class="post-metadata">

**Author:** ![sar2160](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sar2160/32/2370_2.png) [@sar2160](https://discourse.mc-stan.org/u/sar2160)\
**Post date:** [May 22, 2018, 5:26pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/3 "2018-05-22T17:26:12Z")

</div>

Thanks, the model name is sort-of confusing and I could’ve included some more info. Repeat Sales models in real estate try to estimate price appreciation/depreciation using the log price change between sales for properties that have changed hands multiple times to try and and avoid having to control for hedonic factors of each property.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [May 22, 2018, 5:35pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/4 "2018-05-22T17:35:19Z")

</div>

There are lots of problems with this model, starting with priors that we don’t recommend on either statistical or computational grounds (we have a Wiki on prior recommendations and also the regression chapter in the manual).

The problem may be with the parameterization of `beta`, which I don’t quite follow. There’s this in the model block

```
beta ~ normal(mu_beta, tau_eta);

```

where `mu_beta` is a transformed parameter and `tau_eta` a parameter. You want to use the non-centered paramterization on `beta, and make that`beta\_raw ~ normal(0, 1);`and then define`beta = mu\_beta + beta\_raw \* tau\_eta;`. That naming scheme is confusing. We'd usually use`sigma\_beta`as the name of the parameter rather than`tau\_eta`.

Also, do you realize stan uses standard deviation, not variance?

Then you may also need a non-centered parameerization of the time series, but I’d start with `beta` and by tighteing the priors.

---

<div class="post-metadata">

**Author:** ![sar2160](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sar2160/32/2370_2.png) [@sar2160](https://discourse.mc-stan.org/u/sar2160)\
**Post date:** [May 30, 2018, 4:41pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/5 "2018-05-30T16:41:57Z")

</div>

I blindly copied the priors from the original BUGS model without realizing they are on inverse-variance, Ive changed them all based on the wiki guidelines. Sorry for the confusing `tau` names, I did it to mostly preserve the syntax of the BUGS model and make comparison a little easier. I can’t post the BUGS model in text because it’s trapped in a pdf, but I attached it to the original post.

Anyway, I did the non-centered parameterization and modified all the priors, see below. I’m still getting lots of (but less than before) iterations that exceed max treedepth. Considering this model presumably worked well in BUGS could there be something that makes it unsuitable for use in Stan?

Thanks,

```stan
data {
 int<lower=1> N; // N observations
 int<lower=1> Na; // N areas
 int<lower=1> Nt; //N timeperiods
 vector[N] y; // log price diff
 int<lower=1> month1[N]; //month of repeat trxn
 int<lower=1> month0[N]; //month of first trxn
 int<lower=1> area_id[N]; // area id
 real<lower=0> adf; // parameter for DoF
  
}

parameters {
  real<lower=0> sigma_y;
  real beta_raw[Nt];
  real inc[Nt];
  real kappa[Nt];
  real<lower=0> tau_alpha;
  real<lower=0> tau_kappa;
  real<lower=0> tau_eta;
  matrix[Na,Nt] alpha; 
  real<lower=1> nu; // degrees of freedom for t
  
}

transformed parameters {
  real beta[Nt];
  real mu_beta[Nt];
  real mu[N];
  

  beta[1] = 0;
  mu_beta[1] = 0;
  for (t in 2:Nt){
        mu_beta[t] = beta[t-1]+inc[t]*kappa[t]; // 'matt trick?'
        beta[t] = mu_beta[t] + beta_raw[t] * tau_eta; // non-centered reparameterization

   }      
  for (i in 1:N){
       mu[i] = (beta[month1[i]] - beta[month0[i]]) + 
     (alpha[area_id[i], month1[i]] - alpha[area_id[i], month0[i]]);
   }
  
}

model {
  // Priors
  sigma_y ~ student_t(2,0,10);
  tau_alpha ~ normal(0,10);
  tau_kappa ~ normal(0,10);
  tau_eta ~ normal(0,10);
  nu ~ exponential(adf);
  beta_raw ~ normal(0,1);
  inc ~ normal(0,1);

  for (t in 2:Nt) {
    kappa[t] ~ normal(kappa[t-1], tau_kappa);
   }
   
   for (a in 1:Na){
     for (t in 2:Nt){
       alpha[a,t] ~ normal(alpha[a,t-1], tau_alpha);
     }
   }
   
y ~ student_t(nu, mu, sigma_y);
}

```

---

<div class="post-metadata">

**Author:** ![aaronjg](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/aaronjg/32/3768_2.png) [@aaronjg](https://discourse.mc-stan.org/u/aaronjg)\
**Post date:** [May 30, 2018, 7:38pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/6 "2018-05-30T19:38:00Z")

</div>

Could you post the pairs plot for `sigma_y, tau_alpha, tau_kappa, tau_eta, nu, beta_raw[i], inc[i], kappa[i], alpha[j,i]` for a a few `i,j`?

The model may not actually have worked well in BUGS. BUGS has fewer diagnostics than Stan, so many models that appear to work in BUGS and failed in Stan, actually were never working in BUGS to begin with.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [June 5, 2018, 1:57am UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/7 "2018-06-05T01:57:26Z")

</div>

> [@sar2160](#):
>
> I can’t post the BUGS model in text because it’s trapped in a pdf

i didn’t see a link.

can you cut and paste from the pdf?

kappa still has a centered parameterization. do you get divergences? does it work with simulated data?

---

<div class="post-metadata">

**Author:** ![sar2160](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sar2160/32/2370_2.png) [@sar2160](https://discourse.mc-stan.org/u/sar2160)\
**Post date:** [June 6, 2018, 9:04pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/8 "2018-06-06T21:04:15Z")

</div>

I basically never get divergences, just exceeding treedepth. Haven’t tried simulated data yet.

Here’s the BUGS code. Note I’m not including delta in Stan.

 ![43%20PM](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/a/af52d6268803241ca6c54dd0d5fc5b13a6c496eb.png)

Here is a pairs plot, I’m not sure what exactly to look for but it seems like alpha being bimodal is worrisome? And alpha is correlated with energy. Pairs for other (i,j) look similar, although sometimes alpha will be unimodal or sometimes even trimodal.

 ![44%20PM](https://canada1.discourse-cdn.com/flex030/uploads/mc_stan/original/2X/1/1eb573a37b72c425e0d964f7a391ec2716903d6f.jpg)

Thanks,

---

<div class="post-metadata">

**Author:** ![aaronjg](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/aaronjg/32/3768_2.png) [@aaronjg](https://discourse.mc-stan.org/u/aaronjg)\
**Post date:** [June 7, 2018, 1:40am UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/9 "2018-06-07T01:40:09Z")

</div>

What is alpha in the pairs plot? In your model you don’t have a scalar quantity alpha.

---

<div class="post-metadata">

**Author:** ![sar2160](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sar2160/32/2370_2.png) [@sar2160](https://discourse.mc-stan.org/u/sar2160)\
**Post date:** [June 7, 2018, 3:21pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/10 "2018-06-07T15:21:43Z")

</div>

Alpha is an _[i,j]_ matrix, I am plotting just one representative _i_ and _j_ here.

Thanks.

---

<div class="post-metadata">

**Author:** ![aaronjg](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/aaronjg/32/3768_2.png) [@aaronjg](https://discourse.mc-stan.org/u/aaronjg)\
**Post date:** [June 7, 2018, 6:00pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/11 "2018-06-07T18:00:25Z")

</div>

Try using the ‘Matt trick’ on the alphas. Parameterize the time series in terms of unit normal innovation terms, and then get the alphas in the transformed parameter block.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [June 7, 2018, 11:17pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/12 "2018-06-07T23:17:51Z")

</div>

That is,

```
parameters {
  matrix[Na, Nt] alpha_std;

transformed parameters {
  matrix[Na, Nt] alpha;
  for (a in 1:Na) {
    // implicit: alpha[a, 1] ~ normal(mu_alpha0, tau_alpha0)
    alpha[a, 1] = mu_alpha0 + tau_alpha0 * alpha_std[a, 1]; 
    for (t in 2:Nt)
      // implicit: alpha[a, t] ~ normal(alpha[a, t - 1], tau_alpha)
      alpha[a, t] = alpha[a, t - 1] + tau_alpha * alpha_std[a, t - 1]
  }
    
model {
  to_vector(alpha_std) ~ normal(0, 1);

```

It would be a bit more efficient to compute the alpha term by scaling all  
the `alpha_std` at once with a scalar-matrix multiply and then putting `alpha` together using `cumulative_sum`.

---

<div class="post-metadata">

**Author:** ![sar2160](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/sar2160/32/2370_2.png) [@sar2160](https://discourse.mc-stan.org/u/sar2160)\
**Post date:** [June 12, 2018, 1:52pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/13 "2018-06-12T13:52:13Z")

</div>

Thanks, that seems to have solved the problem, no more iterations exceeding max\_treedepth as long as I set the max to 11.

---

<div class="post-metadata">

**Author:** ![changmoo](https://avatars.discourse-cdn.com/v4/letter/c/7ba0ec/32.png) [@changmoo](https://discourse.mc-stan.org/u/changmoo)\
**Post date:** [September 18, 2019, 10:36pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/14 "2019-09-18T22:36:53Z")

</div>

I am having a similar problem in converting BUGS code into STAN. Could you upload a full version of the hierarchical trend repeat sales model? I could follow the piecemeal inputs. It will be helpful a lot.

---

<div class="post-metadata">

**Author:** ![martinmodrak](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/martinmodrak/32/133_2.png) [@martinmodrak](https://discourse.mc-stan.org/u/martinmodrak)\
**Post date:** [September 24, 2019, 8:22am UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/15 "2019-09-24T08:22:28Z")

</div>

Tagging @sar2160 to make sure they are the post (and hopefully can share the Stan code)

---

<div class="post-metadata">

**Author:** ![changmoo](https://avatars.discourse-cdn.com/v4/letter/c/7ba0ec/32.png) [@changmoo](https://discourse.mc-stan.org/u/changmoo)\
**Post date:** [October 11, 2019, 12:49pm UTC](https://discourse.mc-stan.org/t/repeat-sales-regression-ported-from-bugs/4267/16 "2019-10-11T12:49:42Z")

</div>

I am trying to solve the same issue but not very successful. Could upload the revised and working Stan code for the Hierarchical Trend Repeat Sales Model with Local Linear Trend and t-distribution? It is very difficult to follow the piecemeal codes.
