# Function to return number of non-zero elements in a matrix

**URL:** <https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123>\
**Category:** Modeling\
**Created:** [January 27, 2022, 2:52pm UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123 "2022-01-27T14:52:54Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![tds151](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/tds151/32/5300_2.png) [@tds151](https://discourse.mc-stan.org/u/tds151)\
**Post date:** [January 27, 2022, 2:52pm UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123/1 "2022-01-27T14:52:54Z")

</div>

I’d like to use Stan’s sparse matrix functions, such as `csr_extract_w`, but I need to known the number of non-zero values in order to declare this vector inside a Stan script.

---

<div class="post-metadata">

**Author:** ![jsocolar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jsocolar/32/2486_2.png) [@jsocolar](https://discourse.mc-stan.org/u/jsocolar)\
**Post date:** [January 27, 2022, 2:58pm UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123/2 "2022-01-27T14:58:06Z")

</div>

Could you elaborate a little bit on your use case? I’m pretty sure (but not completely positive) that because Stan does not support discrete parameters, the only way for a matrix element to be a literal zero is if either

- the matrix is data or transformed data
- the matrix is of some constrained data type where certain elements are constrained to zero _a priori_.

In either case, the number of zeros should be _a priori_ known.

---

<div class="post-metadata">

**Author:** ![tds151](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/tds151/32/5300_2.png) [@tds151](https://discourse.mc-stan.org/u/tds151)\
**Post date:** [January 27, 2022, 3:25pm UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123/3 "2022-01-27T15:25:26Z")

</div>

I have a sparse (random effects) matrix mostly filled with 0’s, by construction that I will multiply by a vector of regression coefficients under a sparse configuration. For reasons of how the associated regression coefficients are indexed, I am not able to render the sparse (w,v,u) representation outside of Stan (in R). I must do so within Stan and it has explicit functions for this purpose; e.g. `csr_extract_w()`. The problem is that to declare a Stan vector to hold the output of `csr_extract_w(A)` where `A` is some sparse matrix, I need to know the number of non-zero elements in `A` (to dimension the length of the output vector). Now, I can’t use R to get the number of non-zero elements in `A` because I will be embedding the sparse matrix multiplication within a `reduce_sum()` function.

---

<div class="post-metadata">

**Author:** ![jsocolar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/jsocolar/32/2486_2.png) [@jsocolar](https://discourse.mc-stan.org/u/jsocolar)\
**Post date:** [January 27, 2022, 6:17pm UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123/4 "2022-01-27T18:17:34Z")

</div>

Got it. Yeah, that seems like a tough indexing problem to solve. I imagine there is a way to handle all the indices in a super clever way, but also that it might not be worth anyone’s time to figure it out (note that it would be a potential efficiency gain, however, to slice the sparse representation rather than the full matrix in `reduce_sum`, because it would ensure that each chunk gets a similar number of nonzero elements).

Anyway, here is a function for the number of nonzero elements. Whether to double-loop over the matrix elements or (as I did) to flatten the matrix and loop once is a matter of taste.

```stan
  int nnz(matrix x) {
    int N = num_elements(x);
    vector[N] x2 = to_vector(x);
    int out = 0;
    for(i in 1:N){
      if(x2[i] != 0) {
        out += 1;
      }
    }
    return(out);
  }

```

---

<div class="post-metadata">

**Author:** ![nikunj410](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/nikunj410/32/13085_2.png) [@nikunj410](https://discourse.mc-stan.org/u/nikunj410)\
**Post date:** [February 28, 2025, 8:11pm UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123/5 "2025-02-28T20:11:37Z")

</div>

I was stuck in a similar problem. I came up with a different solution that does not require loops.  
`non_zero = rank(append_row(to_vector(abs(x)), machine_precision()), num_elements(x)+1)`

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [March 28, 2025, 7:02pm UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123/6 "2025-03-28T19:02:18Z")

</div>

> [@nikunj410](#):
>
> I came up with a different solution

This is a super cool approach.

> [@jsocolar](#):
>
> Whether to double-loop over the matrix elements or (as I did) to flatten the matrix and loop once is a matter of taste.

It’s also a matter of memory pressure. Your approach will be suboptimal because it’s allocating a new `N`-vector and filling it, which will take about as long as the rest of the program, which also goes over a loop just computing some arithmetic. I’d suggest:

```stan
int nnz(matrix x) {
  int nz = 0;
  for (n in 1:cols(x)) {
    for (m in 1:rows(x)) {
      nz += x[m, n] != 0.0;
    }
  }
  return nz;
}

```

It’d probably be faster with `nz += !x[m, n];` to avoid the comparison and then `return rows(x) * cols(x) - nz;`.

---

<div class="post-metadata">

**Author:** ![Corey.Plate](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/corey.plate/32/20600_2.png) [@Corey.Plate](https://discourse.mc-stan.org/u/Corey.Plate)\
**Post date:** [March 29, 2025, 1:20am UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123/7 "2025-03-29T01:20:52Z")

</div>

For large arrays, what about:

```stan
int nz = to_int(sum(fmin(machine_precision(),abs(x)) ./ machine_precision()));

```

I tried to figure out a construction that avoided `to_vector()` and for-looping. It incurs a cost per infix operation but uses only operations natively vectorized in Stan.

---

<div class="post-metadata">

**Author:** ![Bob\_Carpenter](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.mc-stan.org/bob_carpenter/32/9230_2.png) [@Bob\_Carpenter](https://discourse.mc-stan.org/u/Bob_Carpenter)\
**Post date:** [March 31, 2025, 9:39pm UTC](https://discourse.mc-stan.org/t/function-to-return-number-of-non-zero-elements-in-a-matrix/26123/8 "2025-03-31T21:39:48Z")

</div>

> [@Corey.Plate](#):
>
> For large arrays, what about:

I try to stay away from things that are this clever.

The way you coded it in Stan will have a problem because Stan’s not clever enough to use expression templates for all the intermediates, so there is going to be a lot of implicit allocation and copying overhead. Both `fmin()` and `abs()` will allocate new vectors the size of `x`.

That `./` can just be `/` as the value on the LHS is a scalar.

> [@Corey.Plate](#):
>
> only operations natively vectorized in Stan.

Because Stan code gets compiled to C++, there’s no difference between writing a loop in Stan and writing in natively in C++. The only difference is when there is autodiff and we can reduce the total number of operations or the size of the autodiff graph.

The code I wrote doesn’t do any allocation or any autodiff, so it should be as fast as if you had written the code directly in C++. All our vectorized code does is code the loops in C++ (sometimes delegating to Eigen, which can vectorize at the CPU level).
