I’m writing to ask if people have recommendations for models of profiling/benchmarking.
The data I have at hand is 25K paired observations of new code vs. old code timing (or old code vs. old code for a control). There is also a new code vs. new code data set to measure false positives. Each experimental condition comes with covariates about things like whether the test was single, double, or quadruple precision, whether it involves complex numbers or not, whether it is a matrix, array, or scalar operation, etc. etc. There are 12 paired replicates of each experimental condition. Other benchmarking setups may vary, but the general organization will be the same.
The covariates I have are either binary (e.g., does it involve a complex number, is the operation stripped over matrices) or finite categorical (which test suite it is in, which position it was among the 12 replicates, etc.). Some can be treated as continuous, like problem dimension.
There are two obvious choices here. We can model each system independently and then compare in the posterior, or we can model the differences in the paired comparisons. Modeling differences is a lot easier because the log of the ratio of timings is a natural scale that removes the actual timing baseline and makes the units ratios of performance. Modeling the actual time that all of these processes take is possible and I have the data broken down for it, but it’s going to be a lot harder, I think. So the model I’m thinking about is something like this for n pairwise observations of the A code timing, t^\text{A}_n and B code timing, t^\text{B}_n.
\log(t^\text{A}_n / t^\text{B}_n) \sim \textrm{Student-t}(\nu, \alpha + \beta \cdot x_n + ..., \sigma)
The \ldots is for the varying effects I don’t know how to easily write down in a sketchy way.
Edit: (oops, fat-fingered that out too early). Here’s some of the literature I’ve found so far:
-
Robust benchmarking in noisy environments. Jiahao Chen, Jarrett Revels and Alan Edelman. [This is how Julia does this evaluation]
-
Rigorous benchmarking in reasonable time. Tomas Kalibera and Richard Jones. ACM SIGPlan. 2013.
-
Performance evolution of configurable software systems: an empirical study. Christian Kaltenecker, et al. Springer Nature. 2023.
I would love to get feedback on what I should actually be reading, especially if there’s someone who understands Bayesian statistics and has worked in this area. I’m OK reading the frequentist stuff, too.