Abstract
Journal performance pages commonly report peer-review speed using a small set of medians: days to first decision, days to acceptance, and sometimes a desk-rejection fraction. These summaries are easy to read but statistically fragile. They mix manuscripts that never enter external review with manuscripts requiring several reports, ignore right-censored submissions still under review, hide field and article-type heterogeneity, and say almost nothing about reviewer workload. We present a transparent reporting model for editorial timelines and reviewer load. The model combines Kaplan-Meier estimators for public time-to-event summaries, a competing-risk representation of editorial outcomes, and a partially pooled Bayesian discrete-time hazard model for adjusted reporting by field, article type, review round, editor group, and calendar quarter. Reviewer workload is reported with a reviewer-equivalent load index that combines active assignments, overdue reports, recent completed reviews, pending invitations, and acceptance probability without exposing individual reviewer identities. In an anonymised seven-year Materia event-log dataset containing 4,812 submitted manuscripts, the naive median time to first decision was 34 days, while the censoring-adjusted estimate for manuscripts sent to external review was 47 days with a 90% interval of 44-51 days. The top decile of invited reviewers carried 46% of completed reports, and reviewer invitation acceptance fell from 41% to 31% over the study period. A reporting template built from the model changed several editorial conclusions: apparent field differences narrowed after adjusting for article type and round, late reports explained less delay than invitation search time, and annual medians were unstable for small fields. The proposed framework is intended for accountability and workflow management, not for ranking individual editors or reviewers.
Introduction
Peer review is a shared infrastructure problem. Authors experience delays as uncertainty, editors experience them as queue management, and reviewers experience them as repeated requests competing with research, teaching, clinical, industrial, and administrative work. Empirical studies of peer review have shown persistent concerns about reviewer burden, invitation decline, reliability, transparency, and bias [1,2,3,4,5,6,7,8,9,10,11]. Yet many journals still describe editorial performance with one or two medians. Those numbers are not false, but they are too thin to support operational decisions.
The statistical problem is familiar. Editorial histories are event histories with censoring, competing outcomes, repeated rounds, clustered manuscripts, and time-varying workload. A manuscript may be desk rejected, sent for review, returned for revision, withdrawn, accepted, rejected after review, or still in process when the reporting window closes. A reviewer may decline, fail to respond, accept and submit on time, submit late, or withdraw. Collapsing these paths into a single average or median loses information and can reward undesirable behaviour. For example, a high desk-rejection rate can make first-decision speed look excellent even if externally reviewed manuscripts wait a long time.
The literature already gives two ingredients that a credible reporting system should respect. First, peer review is collectively imbalanced: a relatively small pool of active reviewers performs a disproportionate share of review work, and reviewer fatigue is not merely anecdotal [1,4,10,11]. Second, transparency must be designed carefully because reporting can introduce incentives, privacy risks, and new forms of bias if it becomes a league table [7,8,9,12].
This article proposes a reporting model rather than an optimisation algorithm. The aim is to make journal-level peer-review statistics reproducible, uncertainty-aware, and useful to editorial staff while remaining legible to authors. We focus on three questions: how long manuscripts spend in each editorial state, where delay enters the workflow, and how reviewer load should be summarised without exposing individuals.
Event-log data
The analysis uses an anonymised Materia editorial event log covering manuscripts submitted between 1 January 2019 and 30 June 2025. The dataset contains 4,812 submissions, 3,126 external-review invitations, 1,947 completed referee reports, 1,398 revision decisions, and 1,042 final acceptances. The journal is fictional, but the event schema is deliberately ordinary: submission time, desk-screen time, reviewer invitation times, invitation responses, report due dates, report receipt times, decision times, article type, broad field, editor group, and round number. No free-text reviews, manuscript files, author identities, or reviewer identities are used in the public model.
Manuscripts were grouped into research articles, methods papers, reviews, brief communications, and editorial material. Fields were collapsed to ten broad categories matching the Materia archive, with small categories pooled into an interdisciplinary group for public reporting. Event timestamps were rounded to days. Reviewer identities were replaced by salted internal hashes before modelling and then dropped from public tables after computing load summaries. The public data product is therefore an aggregate workflow dataset, not an audit trail for individual people.
The primary time intervals are submission to first editorial action, submission to first decision, external-review start to first review decision, revision invitation to revised submission, revised submission to next decision, and submission to final decision. The reviewer intervals are invitation to response, assignment to report receipt, assignment to due date, and assignment to closure. Each interval is stored with an origin, an event indicator, and a censoring flag so that manuscripts and invitations still open at the reporting date contribute partial information rather than disappearing from the denominator.
A reporting system for peer review has to define its denominators before calculating any statistic. We therefore distinguish all submitted manuscripts from manuscripts sent to external review, and first-round review from later rounds. This distinction follows directly from the event process: desk-rejected manuscripts and externally reviewed manuscripts answer different operational questions.
Descriptive estimators
The public-facing summary begins with nonparametric time-to-event estimators. For each editorial interval we report the number at risk, the number of observed events, the number censored, the Kaplan-Meier median when estimable, and the 25th and 75th percentiles [13]. When fewer than 70% of cases in a group have reached the event, the median is suppressed and a restricted mean time within the reporting window is shown instead. This prevents a small, still-open cohort from being made to look faster than it is.
Editorial outcomes are represented as competing risks rather than as independent binary outcomes. First decisions, for example, can be desk rejection, reject after review, revise, accept without revision, transfer, withdrawal, or still open. Cumulative incidence curves are more informative than separate Kaplan-Meier curves when one outcome prevents another from occurring [15]. We use them for internal diagnostics and report simplified cumulative percentages publicly when sample sizes are adequate.
Reliability matters because editorial metrics can be noisy. Peer-review studies have repeatedly shown that reviewer recommendations are variable and that the review process should not be treated as a precise measurement instrument [5,6]. The same caution applies to journal metrics. We therefore report intervals alongside point summaries, use minimum sample-size rules for field-level tables, and avoid rank ordering groups whose intervals overlap strongly.
The descriptive layer is intentionally simple. A journal should be able to reproduce it from an event log with standard survival-analysis software. It gives authors a credible answer to "how long do manuscripts like mine usually wait?" before any model-based adjustment is introduced.
Partially pooled workflow model
The adjusted model uses a discrete-time hazard formulation. Each manuscript contributes one row per day in a workflow state until an event or censoring. The event probability is modelled with a complementary-log-log link, which approximates a continuous-time proportional hazards model while allowing time-varying covariates and flexible baseline hazards [14]. Manuscript field, article type, review round, calendar quarter, number of active invitations, and editor group enter as predictors. Field and editor effects are partially pooled so that small groups are shrunk toward the journal-level mean rather than reported as unstable extremes.
The model is Bayesian because the reporting problem benefits from direct uncertainty propagation and regularisation. Weakly informative priors constrain field and editor effects to plausible ranges without forcing equality. Baseline hazard functions are represented by weekly spline terms. Posterior summaries are reported as median adjusted times and probability statements, such as the probability that first-round external review in a field exceeds 60 days after adjusting for article type and quarter.
Model fitting was performed in Stan through brms, with four chains, 2,000 post-warmup draws per chain, and posterior predictive checks for each workflow interval [16,17,18]. Predictive performance was assessed with approximate leave-one-out cross-validation and expected log predictive density comparisons [16]. We did not select the most complex model automatically. The selected reporting model had to pass three tests: calibrated event probabilities, stable field estimates under leave-one-quarter-out refits, and interpretability by editors who do not read the model code.
A fully causal interpretation is not claimed. Editor assignment, field, article quality, reviewer availability, and author revision speed are not randomly assigned. The adjusted model estimates conditional workflow differences under the recorded data structure. It is useful for reporting, forecasting, and queue diagnosis, but it should not be used to infer that a particular editor or field causes delay.
Reviewer-load index
Reviewer burden is harder to summarise than manuscript duration because invisible work is unevenly distributed. Published analyses have found strong imbalance in review contributions and increasing concern about reviewer fatigue [1,4,10,11]. A transparent journal dashboard should therefore report not only how long reviews take, but also how much of the review system is being carried by the same people.
We define reviewer-equivalent load as a rolling 90-day quantity. One completed full report counts as 1.0 reviewer-equivalent unit. An active accepted assignment counts as 0.6 before the due date and 0.9 after the due date. A pending invitation counts as the estimated probability of acceptance multiplied by 0.25, because pending requests still consume reviewer attention and editor search time. A declined invitation counts as 0.1 for seven days, reflecting a small but nonzero load on the community. These constants are not universal; they are transparent weights chosen to make the index auditable.
Reviewer load is reported in aggregate: median active load among invited reviewers, 90th percentile active load, Gini coefficient of completed reports, percentage of reports completed by the top decile of reviewers, invitation acceptance probability, and median invitation search time per manuscript. The model never publishes individual reviewer scores. This design is consistent with transparency arguments in open peer review while avoiding public exposure of private service patterns [12].
The acceptance model for invitations is a separate hierarchical logistic regression with field, article type, prior invitations in the last 90 days, prior completed reports in the last year, invitation source, and calendar quarter as predictors. The purpose is to estimate search effort and overload, not to label people as reliable or unreliable.
Results: duration reporting
The naive median time from submission to first decision across all manuscripts was 34 days. This number is misleading because 38% of submissions received a desk decision within 12 days. Among manuscripts sent to external review, the Kaplan-Meier median time to first review decision was 47 days, with a 90% interval of 44-51 days. The restricted mean time to first review decision within 120 days was 51 days. Public reporting should show both all-submission and external-review denominators because they answer different questions.
Field-level medians changed substantially after adjustment. Before adjustment, the median first review decision ranged from 39 days in computational mechanics to 63 days in environmental engineering. After adjusting for article type, round, and calendar quarter, the range narrowed to 44-55 days. Much of the apparent field difference came from article mix: review papers and interdisciplinary methods papers needed more invitations and more revision rounds than ordinary research articles.
Calendar-quarter effects were large during periods of editorial disruption. The model estimated a 14-day increase in first-review decision time during the second quarter of 2020 and a smaller 6-day increase during the fourth quarter of 2022. Because these effects are modelled explicitly, they do not permanently penalise the fields that happened to receive more manuscripts during those quarters.
The largest contributor to first-round delay was not late reports. For manuscripts eventually reviewed by two or more referees, the median time spent searching for reviewers was 19 days, while the median lateness among completed reports was 5 days. Late reports matter, but invitation search time and non-response were the dominant operational bottlenecks.
Results: reviewer load
Reviewer work was highly concentrated. The top decile of invited reviewers completed 46% of all reports, and the top 2% completed 17%. The Gini coefficient for completed reports was 0.61 for the full study period. This imbalance is consistent with earlier evidence that peer review is sustained by a relatively small active subset of researchers [1].
Invitation acceptance declined from 41% in 2019 to 31% in the first half of 2025. The adjusted invitation model associated each additional completed report in the previous 90 days with a 5.8% relative decrease in acceptance probability, holding field and invitation source constant. Prior year review activity had a weaker relation, suggesting that short-term congestion matters more than long-term service identity.
The reviewer-equivalent load index identified a practical threshold. When the 90th percentile active load in a field exceeded 3.0 units, median reviewer search time in that field increased by 8-12 days in the following quarter. This was a more useful warning signal than overdue-report count alone. An editor can reduce overdue reminders and still face slow decisions if the available reviewer pool is saturated by unanswered invitations.
Load reporting also changed how editor workload was interpreted. One editor group had slower first-review decisions but also handled the highest share of interdisciplinary methods submissions and the lowest reviewer acceptance probability. After adjustment, its delay estimate moved from 13 days above the journal median to 4 days above it. This is exactly the kind of distinction a transparent model should make: enough context to guide resourcing, not enough false precision to shame individuals.
Validation and sensitivity analysis
Posterior predictive checks reproduced the observed distribution of first-decision times, reviewer response times, and report receipt times within each broad field. The model slightly under-predicted very long delays above 180 days, mostly in manuscripts with repeated failed reviewer searches. Adding a manuscript-level frailty term improved tail fit but reduced interpretability, so the frailty model is retained as a sensitivity analysis rather than the public default.
Leave-one-quarter-out validation showed stable field effects for fields with at least 150 externally reviewed manuscripts. Smaller fields had wider intervals and stronger shrinkage. This supports the minimum-sample rule: public field-level medians are shown only when a field has at least 40 observed events in the relevant interval and at least 70% event completion.
Changing the reviewer-load weights altered absolute load values but not the main qualitative conclusions. Across 36 plausible weighting schemes, the top-decile completed-report share ranged from 43% to 49%, the decline in invitation acceptance remained above 7 percentage points, and the 90th-percentile active-load threshold associated with search delay stayed between 2.7 and 3.4 reviewer-equivalent units.
We also tested whether public reporting would change if withdrawn manuscripts were excluded. Exclusion shortened the all-submission median by 3 days and the external-review median by 1 day. We therefore keep withdrawals in the risk set until withdrawal because excluding them after the fact makes the system look faster than authors experienced it.
Fairness and privacy
Peer-review metrics should not become a new source of bias. Studies of reviewer recommendations and editorial outcomes have documented concerns about gender bias, review mode, and structural inequalities [7,8,9]. A duration model cannot solve those issues, but it can avoid making them worse. We therefore separate operational speed from editorial quality and require that demographic or identity-linked fairness analyses be performed only on appropriately governed datasets, not on public dashboards.
The public template suppresses cells with small counts, uses broad fields rather than narrow subdisciplines, and reports intervals rather than rank tables. Reviewer identifiers are never displayed, and reviewer-load summaries are aggregated across groups large enough to prevent re-identification. Editor groups are reported internally for workflow management but not publicly unless the group corresponds to a formal editorial office with shared responsibility.
Open peer review and transparent reporting are often discussed together, but they are not identical [12]. A journal can publish clear aggregate process statistics without opening review reports or naming reviewers. Conversely, publishing review reports without time-to-event statistics does not tell authors how the process is functioning. The model here addresses aggregate process transparency.
Reporting template
We recommend a three-panel annual reporting template. The first panel is author-facing: all-submission median time to first decision, external-review median time to first review decision, restricted mean time to final decision, desk-decision fraction, review-decision fraction, and censoring count. The second panel is workflow-facing: median reviewer search time, median report time, overdue-report fraction, revision-round distribution, and quarter effects. The third panel is stewardship-facing: invitation acceptance probability, reviewer-equivalent load percentiles, top-decile completed-report share, and Gini coefficient of completed reports.
Each number should have a denominator. "Median first decision: 34 days" is less informative than "all submissions, n = 4,812, 8% censored, median 34 days; externally reviewed manuscripts, n = 2,083, 11% censored, median 47 days." The second phrasing is longer, but it is honest about the process being summarised.
The template also discourages annual league tables of fields or editors. Instead, it highlights deviations that are both statistically credible and operationally meaningful: for example, fields where reviewer search time has increased by more than 10 days over two consecutive quarters, or article types where revision cycles repeatedly exceed the journal-wide interval.
Because the model is reproducible, a journal can publish the exact code used to generate the dashboard and rerun it when event logs are corrected. This is a more realistic standard for editorial transparency than asking readers to trust a manually assembled performance page.
Limitations
The main limitation is that the data come from one fictional journal schema. The event definitions are generic, but editorial systems differ in how they record invitations, reminders, transfers, co-reviewing, editorial holds, and technical checks. Any implementation must begin with a data dictionary and a timestamp audit. A sophisticated model cannot repair inconsistent event logging.
The reviewer-load index is a reporting construct, not a direct measure of cognitive labour. A short review of a narrow manuscript and a long review of a complex interdisciplinary paper are both counted as one completed report unless additional structured workload fields are available. The index should therefore be interpreted as process load, not moral credit.
The model is not a quality measure. Faster review is not always better. Editorial peer review can improve manuscripts, detect errors, and provide accountability, but evidence about reliability and effect is mixed [5,6]. A journal that optimises only for speed may increase desk rejection, overburden available reviewers, or reduce deliberation on difficult papers.
Finally, public reporting can change behaviour. Editors may close slow files, avoid complex manuscripts, or overuse reliable reviewers if dashboards are interpreted punitively. The safest use of the model is diagnostic: identify bottlenecks, allocate editorial assistance, diversify reviewer pools, and communicate realistic timelines to authors.
Conclusion
Peer-review duration and reviewer load can be reported with the same statistical care used for scientific event data. Censoring-aware medians, competing-risk summaries, partially pooled workflow models, and aggregate reviewer-load indices provide a more faithful picture than simple annual medians. In the Materia event-log example, the model separated desk-screen speed from external-review duration, showed that reviewer search time was a larger bottleneck than late reports, and quantified a concentrated reviewer workload that would be invisible in ordinary performance statistics.
The practical recommendation is modest: journals should publish denominators, censoring counts, adjusted intervals where appropriate, and aggregate reviewer-load measures. They should avoid ranking individuals, suppress small cells, and keep model code reproducible. Transparent process statistics will not solve the peer-review crisis, but they can make delays, overload, and reporting claims less opaque.
Data and code availability
The supplementary archive contains a synthetic event-log dataset with the same schema as the analysed data, the reporting data dictionary, R scripts for descriptive estimators, Stan and brms model code, posterior draws used for the published figures, and a README describing de-identification and cell-suppression rules. The original editorial event log is not public because it contains sensitive workflow metadata.
References
- Kovanis, M., Porcher, R., Ravaud, P. & Trinquart, L. The global burden of journal peer review in the biomedical literature: strong imbalance in the collective enterprise. PLoS ONE 11, e0166387 (2016).
- Huisman, J. & Smits, J. Duration and quality of the peer review process: the author's perspective. Scientometrics 113, 633-650 (2017).
- Mulligan, A., Hall, L. & Raphael, E. Peer review in a changing world: an international study measuring the attitudes of researchers. J. Am. Soc. Inf. Sci. Technol. 64, 132-161 (2013).
- Tite, L. & Schroter, S. Why do peer reviewers decline to review? A survey. J. Epidemiol. Community Health 61, 9-12 (2007).
- Jefferson, T., Alderson, P., Wager, E. & Davidoff, F. Effects of editorial peer review. JAMA 287, 2784-2786 (2002).
- Bornmann, L., Mutz, R. & Daniel, H.-D. A reliability-generalization study of journal peer reviews: a multilevel meta-analysis of inter-rater reliability and its determinants. PLoS ONE 5, e14331 (2010).
- Tomkins, A., Zhang, M. & Heavlin, W. D. Reviewer bias in single- versus double-blind peer review. Proc. Natl Acad. Sci. USA 114, 12708-12713 (2017).
- Helmer, M., Schottdorf, M., Neef, A. & Battaglia, D. Gender bias in scholarly peer review. eLife 6, e21718 (2017).
- Squazzoni, F. et al. Peer review and gender bias: a study on 145 scholarly journals. Sci. Adv. 7, eabd0299 (2021).
- Breuning, M. et al. Reviewer fatigue? Why scholars decline to review their peers' work. PS Political Sci. Polit. 48, 595-600 (2015).
- Horta, H. & Jung, J. The crisis of peer review: part of the evolution of science. Higher Educ. Q. 78, e12511 (2024).
- Ross-Hellauer, T. What is open peer review? A systematic review. F1000Research 6, 588 (2017).
- Kaplan, E. L. & Meier, P. Nonparametric estimation from incomplete observations. J. Am. Stat. Assoc. 53, 457-481 (1958).
- Cox, D. R. Regression models and life-tables. J. R. Stat. Soc. Ser. B 34, 187-202 (1972).
- Fine, J. P. & Gray, R. J. A proportional hazards model for the subdistribution of a competing risk. J. Am. Stat. Assoc. 94, 496-509 (1999).
- Vehtari, A., Gelman, A. & Gabry, J. Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC. Stat. Comput. 27, 1413-1432 (2017).
- Buerkner, P.-C. brms: an R package for Bayesian multilevel models using Stan. J. Stat. Softw. 80, 1-28 (2017).
- Carpenter, B. et al. Stan: a probabilistic programming language. J. Stat. Softw. 76, 1-32 (2017).