Judge My ReviewersBack to reviews

Trust the receipts

Methodology & limitations

How the scores are produced, what they can tell you, and where the machine should sit down.

What we score

The current public dataset covers 26,749 papers from six 2025 OpenReview venues: ICLR, ICML, NeurIPS, TMLR, COLM, and MIDL. It contains 128,723 public official reviews; 99,671 have completed AI scoring.

Inputs are limited to paper title, abstract, decision metadata, and public official review text. PDF full text, author rebuttals, discussion replies, and administrative comments are excluded from the public ranking pipeline.

What makes the board

We publish one thing: outrageous public peer reviews. An internal 0-to-100 score and a small set of behavior types identify comments that cross a clear professional boundary. Ordinary rejection, technical disagreement, requests for experiments, and merely unhelpful prose do not qualify.

The score is not displayed as an objective verdict. Readers use the like and dislike controls to agree or disagree with a thread's placement, and those community votes determine the durable all-time ranking.

Model and feed

The current outrage pass uses gpt-5.6-luna with versioned structured output. Luna re-reviewed 1,095 candidates from the original 99,671 scored reviews; 66 concise, feed-worthy entries are currently published. Each entry may include one short editorial aside written to set the tone, not to explain the judgment or replace the discussion.

Known limitations

This is an experimental text audit, not an objective measure of a person. AI selection and editorial asides can be wrong.
  • There is not yet a representative human gold-standard set or published inter-annotator agreement result for this exact outrage scorer.
  • Model judgments may react unevenly to non-native English, field-specific style, sarcasm, or terse venue conventions.
  • An excerpt can omit context found elsewhere in the same review. Always check the linked OpenReview source.
  • Venue forms and reviewing cultures differ. Cross-venue comparisons should be treated as exploratory.
  • The system evaluates public text. It does not infer, identify, or rank anonymous reviewers as people.

Corrections

Every detail view links to its OpenReview source and includes a score-report form. Reports should identify the paper, review, disputed selection, and reason. Credible privacy, attribution, or safety concerns are prioritized for hiding while they are reviewed.

Last updated: August 8, 2026.