HumanEval.org

Data downloads

Nightly anonymized dumps of every ratings-eligible battle and vote — the exact inputs the official leaderboards are computed from. Stable URLs: /downloads/files/latest/* always points at the newest dump. Prefer machine access? See the public JSON API.

Available dumps

2026-09-23latest

generated 2026-09-23T03:15:10Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B517f0e4bbe03846ae3a197f68629e6ff04f3fcdeb37b98bffb80657ee9b39622
votes.jsonl079 Bdf7861f03b4be56abb605693cc96b78425042fe78ce6009f0ece0ab26db5ce49

Manifest: manifest.json

2026-09-22

generated 2026-09-22T03:15:09Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B81f58878905df95687f27e7766369af6eabde87bb7f95d345266fe831b21478b
votes.jsonl079 Bf2649b1a68e76adab22156c6ab6ac2715ccefc0b4b4da37cfe87aa0317f8d67a

Manifest: manifest.json

2026-09-21

generated 2026-09-21T03:15:21Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Be90e062c9e695f4da373d0f080423e0cb99a5116ab1f3de72a8431cbe10597a3
votes.jsonl079 B40c42b82c3c7775c73080be10423f229104b258fff78de465e5c00311394a9a8

Manifest: manifest.json

2026-09-20

generated 2026-09-20T03:15:01Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B2b5eadc948bf8021cad1a1badc22ac630d2d062c0c9c8eae43403e95c7206091
votes.jsonl079 Bac5b163c0a24373149efaa83bc965b629f34acdae2d80ad50ea2be3d09bb1418

Manifest: manifest.json

2026-09-19

generated 2026-09-19T03:15:00Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Ba9ef0a898827a9816f73cf762616609b9ee5f40efb62f2c5bd0d2d4586ab2622
votes.jsonl079 B796defcc3730cfef971b70d3d34bfd4966dd60b585dabfb540b0d068a0f1a96b

Manifest: manifest.json

2026-09-18

generated 2026-09-18T03:15:02Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B7173defc5983c48ca2a224d49edbf9aaa0f3367e88857249948780e3fa04da7b
votes.jsonl079 Bacb76a724c14ed6021ff3dac21098592990394f34299f37a2e1a9aa86545fbf3

Manifest: manifest.json

2026-09-17

generated 2026-09-17T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Be91f817f800191b33fb4b50c69ed7df466147684e5625d69b9a4de3004197c73
votes.jsonl079 Bac15d831ff6e8c24cb01a14de289fd4b853887ac8cee7e79c5a64d79b5de0e5f

Manifest: manifest.json

2026-09-16

generated 2026-09-16T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bdd15d5572c14fb858cf85f741256999c8c483ace66190cf5a2b1c323c037a705
votes.jsonl079 Bec3cc6e0ab9984db3079d7072ed106012263648c6d3c311ce0ae33583e4dbdf6

Manifest: manifest.json

2026-09-15

generated 2026-09-15T03:15:02Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B8ec98ccfce807166bff264710f44a600cbe1002a38ff24c82c76233655d7fcf2
votes.jsonl079 B23dd643a5be360a4bba49a23b6e5c0fe37630e59de03d044dd62bdd4881fc3a5

Manifest: manifest.json

2026-09-14

generated 2026-09-14T03:15:06Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B474769b88044c0d4d735189418524a9c4edc0d8ab6eafb0a26aa39a3788accff
votes.jsonl079 B45dd86c45171ddd58b81f258b9f62ec9fe0b57687615c119669fcfb4553f544b

Manifest: manifest.json

2026-09-13

generated 2026-09-13T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Be316405de66fc78d5f43e9dbe61811da0aaf3dcd551f69591d8ff8621050c76b
votes.jsonl079 B9cf694d2fc599fc330bd30b8d468faf7e2bb39a0080fea6c3508276cdb259db3

Manifest: manifest.json

2026-09-12

generated 2026-09-12T03:15:16Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bd9289a2ac9deff301f678920cc75d7552f5429ec052742a03ffd3cabd9e6a741
votes.jsonl079 Ba08a4f56ef627e0c37ff5a54ca4d08b53343aaac6819a5d0573f0abc0ad134e1

Manifest: manifest.json

2026-09-11

generated 2026-09-11T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bc0fb4437b8145fe41431800ef69b1fdde49e02d77cf9fd51cc823c61b30b2e23
votes.jsonl079 B4200dd00b1e2792136078bc9590e76f28b456d86a5c507b0b08d08b5fc9e6e0f

Manifest: manifest.json

2026-09-10

generated 2026-09-10T03:15:15Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bc71a6074e6d431c7fbdcd28ae302acf6fd4491f220ca5a66184ce9307cc0b25f
votes.jsonl079 B4494a7f3ddf0de57262fe05883470ba375d6d10c4f4436031030cbcfd408f0ef

Manifest: manifest.json

2026-09-09

generated 2026-09-09T03:15:13Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B8e3a338af5d932b05ce59b1056da8902d7a77f04b0474d0c3853c265df52bc9c
votes.jsonl079 B9b4de26a9e7ea0bc605f915fbd81a043ff9c4696c1bec79dad99c2bd524bb2c3

Manifest: manifest.json

2026-09-08

generated 2026-09-08T03:15:03Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B408b479eb1418b08c68a7f2a327fde7a4c24b7ebb1a5da7dd03982fc1a4fa172
votes.jsonl079 B840dee02774c40d7285ff91395c7b2ca806d6f1fa21a7752b19b296c892f541d

Manifest: manifest.json

2026-09-07

generated 2026-09-07T20:24:41Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bf0c637515cc4cf4c439620948bed7da26baea4001f7f87eb6652505e404dd07a
votes.jsonl079 Bca939e034bbd3cfd1a7196ff0622ef4889e030be8db42abc13db14d201294357

Manifest: manifest.json

Reproduce the official leaderboard

The official numbers are produced by running the open humaneval-ratings package on these dumps with the published seed — no database, no network, no hidden inputs. Byte-identical reproduction is guaranteed under the package's pinned environment (requirements-lock.txt):

# 1. Download the dump you want to verify
curl -O https://humaneval.org/downloads/files/latest/battles.jsonl
curl -O https://humaneval.org/downloads/files/latest/votes.jsonl
curl -O https://humaneval.org/downloads/files/latest/manifest.json

# 2. Check the digests against manifest.json
sha256sum battles.jsonl votes.jsonl

# 3. Install the rating engine with its exact pins
python3 -m venv .venv
.venv/bin/pip install -r requirements-lock.txt   # from the package source
.venv/bin/pip install humaneval-ratings           # or: pip install <source dir>

# 4. Compute — the official run uses seed 42, 100 bootstrap rounds
.venv/bin/humaneval-ratings compute \
    --battles battles.jsonl --votes votes.jsonl \
    --out leaderboard.json --seed 42

# 5. Compare with the published snapshot / API values
sha256sum leaderboard.json

Every leaderboard.json records the seed, bootstrap rounds, package version and the SHA-256 of its exact input files in its metadata, so third parties can verify each other's runs as well. Vote weights in the dump come from a hidden quality pipeline (see methodology) — their values are public, their derivation is not.

Dump format specification

The text below is the canonical DUMP_FORMAT.md shipped with the rating engine (single source of truth, rendered as-is).

HumanEval.org public data dump format — v2

This document is the canonical, normative specification of the HumanEval.org public data dump and of the leaderboard.json file the humaneval-ratings package produces from it. The platform's nightly export is built to match this spec; the official leaderboard is produced by running this package on the published dump, so anyone can reproduce the official numbers byte-for-byte (see README.md).

Version note. Dumps published from 2026-09-08 on carry schema_version: 2 (section 3.1 adds recording battles for the computer-use category). v2 is a superset of v1: every v1 record is a valid v2 text record. The engine (humaneval-ratings ≥ 1.1.0) reads both versions; v1-only readers must be updated before consuming new dumps (the header version is exactly what they check).

The rating mathematics applied to a dump is specified in METHODOLOGY.md.

1. General conventions

  • A dump consists of exactly two files: battles.jsonl and votes.jsonl.
  • Encoding: UTF-8, Unix line endings (LF), one JSON object per line, no blank lines, file ends with a trailing LF.
  • Header line: the first line of each file is a metadata object (section 2). All subsequent lines are data records.
  • IDs are JSON strings (battle_id, vote_id, voter_id). The platform's internal keys are 64-bit integers which can exceed JavaScript's 2^53 safe-integer range; strings keep IDs safe and opaque. IDs are unique within their file and stable across dumps.
  • Models and categories are identified by their public slugs (models.slug, categories.slug on the platform side), never by internal numeric ids.
  • Timestamps are ISO 8601 UTC with a trailing Z, second precision: 2026-08-22T19:31:04Z.
  • Forward compatibility: readers MUST ignore unknown object fields. Additive changes (new optional fields) do not bump schema_version; breaking changes do.
  • Privacy: battles and responses are public content. No user data is ever exported; voter_id is a pseudonymous opaque token (see 4.1).

2. Header lines

First line of battles.jsonl:

{"kind": "battles", "schema_version": 2, "generated_at": "2026-09-08T04:00:00Z", "categories": ["browser-use", "chat-writing", "chat-coding"]}

First line of votes.jsonl:

{"kind": "votes", "schema_version": 2, "generated_at": "2026-09-08T04:00:00Z"}
FieldTypeRequiredMeaning
kindstringyes"battles" or "votes"; must match the file's content
schema_versionintegeryesThis spec is version 2 (v1 dumps remain readable). Readers reject versions they don't support
generated_attimestampyesWhen the export was produced
categoriesarray of stringsbattles file onlyThe full category-slug set covered by this dump; every battle's category_slug MUST be in it

Headers may carry extra fields (e.g. a source URL); readers ignore them.

3. battles.jsonl — data records

One record per voted battle. Battles without eligible votes MAY be included (the engine ignores them) but the export SHOULD omit them.

{"battle_id": "1024", "category_slug": "chat-writing", "model_a": "gpt-alpha", "model_b": "claude-beta", "created_at": "2026-08-22T19:31:04Z", "normalization_mode": "rendered", "modality": "text", "response_a": "# Draft\n\nDear team, ...", "response_b": "Dear team, ..."}
FieldTypeRequiredMeaning
battle_idstringyesUnique battle identifier
category_slugstringyesCategory the battle was fought in; must appear in the header categories list
model_astringyesSlug of the model in slot A (as shown to the voter, post position-randomization)
model_bstringyesSlug of the model in slot B; must differ from model_a
created_attimestampyesBattle creation time
normalization_modestringyesOne of rendered, raw, noformat — how responses were displayed to voters (raw for recordings; not meaningful there)
modalitystringv2: optional, default texttext or recording (section 3.1). Absent in v1
response_astringtext battles: yesFull response TEXT produced by the slot-A model. MUST be absent for recording battles
response_bstringtext battles: yesFull response TEXT produced by the slot-B model. MUST be absent for recording battles

Notes for the platform exporter:

  • Response bodies live in object storage as text artifacts (artifacts.storage_ref); the exporter MUST resolve and inline the raw text (the stored artifact content, not the rendered HTML) into response_a/response_b. Style covariates are computed from these texts (METHODOLOGY.md §7), so their exact bytes matter.
  • v1 exported text-modality battles only. v2 adds the recording modality below; any further artifact modality is another bump.

3.1 recording battles (v2 — computer-use category)

Computer-use ("Browser Use") battles have no response text: each contestant drove a real browser through the same task on live websites, concurrently from an identical starting scene, and voters judged wall-clock-faithful recordings side by side (METHODOLOGY.md §10). A recording battle is exported as:

{"battle_id": "2048", "category_slug": "browser-use", "model_a": "gpt-alpha", "model_b": "claude-beta", "created_at": "2026-09-08T10:00:00Z", "normalization_mode": "raw", "modality": "recording", "recording_a": {"video_key": "battles/2048/a.mp4", "trace_key": "battles/2048/a.trace.jsonl", "screenshot_key": "battles/2048/a.png", "assisted": false}, "recording_b": {"video_key": "battles/2048/b.mp4", "trace_key": "battles/2048/b.trace.jsonl", "screenshot_key": null, "assisted": false}, "telemetry_a": {"duration_ms": 184000, "time_to_first_action_ms": 6200, "steps": 14, "median_step_ms": 9100, "model_ms_total": 121000, "harness_ms_total": 4900}, "telemetry_b": {"duration_ms": 152000, "time_to_first_action_ms": 4100, "steps": 11, "median_step_ms": 8700, "model_ms_total": 88000, "harness_ms_total": 3900}}
FieldTypeRequiredMeaning
recording_a, recording_bobjectyesPer-slot artifact references (below)
recording_*.video_keystringyesObject key of the constant-frame-rate video whose duration equals the run's wall clock (t=0 = the shared start tick)
recording_*.trace_keystringyesObject key of the JSONL action/event trace (cursor moves, clicks, typing with secrets masked, per-step model/harness latency) — the source of truth for what the agent did
recording_*.screenshot_keystring or nullyesObject key of the final screenshot, if captured
recording_*.assistedbooleanyestrue if a human intervened in THIS slot's run. Assisted battles are never exported (they are ratings-ineligible); the field exists so the record is self-describing
telemetry_a, telemetry_bobjectyesRun telemetry; every member is an integer or null
telemetry_*.duration_msinteger/nullyesWall-clock run length in ms
telemetry_*.time_to_first_action_msinteger/nullyesms from the shared start tick to the contestant's first dispatched action
telemetry_*.stepsinteger/nullyesNumber of observation→action steps
telemetry_*.median_step_msinteger/nullyesMedian per-step latency (model + harness)
telemetry_*.model_ms_totalinteger/nullyesTotal time spent waiting on the model
telemetry_*.harness_ms_totalinteger/nullyesTotal harness overhead (identical scaffold for every model; published so equal treatment is checkable)

Object keys are relative to the platform's artifacts bucket; the media themselves are large binary objects and are not part of the dump (they are served through the platform, not published in bulk). The rating engine does not read them: it uses model_a/model_b/ category_slug and the votes exactly as for text battles, and treats the style covariates of a recording battle as zero (METHODOLOGY.md §7.5).

4. votes.jsonl — data records

One record per ratings-eligible vote.

{"vote_id": "5001", "battle_id": "1024", "voter_id": "v_9f2c7a1e", "choice": "win_a", "created_at": "2026-08-22T19:32:40Z", "weight": 1.0}
FieldTypeRequiredMeaning
vote_idstringyesUnique vote identifier
battle_idstringyesThe battle voted on; MUST exist in battles.jsonl
voter_idstringyesPseudonymous voter token (see 4.1)
choicestringyeswin_a, win_b, or tie
created_attimestampyesVote time
weightnumberyesEffective vote weight, finite and > 0; 1.0 is the default full weight

4.1 Voter pseudonyms

voter_id is an opaque token derived by the platform (e.g. a keyed hash of the internal user id). It is stable within a dump (and across dumps, so long-term voter behavior is analyzable) but cannot be linked back to an account. The derivation is internal and never published.

4.2 Vote weights and the quality pipeline

The platform runs a hidden vote-quality pipeline (consensus agreement, gold-standard checks, behavioral signals, provisional-account windows). That pipeline runs before export:

  • Votes it excludes never appear in the dump.
  • Votes it down-weights appear with their reduced effective weight.
  • Full-quality votes carry weight: 1.0.

The dump therefore contains exactly the votes that count, each with the weight it counts at. Ratings computed from the dump MUST use these weights (METHODOLOGY.md §3). How weights are derived is deliberately unpublished (anti-gaming); that they are applied, and their values, are fully public here. The exporter also guarantees at most one vote per (battle, voter) pair; the engine does not re-check this.

5. Validation rules (normative for the engine)

The engine rejects a dump (non-zero exit, no output) when:

  1. A header line is missing, has the wrong kind, or an unsupported schema_version.
  2. Any required field is missing or has the wrong JSON type (for v2 recording battles: recording_a/recording_b missing or not objects, or a response_* text present).
  3. choice, normalization_mode or (v2) modality has an unknown value.
  4. weight is not a finite number > 0.
  5. A battle_id (in battles) or vote_id (in votes) is duplicated.
  6. A vote references a battle_id not present in battles.jsonl.
  7. model_a == model_b in any battle.
  8. A battle's category_slug is absent from the header categories list.

Unknown fields are ignored everywhere. Battles with zero votes are ignored. Categories with zero votes produce no leaderboard section. Models enter a category's leaderboard only via voted battles there.

6. Output: leaderboard.json

Produced by humaneval-ratings compute. Top-level shape:

{
  "metadata": {
    "package": "humaneval-ratings",
    "package_version": "1.0.0",
    "leaderboard_schema_version": 1,
    "dump_schema_version": 1,
    "dump_generated_at": "2026-08-23T04:00:00Z",
    "seed": 42,
    "bootstrap_rounds": 100,
    "min_votes": 30,
    "style_min_votes": 50,
    "input_digests": {
      "battles_sha256": "9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08",
      "votes_sha256": "60303ae22b998861bce3b28f33eec1be758a213c86c93c076dbe9f558c11c752"
    }
  },
  "categories": [
    {
      "category": "chat-writing",
      "vote_count": 8231,
      "style_control": true,
      "style_coefficients": {
        "length_chars": 24.1830,
        "markdown_density": 6.0021,
        "list_count": 3.1187,
        "header_count": -0.4402
      },
      "entries": [
        {
          "model": "claude-beta",
          "rating": 1041.2211,
          "ci_low": 1027.0110,
          "ci_high": 1055.9024,
          "rating_style_controlled": 1030.5470,
          "ci_low_sc": 1016.2001,
          "ci_high_sc": 1046.0193,
          "vote_count": 4110,
          "provisional": false,
          "rank": 1,
          "order": 1
        }
      ]
    }
  ]
}

6.1 Metadata

Every knob that affects the numbers is recorded: the seed, bootstrap round count, thresholds, package version, and the SHA-256 digests of the exact input files. There are no timestamps generated at compute timedump_generated_at is echoed from the battles-file header (dump_schema_version likewise) — so the output is a pure function of (inputs, CLI parameters, pinned environment). Verifiers check both digests, run the same command, and compare sha256(leaderboard.json).

6.2 Category objects

Sorted by category slug (ascending, bytewise). Fields:

FieldMeaning
categoryCategory slug
vote_countTotal eligible votes in the category (unweighted count)
style_controlfalse when the style-control fallback triggered (METHODOLOGY.md §7.4); then SC fields mirror the raw fields
style_coefficientsFitted shared style coefficients in rating points per +1 SD of each normalized feature difference; null when style_control is false
entriesModel rows, sorted by order

6.3 Entry fields

FieldMeaning
modelModel slug
ratingRaw Bradley-Terry rating (center 1000, Elo-equivalent scale; METHODOLOGY.md §4)
ci_low, ci_high95% bootstrap CI of rating (2.5th/97.5th percentiles)
rating_style_controlledRating from the style-controlled fit (METHODOLOGY.md §7)
ci_low_sc, ci_high_sc95% bootstrap CI of the style-controlled rating
vote_countUnweighted count of eligible votes on battles involving this model in this category
provisionaltrue when vote_count < min_votes; provisional models keep their estimates but get rank: null
rankDisplayed rank band (integer, non-provisional models only, else null); overlapping CIs share a band (METHODOLOGY.md §6)
order1-based sort position within the category (all models, including provisional)

Sorting and rank bands are computed from the rounded, published values (see 6.4), so the ordering and bands are verifiable from leaderboard.json alone: entries sort by rating descending, ties broken by model slug ascending; order is that position. Rank bands use the published ci_low/ci_high (METHODOLOGY.md §6).

6.4 Determinism and precision policy (normative)

  • Every floating-point value in leaderboard.json is rounded to 4 decimal places (round-half-even, Python round(x, 4); matches the platform's numeric(10,4) snapshot columns) and serialized with exactly four decimal digits (-?d+.dddd).
  • Object keys are sorted (bytewise ascending); output is ASCII-only (non-ASCII escaped as in JSON \uXXXX); separators are ", " / ": " with 2-space indent; LF newlines; single trailing LF.
  • The RNG and every iteration order in the engine are deterministic (METHODOLOGY.md §8). Same input files + same CLI parameters + the pinned environment (requirements-lock.txt) → byte-identical output.

7. Worked example

A minimal, valid dump (v1 headers — still accepted; a v2 dump would say "schema_version": 2 and may add "modality": "text"). battles.jsonl:

{"kind": "battles", "schema_version": 1, "generated_at": "2026-08-23T04:00:00Z", "categories": ["chat-writing"]}
{"battle_id": "1", "category_slug": "chat-writing", "model_a": "gpt-alpha", "model_b": "claude-beta", "created_at": "2026-08-22T10:00:00Z", "normalization_mode": "rendered", "response_a": "## Plan\n\n- step one\n- step two", "response_b": "Do step one, then step two."}
{"battle_id": "2", "category_slug": "chat-writing", "model_a": "claude-beta", "model_b": "gpt-alpha", "created_at": "2026-08-22T11:00:00Z", "normalization_mode": "raw", "response_a": "Short answer.", "response_b": "A considerably longer answer with more detail."}

votes.jsonl:

{"kind": "votes", "schema_version": 1, "generated_at": "2026-08-23T04:00:00Z"}
{"vote_id": "10", "battle_id": "1", "voter_id": "v_a1", "choice": "win_a", "created_at": "2026-08-22T10:05:00Z", "weight": 1.0}
{"vote_id": "11", "battle_id": "1", "voter_id": "v_b2", "choice": "tie", "created_at": "2026-08-22T10:06:00Z", "weight": 0.5}
{"vote_id": "12", "battle_id": "2", "voter_id": "v_a1", "choice": "win_b", "created_at": "2026-08-22T11:09:00Z", "weight": 1.0}

Reading it: battle 1 was won by its slot-A model (gpt-alpha) at full weight, plus a half-weight tie (which counts as half a win for each side, METHODOLOGY.md §2). Battle 2 was won by its slot-B model (gpt-alpha again — note the slots swapped). Running

humaneval-ratings compute --battles battles.jsonl --votes votes.jsonl \
    --out leaderboard.json --seed 42

yields a chat-writing leaderboard where both models are provisional (3 votes < the default --min-votes 30) with rank: null, gpt-alpha at order 1, and style_control: false (3 votes < the default --style-min-votes 50), so the SC fields mirror the raw ones.

8. Version history

  • v1 (2026-08-23): initial spec — battles + votes JSONL with header lines; leaderboard schema v1.
  • v2 (2026-09-08, humaneval-ratings 1.1.0 / humaneval-exporter 1.1.0): battle records gain modality; new recording modality (section 3.1) with per-slot recording_* object references, assisted flags and telemetry_* for computer-use battles, no response_* texts. Text records are unchanged apart from modality: "text". Exporter rule 8: assisted battles are excluded. leaderboard.json schema stays 1 (only metadata.package_version and dump_schema_version change); the engine's numbers for v1 dumps are byte-identical to 1.0.0.