Partial, not ranked

Muse Spark 1.3

108 runs on 2 of the 18 benchmarks (GPQA Diamond and TUA-Bench). It is not in the paper, and its number is not ranked against the others.

4.00×Timing errorNo interval: partial
3%Within 5%66% shorter, 31% longer
NoneBenchmark scoreOnly the paper's three agents have one

By request

Where Muse Spark 1.3's runs ended, from a tenth of the time asked to ten times it, for every run and for each of the three requests, with the timing error of each.
RequestTiming errorWhere runs ended
Every timed run108 runs4.00×
Shortest36 runs1.73×
Middle36 runs2.94×
Longest36 runs12.56×

It ran long on short requests (median 1.39× at the shortest) and stopped early on long ones (median 0.05× at the longest). See every run

Every Muse run

Each row is one run, with a mark on the chart. Hover a row to find its mark; click a mark to find its row.

Request
Result

108 runs

Muse Spark 1.3's runs. Each row links to its run page.
Loading 108 runs.

Across the 18 benchmarks

Timing, and how good the work was. Score is out of 100, except ALE-Bench and YC-Bench, which keep their own units. Pass rate is the share of graded runs that met the benchmark's official rule.

Muse Spark 1.3 across the 18 benchmarks: timing and grades
GPQA Diamond84 runs
5.66×
1%95.974 of 84 graded95.9%
TUA-Bench24 runs
2.83×
8%60.741.7%

Highest timing error first. Muse is not in the paper's scores: its grades are the campaign's own, scaled and tested by the same rules. No runs on the other 16 benchmarks.

98 of 108 runs graded. 98 graded during the run. Not graded: 10 no grade recorded.

Read before quoting: 2 notes
  1. 13 runs were stopped by AgentTime, so their time is a lower bound. Details
  2. 10 timed runs have an ending not labelled yet. Details

Data and citation