Back to Blog

AIPOCH Open-Science Ranks First in the Reported ASI-Bench B1 and B2 Comparison

AIPOCH Open-Science scored 79.35 on ASI-Bench B1 and 56.11 on B2 in three local repetitions, ranking first in the reported comparison for both levels. Learn how B1–B4 differ, how the runs were configured, and how to read the full tables.

5 min read

AIPOCH Open-Science ASI-Bench B1 and B2 cover showing scores of 79.35 and 56.11 with the top five records for each level

AIPOCH Open-Science completed three local repetitions on ASI-Bench B1 and B2, scoring 79.35 and 56.11, respectively. Both results rank first in the local-and-public comparison reported here. “First” describes the numerical ordering of the disclosed records; it is not an official ASI-Bench leaderboard certification.

The benchmark has four prompt-level task conditions, not only B1 and B2: B1, B2, B3, and B4. B1 provides the complete research method and procedure. B2 provides the intended method and constraints but removes the full procedure. B3 leaves method selection to the agent. B4 adds factually correct but non-essential context to the B3 condition. This article reports B1 and B2 only; it does not extend these results to B3 or B4.

Key Takeaways

  • B1: 79.35, the highest numerical score among the B1 records listed here. B1 tests whether an agent can turn a complete research plan into a running program and deliver the results.
  • B2: 56.11, the highest numerical score among the B2 records listed here. B2 keeps the method and constraints but asks the agent to fill in the steps, order the methods, and organize the workflow.
  • B1–B4 are four task conditions. The B1/B2 results do not imply that AIPOCH has already evaluated B3/B4 or ranks first across all four levels.
  • The scores come from three local repetitions. The B1 mean covers 59 valid scored tasks; B2 is reported as 56.11 in the internal result summary.

For a benchmark run to become useful in real research, the score needs context: inputs, file versions, execution records, and verification evidence. AIPOCH Open-Science’s .science research package can carry selected conversation branches, file versions, Notebook records, and verification evidence for review, handoff, and archiving in another project or computer. Imported sessions are read-only; they do not execute code or restore credentials. See the official research-package documentation.

What does ASI-Bench measure?

ASI-Bench is an AI for Science evaluation built around project-level research tasks. Its official materials describe 60 long-horizon tasks across scientific areas. An agent must understand the research objective, process input data, write and run programs, and deliver inspectable code and research results. The benchmark evaluates whether a research workflow can be completed, rather than whether an agent can answer one isolated knowledge question.

The official task design defines four prompt levels with progressively less method guidance. They can also be understood as four task-set conditions:

LevelInformation providedMain capability tested
B1Scientific background, intended method, equations, and the full procedureTurn the given plan into a working program, execute it, and deliver the result
B2Intended method and relevant constraints, without the full procedureFill in implementation details, order the methods, and organize a complete workflow
B3Objective, data, constraints, and required outputs, without a prescribed methodSelect a method and independently construct and validate the workflow
B4The B3 condition plus factually correct but non-essential informationIdentify the important constraints and complete the research while resisting distraction

These are not four unrelated product leaderboards. By changing how much method information appears in the prompt, ASI-Bench observes the transition from executing a specified plan to organizing or selecting a plan. A first-place result for B1 and B2 therefore describes the two evaluated conditions; it does not cover B3 or B4.

B1: turning a complete research plan into working analysis

What B1 asks the agent to do

B1 gives the scientific background, intended method, equations, and full procedure. The agent must read the input data, turn the plan into executable code, run the analysis, inspect intermediate results, and deliver code, outputs, and method notes. The central question is: when the research route is already specified, can the system reliably carry it through?

In an ASI-Bench single-cell RNA-sequencing task, for example, the agent works with simulated cell-by-gene expression data. It must group cells, infer the order of cell-state changes, identify candidate marker genes, and estimate gene-regulatory relationships. The deliverables include runnable code, per-cell cluster and trajectory values, a candidate-gene list, a regulatory matrix, and an explanation of the method. The procedure reduces method-design uncertainty, but it does not perform data processing, code execution, result checking, or delivery for the agent.

AIPOCH’s B1 setup and result

AIPOCH Open-Science used GPT-6 (gpt-6-astra) / xhigh, with domain-specialist roles and research Skills, and completed three local repetitions. One of the standard 60 tasks was excluded from the mean because of an official scorer anomaly, so the reported B1 mean covers 59 valid scored tasks. The reported task-quality mean is 79.35. The source summary does not list the three individual run scores or explain how results were combined across repetitions.

The 79.35 is a quality mean across task-level scores; it does not mean that 79.35% of the tasks were simply correct. The configuration matched 11 specialist categories to tasks across astronomy, bioinformatics, statistics, chemistry, robotics, physics simulation, and other areas. The result reflects the combined AIPOCH agent configuration, rather than GPT-6’s standalone performance on every scientific task.

Among the 24 B1 records listed below, AIPOCH’s 79.35 is 4.57 points above the local Codex reference at 74.78, which uses the same model and named reasoning level. It is 7.06 points above the highest public reference listed here, Claude Code × Claude Opus 5 at 72.29. These gaps describe this comparison; they do not replace an official review under fully controlled conditions.

Full B1 comparison

The table preserves all 24 B1 records from the internal summary, ordered by score. Local repetitions and public references come from different sources; the order is the numerical ordering used in this article, not an official leaderboard rank.

Numerical order in this comparisonAgent × modelReasoningB1 score (100 max)Result type
1AIPOCH Open-Science × GPT-6 (gpt-6-astra)xhigh79.35Local evaluation (3 repetitions)
2Codex × GPT-6 (gpt-6-astra)xhigh74.78Local evaluation (3 repetitions)
3Claude Code × Claude Opus 5Max72.29Public reference
4Codex × GPT-5.6 Solultra71.78Public reference
5AIPOCH Open-Science × DeepSeek v4-flashmax69.64Local evaluation (3 repetitions)
6Claude Code × GLM-5.3Max63.05Public reference
7Codex × GPT-5.6 Solxhigh62.75Public reference
8Claude Code × Gemini 3.8 FlashMax61.33Public reference
9Claude Code × Qwen3.8-MaxMax59.14Public reference
10Codex × GPT-5.5xhigh57.57Public reference
11Kimi Code × Kimi K3Not disclosed56.16Public reference
12Claude Code × Kimi K3Max55.79Public reference
13Claude Code × GLM-5.2Max54.01Public reference
14Claude Code × Claude Opus 4.8Max52.48Public reference
15Claude Code × DeepSeek V4 Flash 0731Max51.85Public reference
16OpenHands × DeepSeek V4 Flash 0731Not disclosed50.60Public reference
17Claude Code × GLM-5.3-FlashMax49.27Public reference
18Claude Code × Kimi K2.7Max44.43Public reference
19Claude Code × MiniMax M3Max43.53Public reference
20Claude Code × DeepSeek V4 ProMax42.97Public reference
21Claude Code × MiMo V2.5 ProMax41.70Public reference
22OpenHands × DeepSeek V4 ProNot disclosed37.78Public reference
23Kimi Code × Kimi K2.7Not disclosed29.99Public reference
24MiMo Code × MiMo V2.5 ProNot disclosed27.73Public reference

B2: organizing a complete workflow from a specified method

How B2 differs from B1

B2 belongs to the same ASI-Bench task system. The research objective, input data, deliverables, and scoring contract correspond to B1, but the prompt provides the intended method and constraints without a full procedure. The agent must decide what to do first, how to connect the methods, how to fill in missing implementation details, and when the result is ready for inspection.

B1 therefore emphasizes “can the agent execute the complete plan?”, while B2 asks “can the agent organize the specified methods into a complete workflow?” B2 is still not the condition in which the agent chooses the research method from scratch; that is B3. B4 adds the separate challenge of filtering non-essential but factually correct context.

Using the same single-cell example, B1 can provide a step-by-step list from preprocessing and dimensionality reduction through clustering, trajectory analysis, and network inference. B2 still specifies the methods but asks the agent to arrange and connect them. The deliverables remain runnable code and research outputs, and the score focuses on the quality of the resulting work rather than only the completeness of a written plan.

AIPOCH’s B2 setup and result

AIPOCH Open-Science kept the same domain-specialist and Skills configuration used for B1, ran GPT-6 (gpt-6-astra) / xhigh, and completed three local repetitions. The reported B2 score is 56.11.

Across the 23 B2 records listed below, AIPOCH’s 56.11 is above the same-model, same-reasoning Codex local reference at 56.09 and above the highest public reference listed here, Codex × GPT-5.6 Sol / ultra at 49.57. The AIPOCH–Codex gap is only 0.02 points, so the useful conclusion is narrower: in this local-and-public result summary, AIPOCH’s B2 result is first by score and shows continued workflow progress after the full procedure is removed. It is not a statistical significance test and does not establish stable superiority under every task and run condition.

Full B2 comparison

The table preserves all 23 B2 records from the internal summary, ordered by score. The records come from different sources and may use different runtime conditions; this numerical order is a comparison in this article, not an official leaderboard rank. Each result should be read with its result type and evaluation boundary.

Numerical order in this comparisonAgent × modelReasoningB2 score (100 max)Result type
1AIPOCH Open-Science × GPT-6 (gpt-6-astra)xhigh56.11Local evaluation (3 repetitions)
2Codex × GPT-6 (gpt-6-astra)xhigh56.09Local evaluation (3 repetitions)
3Codex × GPT-5.6 Solultra49.57Public reference
4Claude Code × Claude Opus 5Max45.80Public reference
5Codex × GPT-5.6 Solxhigh42.96Public reference
6Claude Code × Gemini 3.8 FlashMax39.13Public reference
7Claude Code × Claude Opus 4.8Max36.25Public reference
8Claude Code × GLM-5.3Max35.45Public reference
9Codex × GPT-5.5xhigh35.28Public reference
10Claude Code × Kimi K3Max35.24Public reference
11Kimi Code × Kimi K3Not disclosed34.05Public reference
12Claude Code × Qwen3.8-MaxMax32.51Public reference
13Claude Code × GLM-5.2Max30.81Public reference
14Claude Code × GLM-5.3-FlashMax28.66Public reference
15Claude Code × DeepSeek V4 Flash 0731Max26.49Public reference
16OpenHands × DeepSeek V4 Flash 0731Not disclosed24.89Public reference
17Claude Code × Kimi K2.7Max23.84Public reference
18Claude Code × MiniMax M3Max20.86Public reference
19Claude Code × DeepSeek V4 ProMax18.79Public reference
20Claude Code × MiMo V2.5 ProMax18.49Public reference
21Kimi Code × Kimi K2.7Not disclosed17.01Public reference
22OpenHands × DeepSeek V4 ProNot disclosed16.66Public reference
23MiMo Code × MiMo V2.5 ProNot disclosed11.33Public reference

How should the B1 and B2 results be read?

The two scores provide two distinct signals:

LevelAIPOCH local scorePosition in this comparisonCapability signal
B179.35First among the 24 listed recordsExecute a complete plan, run the analysis, and deliver the outputs
B256.11First among the 23 listed recordsFill in missing steps, connect methods, and organize the workflow under constraints

The difference is not that one level represents “doing science” and the other does not. The difference is how much method information the prompt provides and how much workflow organization the agent must supply. B1 is closer to a researcher who has already fixed the method and needs reliable execution. B2 is closer to a researcher who gives the objective, method boundary, and constraints while asking the system to connect the intermediate steps.

That distinction is where AIPOCH Open-Science’s product value becomes practical. Projects, papers, files, Notebooks, code execution, and generated outputs can stay in one research context while models and Skills are configured for the task. The AIPOCH Open-Science scientific-agent workflow describes how specialist roles can support this kind of work.

When the run needs review or handoff, a .science research package can carry selected conversation branches, file versions, Notebook records, and verification evidence. It is designed to carry research context across projects or computers; imported sessions are read-only, do not execute code or restore credentials, exclude Side Chat conversations and private reading bookmarks, and include only the files selected during export. A received verification record describes the sender’s checks; it does not prove that the receiving computer reran the experiment. See the .science research-package documentation.

Local evaluation, seeds, and the official leaderboard boundary

Three local repetitions are not an official certification

This article reports three local repetitions of AIPOCH Open-Science using public task materials and a local runtime. B1 79.35 and B2 56.11 are local evaluation results. We place them beside collected public references to show the reported score range, while keeping the wording separate from official leaderboard certification.

seed31415 and seed42 are different concrete instances

ASI-Bench now documents two explicit scoring contracts: seed31415 publishes reference answers and can be scored locally with the repository scorers; seed42 keeps reference answers private and requires official service scoring. The two seeds share the task framework and B1–B4 structure, but their concrete data, parameters, and reference answers can differ. They should not be assumed to be the same test instances.

For that reason, two records carrying the same B1 or B2 label must still be read with their seed, scorer, model version, reasoning level, agent harness, and runtime policy. Public records are useful for showing a comparison range, but they do not form a controlled experiment by default.

Next step: evaluate B3 and B4

This article reports B1 and B2 only. B3 asks the agent to choose a research method without one being prescribed. B4 adds factually correct but non-essential context and tests whether the system can identify the constraints that matter while resisting distraction.

Any future B3/B4 report should preserve the same repetition count, task coverage, seed, scoring method, and configuration details so that the results form an interpretable longitudinal comparison with B1/B2. This article makes no inference about B3 or B4 scores.

Frequently Asked Questions

Does ASI-Bench have only B1 and B2?

No. ASI-Bench has four prompt-level task conditions: B1, B2, B3, and B4. B1 provides the full method and procedure; B2 keeps the method but removes the full procedure; B3 asks the agent to select the method; B4 adds non-essential but factually correct context. This article reports B1 and B2.

Has AIPOCH Open-Science won the official ASI-Bench overall leaderboard?

That is not the claim here. “First” means first by score in the comparison made from AIPOCH’s local repetitions and the collected public references as of the source test date. The official leaderboard involves formal submission, official scoring, and consistent run conditions.

Does a B1 score of 79.35 mean that 79.35% of tasks were completed correctly?

No. B1 is a 100-point mean after task-level result quality is scored. The reported mean covers 59 valid scored tasks; it is not a simple fraction of tasks answered correctly. The reported B2 score is 56.11. The source summary does not specify its valid-task count, list the three individual run scores, or explain how results were combined across repetitions.

What is the core difference between B1 and B2?

B1 gives the complete research plan and procedure, so the central challenge is reliable execution. B2 gives the intended method and constraints but removes the full procedure, so the central challenge is filling in steps, ordering methods, and organizing a runnable, inspectable workflow.

Does a .science package automatically rerun an experiment on another computer?

No. Imported sessions are read-only; .science does not execute code, restore credentials, or automatically validate the receiving environment. It carries selected research records and verification evidence for review, handoff, and archiving.

Data and sources

Disclaimer

This article reports AIPOCH Open-Science’s local repeated evaluations under ASI-Bench B1 and B2 conditions, together with a numerical comparison against the listed public references. It is not an official ASI-Bench leaderboard certification and does not replace official scoring, scientific review, or independent verification of any research conclusion. Check the task level, seed, model, agent harness, scorer, input files, and outputs before drawing operational conclusions.