
AIPOCH Open-Science completed three local repetitions on ASI-Bench B1 and B2, scoring 79.35 and 56.11, respectively. Both results rank first in the local-and-public comparison reported here. “First” describes the numerical ordering of the disclosed records; it is not an official ASI-Bench leaderboard certification.
The benchmark has four prompt-level task conditions, not only B1 and B2: B1, B2, B3, and B4. B1 provides the complete research method and procedure. B2 provides the intended method and constraints but removes the full procedure. B3 leaves method selection to the agent. B4 adds factually correct but non-essential context to the B3 condition. This article reports B1 and B2 only; it does not extend these results to B3 or B4.
Key Takeaways
- B1: 79.35, the highest numerical score among the B1 records listed here. B1 tests whether an agent can turn a complete research plan into a running program and deliver the results.
- B2: 56.11, the highest numerical score among the B2 records listed here. B2 keeps the method and constraints but asks the agent to fill in the steps, order the methods, and organize the workflow.
- B1–B4 are four task conditions. The B1/B2 results do not imply that AIPOCH has already evaluated B3/B4 or ranks first across all four levels.
- The scores come from three local repetitions. The B1 mean covers 59 valid scored tasks; B2 is reported as 56.11 in the internal result summary.
For a benchmark run to become useful in real research, the score needs context: inputs, file versions, execution records, and verification evidence. AIPOCH Open-Science’s .science research package can carry selected conversation branches, file versions, Notebook records, and verification evidence for review, handoff, and archiving in another project or computer. Imported sessions are read-only; they do not execute code or restore credentials. See the official research-package documentation.
What does ASI-Bench measure?
ASI-Bench is an AI for Science evaluation built around project-level research tasks. Its official materials describe 60 long-horizon tasks across scientific areas. An agent must understand the research objective, process input data, write and run programs, and deliver inspectable code and research results. The benchmark evaluates whether a research workflow can be completed, rather than whether an agent can answer one isolated knowledge question.
The official task design defines four prompt levels with progressively less method guidance. They can also be understood as four task-set conditions:
| Level | Information provided | Main capability tested |
|---|---|---|
| B1 | Scientific background, intended method, equations, and the full procedure | Turn the given plan into a working program, execute it, and deliver the result |
| B2 | Intended method and relevant constraints, without the full procedure | Fill in implementation details, order the methods, and organize a complete workflow |
| B3 | Objective, data, constraints, and required outputs, without a prescribed method | Select a method and independently construct and validate the workflow |
| B4 | The B3 condition plus factually correct but non-essential information | Identify the important constraints and complete the research while resisting distraction |
These are not four unrelated product leaderboards. By changing how much method information appears in the prompt, ASI-Bench observes the transition from executing a specified plan to organizing or selecting a plan. A first-place result for B1 and B2 therefore describes the two evaluated conditions; it does not cover B3 or B4.
B1: turning a complete research plan into working analysis
What B1 asks the agent to do
B1 gives the scientific background, intended method, equations, and full procedure. The agent must read the input data, turn the plan into executable code, run the analysis, inspect intermediate results, and deliver code, outputs, and method notes. The central question is: when the research route is already specified, can the system reliably carry it through?
In an ASI-Bench single-cell RNA-sequencing task, for example, the agent works with simulated cell-by-gene expression data. It must group cells, infer the order of cell-state changes, identify candidate marker genes, and estimate gene-regulatory relationships. The deliverables include runnable code, per-cell cluster and trajectory values, a candidate-gene list, a regulatory matrix, and an explanation of the method. The procedure reduces method-design uncertainty, but it does not perform data processing, code execution, result checking, or delivery for the agent.
AIPOCH’s B1 setup and result
AIPOCH Open-Science used GPT-6 (gpt-6-astra) / xhigh, with domain-specialist roles and research Skills, and completed three local repetitions. One of the standard 60 tasks was excluded from the mean because of an official scorer anomaly, so the reported B1 mean covers 59 valid scored tasks. The reported task-quality mean is 79.35. The source summary does not list the three individual run scores or explain how results were combined across repetitions.
The 79.35 is a quality mean across task-level scores; it does not mean that 79.35% of the tasks were simply correct. The configuration matched 11 specialist categories to tasks across astronomy, bioinformatics, statistics, chemistry, robotics, physics simulation, and other areas. The result reflects the combined AIPOCH agent configuration, rather than GPT-6’s standalone performance on every scientific task.
Among the 24 B1 records listed below, AIPOCH’s 79.35 is 4.57 points above the local Codex reference at 74.78, which uses the same model and named reasoning level. It is 7.06 points above the highest public reference listed here, Claude Code × Claude Opus 5 at 72.29. These gaps describe this comparison; they do not replace an official review under fully controlled conditions.
Full B1 comparison
The table preserves all 24 B1 records from the internal summary, ordered by score. Local repetitions and public references come from different sources; the order is the numerical ordering used in this article, not an official leaderboard rank.
| Numerical order in this comparison | Agent × model | Reasoning | B1 score (100 max) | Result type |
|---|---|---|---|---|
| 1 | AIPOCH Open-Science × GPT-6 (gpt-6-astra) | xhigh | 79.35 | Local evaluation (3 repetitions) |
| 2 | Codex × GPT-6 (gpt-6-astra) | xhigh | 74.78 | Local evaluation (3 repetitions) |
| 3 | Claude Code × Claude Opus 5 | Max | 72.29 | Public reference |
| 4 | Codex × GPT-5.6 Sol | ultra | 71.78 | Public reference |
| 5 | AIPOCH Open-Science × DeepSeek v4-flash | max | 69.64 | Local evaluation (3 repetitions) |
| 6 | Claude Code × GLM-5.3 | Max | 63.05 | Public reference |
| 7 | Codex × GPT-5.6 Sol | xhigh | 62.75 | Public reference |
| 8 | Claude Code × Gemini 3.8 Flash | Max | 61.33 | Public reference |
| 9 | Claude Code × Qwen3.8-Max | Max | 59.14 | Public reference |
| 10 | Codex × GPT-5.5 | xhigh | 57.57 | Public reference |
| 11 | Kimi Code × Kimi K3 | Not disclosed | 56.16 | Public reference |
| 12 | Claude Code × Kimi K3 | Max | 55.79 | Public reference |
| 13 | Claude Code × GLM-5.2 | Max | 54.01 | Public reference |
| 14 | Claude Code × Claude Opus 4.8 | Max | 52.48 | Public reference |
| 15 | Claude Code × DeepSeek V4 Flash 0731 | Max | 51.85 | Public reference |
| 16 | OpenHands × DeepSeek V4 Flash 0731 | Not disclosed | 50.60 | Public reference |
| 17 | Claude Code × GLM-5.3-Flash | Max | 49.27 | Public reference |
| 18 | Claude Code × Kimi K2.7 | Max | 44.43 | Public reference |
| 19 | Claude Code × MiniMax M3 | Max | 43.53 | Public reference |
| 20 | Claude Code × DeepSeek V4 Pro | Max | 42.97 | Public reference |
| 21 | Claude Code × MiMo V2.5 Pro | Max | 41.70 | Public reference |
| 22 | OpenHands × DeepSeek V4 Pro | Not disclosed | 37.78 | Public reference |
| 23 | Kimi Code × Kimi K2.7 | Not disclosed | 29.99 | Public reference |
| 24 | MiMo Code × MiMo V2.5 Pro | Not disclosed | 27.73 | Public reference |
B2: organizing a complete workflow from a specified method
How B2 differs from B1
B2 belongs to the same ASI-Bench task system. The research objective, input data, deliverables, and scoring contract correspond to B1, but the prompt provides the intended method and constraints without a full procedure. The agent must decide what to do first, how to connect the methods, how to fill in missing implementation details, and when the result is ready for inspection.
B1 therefore emphasizes “can the agent execute the complete plan?”, while B2 asks “can the agent organize the specified methods into a complete workflow?” B2 is still not the condition in which the agent chooses the research method from scratch; that is B3. B4 adds the separate challenge of filtering non-essential but factually correct context.
Using the same single-cell example, B1 can provide a step-by-step list from preprocessing and dimensionality reduction through clustering, trajectory analysis, and network inference. B2 still specifies the methods but asks the agent to arrange and connect them. The deliverables remain runnable code and research outputs, and the score focuses on the quality of the resulting work rather than only the completeness of a written plan.
AIPOCH’s B2 setup and result
AIPOCH Open-Science kept the same domain-specialist and Skills configuration used for B1, ran GPT-6 (gpt-6-astra) / xhigh, and completed three local repetitions. The reported B2 score is 56.11.
Across the 23 B2 records listed below, AIPOCH’s 56.11 is above the same-model, same-reasoning Codex local reference at 56.09 and above the highest public reference listed here, Codex × GPT-5.6 Sol / ultra at 49.57. The AIPOCH–Codex gap is only 0.02 points, so the useful conclusion is narrower: in this local-and-public result summary, AIPOCH’s B2 result is first by score and shows continued workflow progress after the full procedure is removed. It is not a statistical significance test and does not establish stable superiority under every task and run condition.
Full B2 comparison
The table preserves all 23 B2 records from the internal summary, ordered by score. The records come from different sources and may use different runtime conditions; this numerical order is a comparison in this article, not an official leaderboard rank. Each result should be read with its result type and evaluation boundary.
| Numerical order in this comparison | Agent × model | Reasoning | B2 score (100 max) | Result type |
|---|---|---|---|---|
| 1 | AIPOCH Open-Science × GPT-6 (gpt-6-astra) | xhigh | 56.11 | Local evaluation (3 repetitions) |
| 2 | Codex × GPT-6 (gpt-6-astra) | xhigh | 56.09 | Local evaluation (3 repetitions) |
| 3 | Codex × GPT-5.6 Sol | ultra | 49.57 | Public reference |
| 4 | Claude Code × Claude Opus 5 | Max | 45.80 | Public reference |
| 5 | Codex × GPT-5.6 Sol | xhigh | 42.96 | Public reference |
| 6 | Claude Code × Gemini 3.8 Flash | Max | 39.13 | Public reference |
| 7 | Claude Code × Claude Opus 4.8 | Max | 36.25 | Public reference |
| 8 | Claude Code × GLM-5.3 | Max | 35.45 | Public reference |
| 9 | Codex × GPT-5.5 | xhigh | 35.28 | Public reference |
| 10 | Claude Code × Kimi K3 | Max | 35.24 | Public reference |
| 11 | Kimi Code × Kimi K3 | Not disclosed | 34.05 | Public reference |
| 12 | Claude Code × Qwen3.8-Max | Max | 32.51 | Public reference |
| 13 | Claude Code × GLM-5.2 | Max | 30.81 | Public reference |
| 14 | Claude Code × GLM-5.3-Flash | Max | 28.66 | Public reference |
| 15 | Claude Code × DeepSeek V4 Flash 0731 | Max | 26.49 | Public reference |
| 16 | OpenHands × DeepSeek V4 Flash 0731 | Not disclosed | 24.89 | Public reference |
| 17 | Claude Code × Kimi K2.7 | Max | 23.84 | Public reference |
| 18 | Claude Code × MiniMax M3 | Max | 20.86 | Public reference |
| 19 | Claude Code × DeepSeek V4 Pro | Max | 18.79 | Public reference |
| 20 | Claude Code × MiMo V2.5 Pro | Max | 18.49 | Public reference |
| 21 | Kimi Code × Kimi K2.7 | Not disclosed | 17.01 | Public reference |
| 22 | OpenHands × DeepSeek V4 Pro | Not disclosed | 16.66 | Public reference |
| 23 | MiMo Code × MiMo V2.5 Pro | Not disclosed | 11.33 | Public reference |
How should the B1 and B2 results be read?
The two scores provide two distinct signals:
| Level | AIPOCH local score | Position in this comparison | Capability signal |
|---|---|---|---|
| B1 | 79.35 | First among the 24 listed records | Execute a complete plan, run the analysis, and deliver the outputs |
| B2 | 56.11 | First among the 23 listed records | Fill in missing steps, connect methods, and organize the workflow under constraints |
The difference is not that one level represents “doing science” and the other does not. The difference is how much method information the prompt provides and how much workflow organization the agent must supply. B1 is closer to a researcher who has already fixed the method and needs reliable execution. B2 is closer to a researcher who gives the objective, method boundary, and constraints while asking the system to connect the intermediate steps.
That distinction is where AIPOCH Open-Science’s product value becomes practical. Projects, papers, files, Notebooks, code execution, and generated outputs can stay in one research context while models and Skills are configured for the task. The AIPOCH Open-Science scientific-agent workflow describes how specialist roles can support this kind of work.
When the run needs review or handoff, a .science research package can carry selected conversation branches, file versions, Notebook records, and verification evidence. It is designed to carry research context across projects or computers; imported sessions are read-only, do not execute code or restore credentials, exclude Side Chat conversations and private reading bookmarks, and include only the files selected during export. A received verification record describes the sender’s checks; it does not prove that the receiving computer reran the experiment. See the .science research-package documentation.
Local evaluation, seeds, and the official leaderboard boundary
Three local repetitions are not an official certification
This article reports three local repetitions of AIPOCH Open-Science using public task materials and a local runtime. B1 79.35 and B2 56.11 are local evaluation results. We place them beside collected public references to show the reported score range, while keeping the wording separate from official leaderboard certification.
seed31415 and seed42 are different concrete instances
ASI-Bench now documents two explicit scoring contracts: seed31415 publishes reference answers and can be scored locally with the repository scorers; seed42 keeps reference answers private and requires official service scoring. The two seeds share the task framework and B1–B4 structure, but their concrete data, parameters, and reference answers can differ. They should not be assumed to be the same test instances.
For that reason, two records carrying the same B1 or B2 label must still be read with their seed, scorer, model version, reasoning level, agent harness, and runtime policy. Public records are useful for showing a comparison range, but they do not form a controlled experiment by default.
Next step: evaluate B3 and B4
This article reports B1 and B2 only. B3 asks the agent to choose a research method without one being prescribed. B4 adds factually correct but non-essential context and tests whether the system can identify the constraints that matter while resisting distraction.
Any future B3/B4 report should preserve the same repetition count, task coverage, seed, scoring method, and configuration details so that the results form an interpretable longitudinal comparison with B1/B2. This article makes no inference about B3 or B4 scores.
Frequently Asked Questions
Does ASI-Bench have only B1 and B2?
No. ASI-Bench has four prompt-level task conditions: B1, B2, B3, and B4. B1 provides the full method and procedure; B2 keeps the method but removes the full procedure; B3 asks the agent to select the method; B4 adds non-essential but factually correct context. This article reports B1 and B2.
Has AIPOCH Open-Science won the official ASI-Bench overall leaderboard?
That is not the claim here. “First” means first by score in the comparison made from AIPOCH’s local repetitions and the collected public references as of the source test date. The official leaderboard involves formal submission, official scoring, and consistent run conditions.
Does a B1 score of 79.35 mean that 79.35% of tasks were completed correctly?
No. B1 is a 100-point mean after task-level result quality is scored. The reported mean covers 59 valid scored tasks; it is not a simple fraction of tasks answered correctly. The reported B2 score is 56.11. The source summary does not specify its valid-task count, list the three individual run scores, or explain how results were combined across repetitions.
What is the core difference between B1 and B2?
B1 gives the complete research plan and procedure, so the central challenge is reliable execution. B2 gives the intended method and constraints but removes the full procedure, so the central challenge is filling in steps, ordering methods, and organizing a runnable, inspectable workflow.
Does a .science package automatically rerun an experiment on another computer?
No. Imported sessions are read-only; .science does not execute code, restore credentials, or automatically validate the receiving environment. It carries selected research records and verification evidence for review, handoff, and archiving.
Data and sources
- Scores and full B1/B2 comparison tables: AIPOCH internal summary “AIPOCH Open-Science benchmark results — 2026-09-29”; B1 and B2 each completed three local repetitions.
- Official ASI-Bench project documentation: GitHub README.
- Official run guidance: Getting Started.
- Official scoring explanation: How scoring works.
- Official results page: ASI-Bench Leaderboard.
- Product and research workflow: AIPOCH Open-Science.
.sciencecapability boundary: AIPOCH Open-Science research packages.
Disclaimer
This article reports AIPOCH Open-Science’s local repeated evaluations under ASI-Bench B1 and B2 conditions, together with a numerical comparison against the listed public references. It is not an official ASI-Bench leaderboard certification and does not replace official scoring, scientific review, or independent verification of any research conclusion. Check the task level, seed, model, agent harness, scorer, input files, and outputs before drawing operational conclusions.