
AIPOCH Open-Science completed three local repetitions on the official AstaBench LitQA2-FullText-Search validation split. All 10 retrieval questions returned the target paper within the top 30 results, producing recall@30 = 100%. That is the highest value in the 30 public and local records listed in this article, so it ranks first in this numerical comparison; it is not an official AstaBench leaderboard certification.
The result answers one focused question: when a scientific agent receives a research question, can it rank the paper containing the answer among its first 30 retrieved papers? It does not directly show that the agent answered the science question correctly, understood the full paper, wrote a literature review, or reproduced an experiment.
Key Takeaways
- 10/10 target-paper hits and 100% recall@30: AIPOCH Open-Science retrieved the target paper for every question in the official 10-question validation split.
- Three local repetitions: The run used DeepSeek v4-flash / max, six AI specialist roles, and 15 Skills.
- Thirty comparison records: The table contains one AIPOCH local validation result, three public validation records, and 26 public test records. Validation and test are different question sets.
- “First” means first in this comparison: Models, tools, retrieval systems, and runtime conditions differ, so the table should not be presented as an official leaderboard result.
What does LitQA2-FullText-Search measure?
LitQA2-FullText-Search is AstaBench’s retrieval-focused version of the LitQA2-FullText task. The official AstaBench description says that it keeps the literature-retrieval problem but isolates the retrieval step: instead of being scored on the final answer, an agent returns a ranked list of papers likely to contain the answer, and the evaluation checks whether the target paper appears in the top 30 (AstaBench benchmark page; AstaBench task list).
LitQA2-FullText itself concerns scientific questions that require retrieving a paper and reading its full text. The -Search variant makes the first stage explicit: can the system find the right evidence source before it tries to interpret the evidence?
That separation matters. An agent can find the correct paper and still misread a table, method, or result. It can also answer a question in some setting without reliably ranking the correct paper among its first results. LitQA2-FullText-Search measures the retrieval gate alone.
Which questions did this evaluation use?
The AIPOCH Open-Science result uses the official 10-question validation split. Each question maps to a target paper that contains the answer. The agent searches the literature index and returns candidate papers. The validation and test splits belong to the same task family but contain different questions.
| Split | Size | Relationship | Use in this article |
|---|---|---|---|
validation | 10 questions | Does not overlap with test | AIPOCH’s three local repetitions; three public validation records |
test | 75 questions | Does not overlap with validation | Twenty-six public test records |
AstaBench’s public result records include both validation and test entries. This article places them beside the AIPOCH local validation result to show the numerical range and configuration differences. It does not merge the 10 validation questions and 75 test questions into one test set.
How is 100% recall@30 calculated?
Recall@30 checks whether the target paper appears in the first 30 papers returned for each question. A hit receives 1; a miss receives 0; the scores are averaged across the questions. AIPOCH Open-Science hit all 10 validation targets, so 10 / 10 = 1.000000 = 100%.
The metric does not answer these questions:
- whether the first result is the best paper to read;
- whether the agent understood the target paper’s full text;
- whether the agent extracted the correct experimental result;
- whether the agent completed a scientific answer, review, or reproduction.
The precise interpretation is therefore: AIPOCH Open-Science achieved perfect target-paper recall on this 10-question validation split. The score covers the evidence-finding entry point of a research workflow.
How did AIPOCH Open-Science run the local evaluation?
AIPOCH Open-Science used DeepSeek v4-flash / max and completed three local repetitions; the reported result was a target-paper hit for all 10 validation questions. The run was a configured agent system rather than a model-only call: it combined the underlying model with specialist roles and research Skills.
The LitQA2 configuration used six specialist roles for:
- Semantic retrieval: translating the research question into searchable concepts and relations.
- Entity and evidence location: identifying the research object and the paper clues most likely to contain the answer.
- Evidence review: checking whether candidate papers match the question’s evidence need.
- Citation auditing: keeping the source traceable for later research work.
- Data analysis: preserving an entry point for tables, values, and comparisons found later.
- Conclusion review: checking that the retrieval conclusion stays within the evidence.
The system also enabled 15 Skills that supported retrieval, evidence location, citation checks, and research outputs. The score therefore describes the combined Agent configuration, not an unconditional score for DeepSeek v4-flash under every retrieval setup.
For a retrieval result to become part of a research project, the context must also be reviewable and transferable. AIPOCH Open-Science’s .science research package can carry selected conversation branches, file versions, Notebook records, and verification evidence for review, handoff, and archiving in another project or computer. Imported sessions are read-only; they do not execute code or restore credentials. See the official research-package documentation.
For the next step after retrieval, see AIPOCH’s literature-screening and evidence workflow and scientific-agent specialist workflow.
Full 30-record comparison
The table below preserves all 30 records in the source summary and sorts them by recall@30. Row 1 is AIPOCH Open-Science’s local validation result. Rows 2–30 are public records from the AstaBench result set, including three validation entries and 26 test entries.
How to read it: validation and test contain different questions; tied values do not imply a meaningful ordering; the row number is this article’s numerical sort, not an official AstaBench leaderboard position.
| Comparison rank | Agent / record | Model field | recall@30 | Split |
|---|---|---|---|---|
| 1 | AIPOCH Open-Science | deepseek-v4-flash | 1.000000 | Local validation (three repetitions) |
| 2 | Asta Paper Finder | openai/gpt-4o-mini | 0.906667 | Public test |
| 3 | Asta v0 | mockllm/model | 0.906667 | Public test |
| 4 | ReAct | anthropic/claude-sonnet-4-6 | 0.866667 | Public test |
| 5 | ReAct | anthropic/claude-opus-4-6 | 0.866667 | Public test |
| 6 | ReAct | anthropic/claude-opus-4-7 | 0.853333 | Public test |
| 7 | ReAct | openai/gpt-5.5-2026-04-23 | 0.853333 | Public test |
| 8 | ReAct | openai/gpt-5-2025-08-07 | 0.826667 | Public test |
| 9 | ReAct | openai/gpt-5.4-2026-03-05 | 0.800000 | Public test |
| 10 | Asta Paper Finder | openai/gpt-4o-mini | 0.800000 | Public validation |
| 11 | ReAct | google/gemini-3.1-pro-preview | 0.773333 | Public test |
| 12 | ReAct | openai/gpt-4o-2024-08-06 | 0.666667 | Public test |
| 13 | ReAct | openai/gpt-4.1-2025-04-14 | 0.653333 | Public test |
| 14 | ReAct | anthropic/claude-3-5-haiku-20241022 | 0.600000 | Public test |
| 15 | ReAct | google/gemini-2.5-flash-preview-05-20 | 0.573333 | Public test |
| 16 | ReAct | openai/o3-2025-04-16 | 0.573333 | Public test |
| 17 | ReAct | openai/gpt-5-mini-2025-08-07 | 0.560000 | Public test |
| 18 | Smolagents Coder | openai/gpt-5-2025-08-07 | 0.546667 | Public test |
| 19 | Smolagents Coder | anthropic/claude-sonnet-4-20250514 | 0.520000 | Public test |
| 20 | Smolagents Coder | openai/gpt-4.1-2025-04-14 | 0.506667 | Public test |
| 21 | You.com Search API | openai/gpt-4o-mini | 0.500000 | Public validation |
| 22 | Smolagents Coder | openai/gpt-5-mini-2025-08-07 | 0.480000 | Public test |
| 23 | ReAct | anthropic/claude-sonnet-4-20250514 | 0.466667 | Public test |
| 24 | ReAct | together/meta-llama/Llama-4-Scout-17B-16E-Instruct | 0.373333 | Public test |
| 25 | You.com Search API | openai/gpt-4o-mini | 0.360000 | Public test |
| 26 | Smolagents Coder | google/gemini-2.5-flash-preview-05-20 | 0.360000 | Public test |
| 27 | Smolagents Coder | openai/gpt-4o-2024-08-06 | 0.200000 | Public test |
| 28 | ReAct | anthropic/claude-3-5-haiku-20241022 | 0.100000 | Public validation |
| 29 | Smolagents Coder | together/meta-llama/Llama-4-Scout-17B-16E-Instruct | 0.066667 | Public test |
| 30 | Smolagents Coder | anthropic/claude-3-5-haiku-20241022 | 0.026667 | Public test |
What is the fairest same-split comparison?
AIPOCH Open-Science’s 1.000000, Asta Paper Finder’s 0.800000, and You.com Search API’s 0.500000 are the three public or local records in the table that use the validation split. The other public test records use 75 different questions. They show the broader public range but are not a same-question head-to-head comparison with the 10-question local validation run.
How large is the reported numerical lead?
AIPOCH’s 1.000000 is 0.093333, or 9.33 percentage points, above the highest public test value of 0.906667. Against the highest public validation value of 0.800000, the difference is 0.200000, or 20 percentage points. The first comparison crosses validation and test splits; the second stays within the validation split.
What does this result mean for a research workflow?
Literature retrieval is the entry point of a scientific-agent workflow, not the whole research process. The 100% result means that, for these 10 validation requests, AIPOCH Open-Science placed the target paper in the first 30 candidates. Researchers still need to read the paper, verify evidence, run analysis, and review the conclusion.
AIPOCH Open-Science is designed as a persistent research workbench: projects, literature, files, Notebooks, code execution, and generated artifacts can stay in one research context, while models and Skills can be configured for the task. The AIPOCH Open-Science product page and research-workbench overview explain the broader workflow.
The .science package boundary is part of that story. It can carry selected research context and verification evidence for inspection and handoff, but it is not a full-session replay, a credential transfer, or an automatic scientific validation system. A received verification record describes the sender’s checks; it does not prove that the receiving computer reran them.
Limits and next steps
Four boundaries should remain attached to this result:
- The local result is not an official certification. AIPOCH ran the official task and validation split locally, but official leaderboard inclusion is a separate process.
- Validation and test are different question sets. The 30 records combine 10-question validation and 75-question test results, and the split label must stay visible.
- Runtime conditions differ. Records use different models, agents, retrieval tools, and environments; this is a transparent comparison, not a controlled experiment.
- Retrieval hit is not evidence correctness. A target paper in the top 30 only identifies a plausible evidence source. The paper, citations, methods, and conclusions still require review.
The next useful evaluation would report test-split results under a fully documented retrieval stack and preserve each run’s configuration, evidence trail, and verification record in a reviewable research project.
Frequently Asked Questions
What is the difference between LitQA2-FullText-Search and LitQA2-FullText?
LitQA2-FullText asks questions that require retrieving and reading scientific papers. LitQA2-FullText-Search isolates retrieval: the agent returns a ranked list of papers likely to contain the answer, and the score checks whether the target paper appears in the first 30 results.
Is AIPOCH Open-Science’s 100% an official AstaBench first place?
No. It is AIPOCH Open-Science’s result from three local repetitions on the official 10-question validation split. It is the highest value in the 30 public and local records assembled for this article, so it ranks first in this comparison; it is not an official AstaBench leaderboard certification.
What does the “30” in recall@30 mean?
The evaluator checks the first 30 papers returned for each question. A target paper in those 30 counts as a hit. All 10 validation hits produce recall@30 of 100%.
Why can’t 100% and 90.67% be treated as the same exam?
AIPOCH’s 100% uses 10 validation questions, while the highest public 90.67% uses 75 test questions. They are two splits of the same task family, so the values are useful for context but are not identical-question results.
Does this prove that a literature review is correct?
No. It shows target-paper retrieval for this validation set. A literature review still requires checking the original paper, full-text evidence, citations, methods, and conclusions.
Data and sources
- Result data and the 30-record table: AIPOCH internal summary “AIPOCH Open-Science benchmark results — 2026-09-29”; AIPOCH LitQA2 validation run completed three local repetitions.
- Official benchmark description: AstaBench: Benchmarking AI Agents for Science.
- Official task list: allenai/asta-bench README.
- Product workflow: AIPOCH Open-Science.
- Research-package boundaries: AIPOCH Open-Science research packages.
Disclaimer
This article reports a local AIPOCH Open-Science evaluation using an official task and validation split; it is not an official AstaBench leaderboard certification. Recall@30 measures whether the target paper appears in the first 30 retrieved results, not whether the scientific answer, full-text evidence, or research conclusion is correct. Review the original papers, citations, methods, and generated outputs before using them in research.