Back to Blog
5 min read

AIPOCH Open-Science Hits 10/10 on AstaBench LitQA2-FullText-Search

AIPOCH Open-Science completed three local repetitions on the official AstaBench LitQA2-FullText-Search validation split and retrieved all 10 target papers. Read the metric, setup, full 30-record comparison, and limits.

AIPOCH

AIPOCH Open-Science LitQA2-FullText-Search cover showing 10/10 target-paper hits, 100% recall@30, and the top ten of thirty comparison records

AIPOCH Open-Science completed three local repetitions on the official AstaBench LitQA2-FullText-Search validation split. All 10 retrieval questions returned the target paper within the top 30 results, producing recall@30 = 100%. That is the highest value in the 30 public and local records listed in this article, so it ranks first in this numerical comparison; it is not an official AstaBench leaderboard certification.

The result answers one focused question: when a scientific agent receives a research question, can it rank the paper containing the answer among its first 30 retrieved papers? It does not directly show that the agent answered the science question correctly, understood the full paper, wrote a literature review, or reproduced an experiment.

Key Takeaways

  • 10/10 target-paper hits and 100% recall@30: AIPOCH Open-Science retrieved the target paper for every question in the official 10-question validation split.
  • Three local repetitions: The run used DeepSeek v4-flash / max, six AI specialist roles, and 15 Skills.
  • Thirty comparison records: The table contains one AIPOCH local validation result, three public validation records, and 26 public test records. Validation and test are different question sets.
  • “First” means first in this comparison: Models, tools, retrieval systems, and runtime conditions differ, so the table should not be presented as an official leaderboard result.

What does LitQA2-FullText-Search measure?

LitQA2-FullText-Search is AstaBench’s retrieval-focused version of the LitQA2-FullText task. The official AstaBench description says that it keeps the literature-retrieval problem but isolates the retrieval step: instead of being scored on the final answer, an agent returns a ranked list of papers likely to contain the answer, and the evaluation checks whether the target paper appears in the top 30 (AstaBench benchmark page; AstaBench task list).

LitQA2-FullText itself concerns scientific questions that require retrieving a paper and reading its full text. The -Search variant makes the first stage explicit: can the system find the right evidence source before it tries to interpret the evidence?

That separation matters. An agent can find the correct paper and still misread a table, method, or result. It can also answer a question in some setting without reliably ranking the correct paper among its first results. LitQA2-FullText-Search measures the retrieval gate alone.

Which questions did this evaluation use?

The AIPOCH Open-Science result uses the official 10-question validation split. Each question maps to a target paper that contains the answer. The agent searches the literature index and returns candidate papers. The validation and test splits belong to the same task family but contain different questions.

SplitSizeRelationshipUse in this article
validation10 questionsDoes not overlap with testAIPOCH’s three local repetitions; three public validation records
test75 questionsDoes not overlap with validationTwenty-six public test records

AstaBench’s public result records include both validation and test entries. This article places them beside the AIPOCH local validation result to show the numerical range and configuration differences. It does not merge the 10 validation questions and 75 test questions into one test set.

How is 100% recall@30 calculated?

Recall@30 checks whether the target paper appears in the first 30 papers returned for each question. A hit receives 1; a miss receives 0; the scores are averaged across the questions. AIPOCH Open-Science hit all 10 validation targets, so 10 / 10 = 1.000000 = 100%.

The metric does not answer these questions:

  • whether the first result is the best paper to read;
  • whether the agent understood the target paper’s full text;
  • whether the agent extracted the correct experimental result;
  • whether the agent completed a scientific answer, review, or reproduction.

The precise interpretation is therefore: AIPOCH Open-Science achieved perfect target-paper recall on this 10-question validation split. The score covers the evidence-finding entry point of a research workflow.

How did AIPOCH Open-Science run the local evaluation?

AIPOCH Open-Science used DeepSeek v4-flash / max and completed three local repetitions; the reported result was a target-paper hit for all 10 validation questions. The run was a configured agent system rather than a model-only call: it combined the underlying model with specialist roles and research Skills.

The LitQA2 configuration used six specialist roles for:

  1. Semantic retrieval: translating the research question into searchable concepts and relations.
  2. Entity and evidence location: identifying the research object and the paper clues most likely to contain the answer.
  3. Evidence review: checking whether candidate papers match the question’s evidence need.
  4. Citation auditing: keeping the source traceable for later research work.
  5. Data analysis: preserving an entry point for tables, values, and comparisons found later.
  6. Conclusion review: checking that the retrieval conclusion stays within the evidence.

The system also enabled 15 Skills that supported retrieval, evidence location, citation checks, and research outputs. The score therefore describes the combined Agent configuration, not an unconditional score for DeepSeek v4-flash under every retrieval setup.

For a retrieval result to become part of a research project, the context must also be reviewable and transferable. AIPOCH Open-Science’s .science research package can carry selected conversation branches, file versions, Notebook records, and verification evidence for review, handoff, and archiving in another project or computer. Imported sessions are read-only; they do not execute code or restore credentials. See the official research-package documentation.

For the next step after retrieval, see AIPOCH’s literature-screening and evidence workflow and scientific-agent specialist workflow.

Full 30-record comparison

The table below preserves all 30 records in the source summary and sorts them by recall@30. Row 1 is AIPOCH Open-Science’s local validation result. Rows 2–30 are public records from the AstaBench result set, including three validation entries and 26 test entries.

How to read it: validation and test contain different questions; tied values do not imply a meaningful ordering; the row number is this article’s numerical sort, not an official AstaBench leaderboard position.

Comparison rankAgent / recordModel fieldrecall@30Split
1AIPOCH Open-Sciencedeepseek-v4-flash1.000000Local validation (three repetitions)
2Asta Paper Finderopenai/gpt-4o-mini0.906667Public test
3Asta v0mockllm/model0.906667Public test
4ReActanthropic/claude-sonnet-4-60.866667Public test
5ReActanthropic/claude-opus-4-60.866667Public test
6ReActanthropic/claude-opus-4-70.853333Public test
7ReActopenai/gpt-5.5-2026-04-230.853333Public test
8ReActopenai/gpt-5-2025-08-070.826667Public test
9ReActopenai/gpt-5.4-2026-03-050.800000Public test
10Asta Paper Finderopenai/gpt-4o-mini0.800000Public validation
11ReActgoogle/gemini-3.1-pro-preview0.773333Public test
12ReActopenai/gpt-4o-2024-08-060.666667Public test
13ReActopenai/gpt-4.1-2025-04-140.653333Public test
14ReActanthropic/claude-3-5-haiku-202410220.600000Public test
15ReActgoogle/gemini-2.5-flash-preview-05-200.573333Public test
16ReActopenai/o3-2025-04-160.573333Public test
17ReActopenai/gpt-5-mini-2025-08-070.560000Public test
18Smolagents Coderopenai/gpt-5-2025-08-070.546667Public test
19Smolagents Coderanthropic/claude-sonnet-4-202505140.520000Public test
20Smolagents Coderopenai/gpt-4.1-2025-04-140.506667Public test
21You.com Search APIopenai/gpt-4o-mini0.500000Public validation
22Smolagents Coderopenai/gpt-5-mini-2025-08-070.480000Public test
23ReActanthropic/claude-sonnet-4-202505140.466667Public test
24ReActtogether/meta-llama/Llama-4-Scout-17B-16E-Instruct0.373333Public test
25You.com Search APIopenai/gpt-4o-mini0.360000Public test
26Smolagents Codergoogle/gemini-2.5-flash-preview-05-200.360000Public test
27Smolagents Coderopenai/gpt-4o-2024-08-060.200000Public test
28ReActanthropic/claude-3-5-haiku-202410220.100000Public validation
29Smolagents Codertogether/meta-llama/Llama-4-Scout-17B-16E-Instruct0.066667Public test
30Smolagents Coderanthropic/claude-3-5-haiku-202410220.026667Public test

What is the fairest same-split comparison?

AIPOCH Open-Science’s 1.000000, Asta Paper Finder’s 0.800000, and You.com Search API’s 0.500000 are the three public or local records in the table that use the validation split. The other public test records use 75 different questions. They show the broader public range but are not a same-question head-to-head comparison with the 10-question local validation run.

How large is the reported numerical lead?

AIPOCH’s 1.000000 is 0.093333, or 9.33 percentage points, above the highest public test value of 0.906667. Against the highest public validation value of 0.800000, the difference is 0.200000, or 20 percentage points. The first comparison crosses validation and test splits; the second stays within the validation split.

What does this result mean for a research workflow?

Literature retrieval is the entry point of a scientific-agent workflow, not the whole research process. The 100% result means that, for these 10 validation requests, AIPOCH Open-Science placed the target paper in the first 30 candidates. Researchers still need to read the paper, verify evidence, run analysis, and review the conclusion.

AIPOCH Open-Science is designed as a persistent research workbench: projects, literature, files, Notebooks, code execution, and generated artifacts can stay in one research context, while models and Skills can be configured for the task. The AIPOCH Open-Science product page and research-workbench overview explain the broader workflow.

The .science package boundary is part of that story. It can carry selected research context and verification evidence for inspection and handoff, but it is not a full-session replay, a credential transfer, or an automatic scientific validation system. A received verification record describes the sender’s checks; it does not prove that the receiving computer reran them.

Limits and next steps

Four boundaries should remain attached to this result:

  1. The local result is not an official certification. AIPOCH ran the official task and validation split locally, but official leaderboard inclusion is a separate process.
  2. Validation and test are different question sets. The 30 records combine 10-question validation and 75-question test results, and the split label must stay visible.
  3. Runtime conditions differ. Records use different models, agents, retrieval tools, and environments; this is a transparent comparison, not a controlled experiment.
  4. Retrieval hit is not evidence correctness. A target paper in the top 30 only identifies a plausible evidence source. The paper, citations, methods, and conclusions still require review.

The next useful evaluation would report test-split results under a fully documented retrieval stack and preserve each run’s configuration, evidence trail, and verification record in a reviewable research project.

Frequently Asked Questions

What is the difference between LitQA2-FullText-Search and LitQA2-FullText?

LitQA2-FullText asks questions that require retrieving and reading scientific papers. LitQA2-FullText-Search isolates retrieval: the agent returns a ranked list of papers likely to contain the answer, and the score checks whether the target paper appears in the first 30 results.

Is AIPOCH Open-Science’s 100% an official AstaBench first place?

No. It is AIPOCH Open-Science’s result from three local repetitions on the official 10-question validation split. It is the highest value in the 30 public and local records assembled for this article, so it ranks first in this comparison; it is not an official AstaBench leaderboard certification.

What does the “30” in recall@30 mean?

The evaluator checks the first 30 papers returned for each question. A target paper in those 30 counts as a hit. All 10 validation hits produce recall@30 of 100%.

Why can’t 100% and 90.67% be treated as the same exam?

AIPOCH’s 100% uses 10 validation questions, while the highest public 90.67% uses 75 test questions. They are two splits of the same task family, so the values are useful for context but are not identical-question results.

Does this prove that a literature review is correct?

No. It shows target-paper retrieval for this validation set. A literature review still requires checking the original paper, full-text evidence, citations, methods, and conclusions.

Data and sources

Disclaimer

This article reports a local AIPOCH Open-Science evaluation using an official task and validation split; it is not an official AstaBench leaderboard certification. Recall@30 measures whether the target paper appears in the first 30 retrieved results, not whether the scientific answer, full-text evidence, or research conclusion is correct. Review the original papers, citations, methods, and generated outputs before using them in research.