Skip to content

feat: add Serply as optional Google Scholar search backend - #67

Open
googio wants to merge 1 commit into
RUC-NLPIR:mainfrom
googio:feat/serply-search
Open

googio wants to merge 1 commit into
RUC-NLPIR:mainfrom
googio:feat/serply-search

Conversation

@googio

@googio googio commented Sep 19, 2026

Copy link
Copy Markdown

Summary / 概述

Adds serply as an optional keyed search backend: Serply's Google Scholar API,
wired into search.backends exactly like serpbase.

The motivation is corpus coverage for the novelty audit. alphaxiv is the only
papers backend today and it is arXiv-only, so prior art that exists solely in
venue proceedings or a journal is invisible to it. Scholar indexes that
published record, so backends: [alphaxiv, serply] merges and de-duplicates
into a wider prior-art view without changing anything about how the audit runs.

A live query confirms the gap is real rather than theoretical: the top hit for
"iterative hypothesis generation scientific discovery" is an ACL Anthology
paper with no arXiv mirror in the result set.

Linked issue / 关联 issue

Closes #66

Type of change / 改动类型

  • 🐞 Bug fix / 缺陷修复 (fix)
  • ✨ Feature / 新功能 (feat)
  • 📝 Docs / 文档 (docs)
  • ♻️ Refactor / 重构 (refactor)
  • ✅ Tests / 测试 (test)
  • 🔧 Chore / build / CI (chore)

How was this tested? / 如何验证?

Baseline numbers below are from upstream/main in the same virtualenv, so the
comparison is apples to apples.

$ python -m pytest tests/test_search_backends.py
36 passed in 0.20s                      (29 on main, so +7 new)

$ python -m pytest tests/
504 passed, 3 skipped in 7.20s          (main: 497 passed, 3 skipped)

$ python -m ruff check src/ tests/
All checks passed!

$ python -m mypy src/
Found 242 errors in 36 files            (main: 242 -- none introduced)

The unit tests mock requests.get; no network is touched. Separately I ran the
backend against the real API through resolve_backend_names /
build_search_backends to check the response contract:

HTTP 200 | len(articles) = 3 | len(results) = 0
{
  "title": "Hypothesis generation with large language models",
  "link": "https://aclanthology.org/2024.nlp4science-1.10/",
  "description": "Y Zhou, H Liu, T Srivastava, H Mei… - … on NLP for Science  …, 2024 - aclanthology.org"
}

resolve_backend_names(backends=[alphaxiv, serply], key set) -> ['alphaxiv', 'serply']
resolve_backend_names(backends=[alphaxiv, serply], no key)  -> ['alphaxiv']

Two things that response settles, and both are pinned by a test:

  • Scholar results arrive in articles, and the envelope also carries an
    always-empty results list. Keying the wrong array returns zero hits with no
    error, so the array name is passed explicitly rather than guessed.
  • description is the Scholar byline ("Authors - Venue, Year"), not an
    abstract. It maps to snippets because that is the field the contract has,
    but the docs say plainly what it contains so nobody expects abstract text.

Size / 体量

7 files, +195/-22. Slightly above the merged serpbase backend (#65), which
was 5 files, +133/-21, and the extra is accounted for:

Two deliberate omissions, both easy to add if you would rather have them:

  • I did not add serpbase to the .zh docs. It is missing there, but it is
    not my change to make in this PR.
  • I did not touch README.md / README.zh-CN.md. Neither names any keyed
    backend today; both document only the legacy builtin_backend: none | alphaxiv path, so serply has nowhere natural to go. Same for
    examples/research_config.example.yaml, which has no search.backends
    block. serpbase was left out of all three as well.

Happy to split the .zh docs into a follow-up if you prefer this PR at
precedent size.

Notes

  • Serply support is optional, and behavior is unchanged when SERPLY_API_KEY
    is unset: a keyed backend with no key is dropped in resolve_backend_names,
    so backends: [alphaxiv, serply] degrades to [alphaxiv] silently, the same
    as serper and serpbase.
  • Follows SerpBaseBackend in src/core/tools/web/backends.py line for line:
    same _SyncBackend base, same _HTTP_TIMEOUT, same raise_for_status, same
    {url, title, snippets} output, same truncation to max_results. It is a
    GET with an X-Api-Key header where SerpBase is a POST.
  • Key resolution mirrors the existing helpers: serply_api_key on
    SearchConfig, falling back to SERPLY_API_KEY via _WebSearchEnv.
  • Requests identify the caller with User-Agent: Arbor.
  • Tests in tests/test_search_backends.py; network mocked.
  • Docs in docs/search.md, docs/configuration.md and their .zh twins.

Checklist / 检查清单

  • PR title follows Conventional Commits (e.g. fix(config): ...). / PR 标题遵循 Conventional Commits 规范。
  • The change is focused and reasonably small. / 改动聚焦且体量合理。
  • I ran the relevant tests locally (python tests/test_*.py) and they pass. / 我已在本地运行相关测试并通过。
  • Docs / examples updated if behavior or config changed. / 若行为或配置有变更,已同步更新文档/示例(含 README.zh-CN.md)。
    • Both READMEs and the example config are the exception, deliberately: see
      Size for why serply has no place in them today.
  • No secrets, API keys, or tokens are included in the diff. / diff 中不包含任何密钥、API key 或 token。

Disclosure: I work with Serply. Happy to adjust scope, naming, or drop this
entirely if it isn't a direction you want for the project.

Scholar indexes the published record (venue proceedings, journals) that
the arXiv-only alphaxiv backend cannot return, so listing both widens
prior-art coverage for the novelty audit.

Mirrors the serpbase backend: keyed, silently skipped when the key is
absent, key read from serply_api_key or SERPLY_API_KEY.

Closes RUC-NLPIR#66

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: add Serply Google Scholar as an optional papers search backend

1 participant