Google Scholar API for Papers, Citations, and PDFs
Google Scholar has no documented public API. Learn the reliable options for papers, citations, structured JSON, and PDF access.
Short answer: Google Scholar is a web search and discovery product. The official help documentation reviewed for this guide does not document a public Google Scholar API. If you need structured JSON, you must use a third-party extraction service such as SerpApi, or choose a scholarly graph API such as Semantic Scholar when normalized paper and citation data is more important than matching Google’s search results.
That distinction matters. A vendor’s Google Scholar endpoint is the vendor’s extraction layer, not a Google API promise or endorsement. Coverage, freshness, limits, reliability, pricing, and acceptable use must be checked in the current vendor documentation.
What Google Scholar supports directly
Google Scholar can search scholarly literature, find authors and exact titles, restrict results by date, follow Cited by and Related articles, inspect all versions, export formatted citations, and create email alerts. Google’s help page documents these user-facing features: Google Scholar Search Help.
- Use
author:for author searches. - Put an exact paper title in quotation marks.
- Use the date controls to restrict or sort recent results.
- Open Cited by to discover citing works.
- Use All versions to find repository copies, preprints, or publisher pages.
- Export citations in formats such as BibTeX, EndNote, RefMan, or RefWorks from a result’s quotation-mark menu.
- Create an alert for a query when you need new-result notifications.
Does Google Scholar have an official API?
There is no documented public API in the official help material used for this article. Treat claims such as “Google Scholar API” as one of two things:
- A third-party SERP extraction API. The provider retrieves Scholar result pages and returns structured fields.
- An academic graph API. The provider maintains its own scholarly metadata and citation graph, which may not reproduce Google’s ranking or index.
Do not describe a third-party endpoint as Google’s official API unless Google publishes direct documentation that says so. Also review current Google and provider terms with your legal team before automating collection at scale.
How to get Google Scholar results as JSON
SerpApi documents a google_scholar engine. Its documentation describes a required q query for normal searches, citation and cluster modes, date-range parameters, localization, and JSON, HTML, or Markdown output. See the Google Scholar API documentation for current parameters and limits.
cURL
curl -G 'https://serpapi.com/search.json' \
--data-urlencode 'engine=google_scholar' \
--data-urlencode 'q=large language model evaluation' \
--data-urlencode 'api_key=YOUR_SERPAPI_KEY'
Python
import requests
params = {
'engine': 'google_scholar',
'q': 'large language model evaluation',
'api_key': 'YOUR_SERPAPI_KEY',
}
response = requests.get('https://serpapi.com/search.json', params=params, timeout=30)
response.raise_for_status()
data = response.json()
for result in data.get('organic_results', []):
print(result.get('title'))
print(result.get('link'))
print(result.get('publication_info'))
print(result.get('inline_links', {}).get('cited_by'))
Node.js
const params = new URLSearchParams({
engine: 'google_scholar',
q: 'large language model evaluation',
api_key: 'YOUR_SERPAPI_KEY'
});
const response = await fetch(`https://serpapi.com/search.json?${params}`);
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const data = await response.json();
for (const result of data.organic_results ?? []) {
console.log(result.title);
console.log(result.link);
console.log(result.publication_info);
console.log(result.inline_links?.cited_by);
}
These examples use the endpoint and fields documented by SerpApi. Confirm the current endpoint, authentication method, response shape, quota, and pricing before deploying.
Which fields can a Scholar extraction API return?
SerpApi’s organic-result documentation describes fields including the result title, link, publication information, snippet, resources, cited-by information, versions, cached-page link, and related-page link. See its organic-results reference.
| Field or link | Typical use | Important limitation |
|---|---|---|
| Title and link | Store a result and open the source | Ranking and availability can change. |
| Publication information | Display authors, venue, and date | Normalize and validate before deduplication. |
| Snippet | Preview relevance | It is not a complete abstract. |
| Cited-by data | Follow citing works or show a count | Counts are time-dependent and source-specific. |
| Resources | Find PDF or HTML links | A resource may be absent, paywalled, or removed. |
| Versions | Locate alternate copies | Different versions may have different rights and content. |
Design your parser to tolerate missing fields. A result can have a title but no PDF resource, a citation link but no numeric count, or multiple versions with different hosts.
How to find papers and citation counts programmatically
- Send a narrowly scoped query instead of a very broad phrase.
- Persist the provider’s result identifier, title, authors, publication data, source URL, and retrieval timestamp.
- Read citation information from the documented response object rather than scraping text from your own rendered page.
- Store citation counts with a timestamp. They change as Scholar’s index changes.
- Deduplicate by DOI when present, then compare normalized title, authors, and year. Do not assume two identical titles are the same work.
- Keep the original result and source links so a person can verify the record.
For citation discovery, SerpApi documents citation-based searches and cluster or all-version operations. Use those modes when the provider’s current documentation says they fit your query; do not assume every result has a stable identifier or complete citation graph.
How to download PDFs from Google Scholar results
Scholar does not guarantee a downloadable PDF for every result. Google says it may link to an open copy, a repository, a preprint, a publisher copy, or a library subscription. Its help page also states: “Abstracts are freely available for most of the articles. Alas, reading the entire article may require a subscription.”
When a structured response includes a PDF or HTML resource, treat it as a link to another site and check the response status, content type, redirects, and access rights.
Python PDF fetch with validation
from pathlib import Path
from urllib.parse import urlparse
import requests
pdf_url = 'https://example.edu/paper.pdf' # use a resource URL returned by your provider
parsed = urlparse(pdf_url)
if parsed.scheme not in {'http', 'https'}:
raise ValueError('Only HTTP(S) URLs are accepted')
with requests.get(pdf_url, stream=True, timeout=60, allow_redirects=True) as response:
response.raise_for_status()
content_type = response.headers.get('content-type', '').lower()
if 'pdf' not in content_type and not response.url.lower().endswith('.pdf'):
raise ValueError(f'URL did not return a PDF: {content_type}')
with Path('paper.pdf').open('wb') as output:
for chunk in response.iter_content(chunk_size=1024 * 256):
if chunk:
output.write(chunk)
Do not bypass a login, subscription, robots restriction, or access control. If the link requires a library session, send the user to the library or publisher page instead of promising an automated download.
Semantic Scholar versus a Google Scholar extraction API
Semantic Scholar’s Academic Graph API documents paper and author data, citation-related endpoints, and an openAccessPdf field. It is an alternative data source, not a guarantee that you will receive Google’s ranking or exactly the same result set.
| Requirement | Better starting point |
|---|---|
| Match what Google Scholar currently displays | A Scholar extraction service; verify coverage and terms. |
| Persistent paper and author metadata | An academic graph such as Semantic Scholar. |
| Follow Google’s cited-by or related-result workflow | A provider that documents those Scholar operations. |
| Find open-access PDF metadata | Use documented open-access fields, then validate the linked file. |
Compare representative queries across your disciplines, languages, document types, and publication dates. The reviewed Semantic Scholar documentation does not provide enough information to claim broader coverage, lower latency, or higher reliability than a particular Scholar provider.
Production checklist
- Record the provider, endpoint version, query, locale, and retrieval time.
- Pin response parsing to documented fields and handle absent or renamed fields.
- Cache repeat queries where permitted, with a retention policy that fits your use case.
- Implement exponential backoff for transient errors and a bounded retry count.
- Track quota usage, latency, HTTP errors, empty responses, and PDF validation failures.
- Keep source URLs and a human review path for high-impact bibliographies.
- Recheck current pricing, quotas, rate limits, terms, and privacy rules before launch.
- Obtain permission for content storage and redistribution when a publisher or repository requires it.
Troubleshooting
There is no official API key
Cause: Google Scholar’s public help describes the website, not a public developer API.
Fix: Use the website manually, select a documented third-party extraction service, or use an academic graph API.
The API returns an empty result set
Cause: The query is too narrow, the locale differs, a date filter excludes results, or the provider changed its response.
Fix: Log the exact request, test a known broad query, remove filters one at a time, and inspect the provider’s current status and schema.
A result has no PDF
Cause: The underlying record has no accessible full-text link, or access requires a library or publisher subscription.
Fix: Check all versions, repository links, and your library resolver. Preserve the abstract and citation record without claiming full-text availability.
The citation count differs from another service
Cause: Indexes update at different times and apply different deduplication rules.
Fix: Label the source and timestamp, and avoid presenting counts from different indexes as interchangeable.
Downloads save HTML instead of a PDF
Cause: The link redirected to a login page, an error page, or a publisher landing page.
Fix: Check status, final URL, content type, and the first bytes before saving.
Performance, reliability, and cost
Measure your own workload. Query complexity, locale, provider queueing, retries, and downstream PDF hosts all affect latency. Use bounded concurrency, caching, and backoff rather than firing unthrottled requests. Keep raw responses only as long as your terms and privacy policy allow.
There is no verified topic-specific benchmark in the research for coverage, uptime, accuracy, or latency. Compare current vendor quotas and prices directly, then estimate total cost from searches, citation lookups, retries, storage, and PDF transfer. A result-page extraction service and an academic graph may charge or limit different operations.
Or skip the browser setup
If your application also needs a clean visual capture of a Scholar result page, article landing page, or PDF view, ScreenshotNeo provides a screenshot API and MCP server. It is separate from Scholar data extraction: it captures the rendered page rather than returning bibliographic JSON.
One request is enough:
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://scholar.google.com/scholar?q=large+language+model+evaluation -o shot.webp
See the ScreenshotNeo API documentation for options. The same request in Python is:
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://scholar.google.com/scholar?q=large+language+model+evaluation'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://scholar.google.com/scholar?q=large+language+model+evaluation' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, newsletter popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status.
- An MCP server lets Claude, Cursor, and other MCP clients call
take_screenshot,get_page_info, andcapture_pdf. - The free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account.
FAQ
Can I call Google Scholar with a normal REST URL?
You can use Scholar’s website manually, but the reviewed official material does not document a public REST API. A structured response requires a separate provider or data source.
Is a PDF link proof that I may redistribute the paper?
No. A link only indicates where the underlying result points. Check the repository, publisher, library, and license terms.
Should I store citation counts permanently?
Store them with the source and retrieval timestamp. Counts can change, so treat them as observations rather than immutable metadata.
When should I choose an academic graph API?
Choose one when you need normalized paper, author, citation, or open-access metadata and do not need to reproduce Google’s current ranking.


