Google Patents API for Prior Art Search
Google does not document a dedicated Patents REST API. Use Google Patents for interactive discovery and BigQuery for SQL, batch, and semantic search.
Short answer: Google does not publish a dedicated, supported Google Patents REST API for prior-art search. The documented workflow has two layers: use the Google Patents website and Prior Art Finder for interactive discovery, then use Google Patents Public Data in BigQuery when you need SQL, batch processing, embeddings, or semantic search.
That distinction matters. Scraping undocumented endpoints can break without notice and should not be described as an official API. BigQuery is the documented programmatic access path.
What Google officially provides
| Need | Documented route | Best use |
|---|---|---|
| Explore a new invention | Google Patents web search | Freeform phrases, exact phrases, inventor, assignee, date, country, status, language, CPC and family-aware result browsing |
| Generate search terms from disclosure text | Prior Art Finder | Paste a substantial passage and review suggested terms and candidate documents |
| Include papers and other non-patent material | Prior Art Finder with non-patent literature enabled | Bring Google Scholar results into the discovery pass |
| Run repeatable SQL analysis | Google Patents Public Data in BigQuery | Filters, joins, exports, scheduled queries and batch processing |
| Find conceptually similar records | BigQuery embeddings and vector search | Semantic neighbors for abstracts or other patent text |
Google’s help documentation covers freeform and metadata searches and explains that Prior Art Finder can include non-patent literature from Google Scholar. Results can be sorted by relevance or filing date. Google displays one representative from a simple patent family and suppresses other family members, so hit counts are not directly comparable with tools that show every publication. Results are also grouped using Cooperative Patent Classification (CPC) codes.
Start with Google Patents search syntax
- Write down the invention’s essential mechanism, inputs, outputs and constraints.
- Search distinctive technical phrases in quotes, then remove quotes to widen recall.
- Add metadata filters for inventor, assignee, date, country, status or language.
- Use CPC groups to narrow a noisy result set.
- Sort by relevance first, then filing date, and inspect the representative family member.
- Open the publication and read the claims, not only the abstract.
"low-power acoustic event detection" inventor:(Smith) after:priority:20180101 before:priority:20240101 country:US status:GRANT
Operators and exact syntax can change. Confirm the current syntax in Google Patents before putting a query into a production workflow.
Use Prior Art Finder for an invention disclosure
Prior Art Finder works best with a substantial, technical passage rather than a title or a list of keywords.
- Paste the background and detailed-description sections that explain how the system works.
- Review the suggested concepts and remove terms that describe only business context.
- Enable the option to include non-patent literature when papers or standards may be relevant.
- Open promising results and record publication number, priority date, relevant claims and family.
Treat its suggestions as discovery leads. A candidate becomes useful prior art only after you verify dates, public availability, and the disclosure in the actual publication.
Query Google Patents Public Data with BigQuery
Google’s public schema includes publication and application numbers, country and kind codes, worldwide bibliographic data, and US full text. The Google Patents Research Data table adds machine-translated titles and abstracts, extracted top terms, similar documents and forward references. The schema snapshot is dated 2018-11-26, so check the current BigQuery metadata before making completeness or freshness claims.
1. Create a Google Cloud project
- Create or select a Google Cloud project and enable BigQuery.
- Grant the account running queries permission to create jobs and read the public dataset.
- Set a billing project. Public-table bytes and your query processing can incur BigQuery charges under current Google Cloud pricing.
2. Inspect the schema
SELECT
column_name,
data_type,
is_nullable
FROM `patents-public-data.google_patents_research.INFORMATION_SCHEMA.COLUMNS`
WHERE table_name = 'publications'
ORDER BY ordinal_position;
Run this first because field names and available partitions should be checked in the live project.
3. Filter by phrase, date and jurisdiction
SELECT
publication_number,
title,
abstract,
country_code,
publication_date,
priority_date,
cpc
FROM `patents-public-data.google_patents_research.publications`
WHERE publication_date BETWEEN DATE '2018-01-01' AND DATE '2024-12-31'
AND country_code = 'US'
AND (
LOWER(title) LIKE '%acoustic event%'
OR LOWER(abstract) LIKE '%acoustic event%'
)
ORDER BY publication_date DESC
LIMIT 100;
Column names can vary between tables or snapshots. If this query does not compile, use the schema query above and map the live names before proceeding.
4. Export candidates for review
bq query \
--use_legacy_sql=false \
--format=csv \
'SELECT publication_number, title, abstract, priority_date
FROM `patents-public-data.google_patents_research.publications`
WHERE country_code = "US"
LIMIT 1000' > candidates.csv
Run the same query from Python
Install the official BigQuery client and authenticate with Application Default Credentials.
pip install google-cloud-bigquery
from google.cloud import bigquery
client = bigquery.Client(project="YOUR_GCP_PROJECT")
sql = """
SELECT publication_number, title, abstract, priority_date
FROM `patents-public-data.google_patents_research.publications`
WHERE country_code = @country
AND LOWER(abstract) LIKE @term
LIMIT 100
"""
job_config = bigquery.QueryJobConfig(query_parameters=[
bigquery.ScalarQueryParameter("country", "STRING", "US"),
bigquery.ScalarQueryParameter("term", "STRING", "%acoustic event%"),
])
for row in client.query(sql, job_config=job_config).result():
print(row.publication_number, row.title, row.priority_date)
Use query parameters instead of string interpolation. They prevent quoting errors and make repeated searches safer.
Run a BigQuery query from Node.js
npm install @google-cloud/bigquery
const {BigQuery} = require('@google-cloud/bigquery');
const bigquery = new BigQuery({projectId: 'YOUR_GCP_PROJECT'});
const query = `
SELECT publication_number, title, abstract, priority_date
FROM \`patents-public-data.google_patents_research.publications\`
WHERE country_code = @country AND LOWER(abstract) LIKE @term
LIMIT 100`;
const [rows] = await bigquery.query({
query,
params: {country: 'US', term: '%acoustic event%'}
});
console.log(rows);
Build semantic prior-art search with embeddings
Keyword search misses documents that describe the same mechanism with different terminology. Google Cloud’s examples show selecting patent records or abstracts, generating embeddings, and using BigQuery vector functions to retrieve nearest neighbors.
- Select records with usable abstracts and stable identifiers.
- Generate one embedding per abstract with an approved text-embedding model.
- Store the vector beside the publication number and model version.
- Create a vector index when the table size justifies it.
- Embed the invention summary with the same model and distance metric.
- Retrieve nearest neighbors, then rerank by priority date, CPC, jurisdiction and claim relevance.
-- Illustrative pattern: verify model and function names in the current
-- BigQuery documentation before running this in production.
WITH candidates AS (
SELECT publication_number, title, abstract
FROM `patents-public-data.google_patents_research.publications`
WHERE abstract IS NOT NULL
AND country_code = 'US'
)
SELECT publication_number, title, abstract
FROM candidates
WHERE LOWER(abstract) LIKE '%acoustic event%'
LIMIT 100;
The exact embedding model, vector column and VECTOR_SEARCH invocation depend on your Google Cloud configuration. Keep the model name, preprocessing, distance metric and generation date with every vector so results remain reproducible.
Hybrid retrieval is safer than semantic search alone
- Use lexical search for exact components, standards names and rare part numbers.
- Use embeddings to find paraphrases and adjacent terminology.
- Union both result sets, deduplicate by simple patent family, and rerank.
- Read the claims and prosecution history for every material candidate.
Family clustering, CPC groups and hit counts
Google Patents presents one representative from each simple patent family and removes the other family members from the main result list. A semantic or SQL workflow may return every publication unless you explicitly collapse families. Decide whether your analysis counts publications, applications, grants or simple families, and state that choice in reports.
CPC grouping is useful for narrowing a landscape, but a relevant document can sit in an unexpected classification. Start with several neighboring CPC groups and validate classification boundaries against known documents.
Reliability, freshness and cost
| Concern | Practical guidance |
|---|---|
| Freshness | The published schema snapshot is dated 2018-11-26. Check live table metadata and update schedules before claiming current coverage. |
| Geographic coverage | The 2018 launch description said more than 90 million publications from 17 countries. Treat that as a historical launch figure, not a current guarantee. |
| Query cost | Project only needed columns, filter early, use partitions when available, preview bytes processed, and materialize reused subsets. |
| Repeatability | Save SQL, query timestamp, dataset snapshot, model version, filters and family-deduplication rules. |
| Evidence quality | Use search results to discover candidates; verify publication dates, public availability and claim language in the source document. |
Troubleshooting
| Error or symptom | Likely cause | Fix |
|---|---|---|
| No documented REST endpoint | Assuming the website has a supported JSON API | Use the web interface for interactive work and BigQuery for programmatic SQL. |
| Table not found | Wrong project, dataset, table name or region | Open the current public dataset listing and inspect INFORMATION_SCHEMA. |
| Access denied | Missing BigQuery Job User or dataset permissions | Authenticate the active account and grant the minimum project and dataset roles required. |
| Query is expensive | Selecting full text or scanning an unfiltered table | Select only required columns, add date/jurisdiction filters and inspect estimated bytes before running. |
| Too many duplicate hits | Every family member is being returned | Group by a family identifier where available, or deduplicate after export using priority and publication metadata. |
| Semantic results look unrelated | Short or generic query, mismatched embedding model, or no lexical reranking | Use a detailed technical passage, keep preprocessing consistent, combine lexical and vector retrieval, then review claims. |
| Python authentication failure | Application Default Credentials are absent or point to another project | Run gcloud auth application-default login locally or configure the runtime service account. |
| Results appear stale | Assuming the historical schema snapshot is continuously updated | Check current table metadata and record the snapshot date in your report. |
Or skip the browser setup
If your workflow also needs rendered evidence of patent pages, disclosures or comparison reports, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP or PDF. Cookie and consent banners, newsletter popups and chat widgets are removed before capture; bot checks, blank pages and failed loads are not billed. AI agents can call its MCP tools, including take_screenshot, get_page_info and capture_pdf.
See the ScreenshotNeo API documentation for all options. A minimal call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://patents.google.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://patents.google.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://patents.google.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Free usage includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Recommended implementation checklist
- Prototype terms, metadata filters and CPC groups in Google Patents.
- Use Prior Art Finder for a substantial disclosure and include non-patent literature when appropriate.
- Reproduce the search in BigQuery with parameterized SQL.
- Store query text, timestamp, dataset snapshot and family rules.
- Add embeddings for semantic neighbors, then combine them with lexical retrieval.
- Validate priority dates, public availability and claims in the underlying publications.
- Report coverage and freshness with date-qualified language.
FAQ
Is there a free Google Patents API key?
No supported, documented Google Patents REST API key workflow was identified. BigQuery access uses Google Cloud authentication and billing settings.
Can I scrape patents.google.com?
Scraping an undocumented interface is brittle and may conflict with site terms. Prefer the documented web workflow or BigQuery public data.
Does BigQuery contain every patent?
No universal completeness claim should be made. Coverage, jurisdictions and freshness must be checked against the current dataset metadata.
Should semantic search replace keyword search?
No. Use semantic retrieval to expand candidates, then combine it with exact terms, CPC and date filters and manual claim review.
Are Google Patents result counts publication counts?
Usually not. The interface shows one representative from each simple patent family, while a raw data query may return many family members.
Where can I find working examples?
The Google-maintained patents-public-data repository contains examples for landscaping, claim extraction and claim-breadth modeling. It is archived and states that it is not an official Google product, so treat it as an example source rather than a current support channel.


