ScreenshotNeo

BlogAI agents

How to Scrape GitHub and Use Its API with AI Agents

Build a reliable GitHub workflow for AI agents: authenticated API calls, pagination, rate limits, policy boundaries, validation, and safe automation.

By the ScreenshotNeo team1 October 20269 min read

Use GitHub’s documented API for supported data access, then give your AI agent a narrow, authenticated tool with pagination, rate-limit handling, provenance, and human review. Treat website scraping as a separate activity. GitHub’s acceptable-use policy defines scraping separately from API collection, and public visibility alone does not authorize every automated use.

This guide shows a complete workflow: choosing an endpoint, authenticating safely, collecting every page, handling limits, deciding when HTML scraping is appropriate, and validating an agent’s output before it changes anything.

1. Choose the API operation before writing agent code

Start with the REST API endpoint reference. A request consists of an HTTP method and path, plus endpoint-specific headers, authentication, query parameters, and sometimes a body.

Method Typical use
GET Retrieve resources
POST Create resources
PATCH Update selected properties
PUT Replace a resource or collection
DELETE Delete a resource

Define the agent operation in plain language first: for example, “list all open issues in OWNER/REPO and summarize labels.” Map that operation to one endpoint and record the required permission before exposing it to the model.

2. Authenticate with the smallest useful permission

Authenticated requests need a token with the endpoint’s required permissions. GitHub recommends fine-grained personal access tokens for personal use when possible, and GitHub Apps for organizational or on-behalf-of-user integrations. In GitHub Actions, use GITHUB_TOKEN when it is suitable and set workflow permissions explicitly.

Keep credentials out of prompts, logs, source control, browser code, and model-visible tool results. Treat tokens like passwords. Never give an agent a write-capable token for a read-only task.

Required headers

Most endpoints use Accept: application/vnd.github+json. Select an API version with X-GitHub-Api-Version; GitHub’s current documentation example uses 2026-03-10, so confirm the supported version when you implement. Every request must include a valid User-Agent; GitHub says requests without one are rejected.

3. A complete Python collector with pagination and rate-limit handling

The script below lists every issue page, follows GitHub’s returned Link URLs, preserves source URLs, and stops safely when a limit or transient failure occurs.

import os
import time
import requests

API = 'https://api.github.com'
TOKEN = os.environ.get('GITHUB_TOKEN')
OWNER = 'octocat'
REPO = 'Hello-World'

if not TOKEN:
    raise SystemExit('Set GITHUB_TOKEN before running')

session = requests.Session()
session.headers.update({
    'Accept': 'application/vnd.github+json',
    'X-GitHub-Api-Version': '2026-03-10',
    'User-Agent': 'github-agent-example/1.0',
    'Authorization': f'Bearer {TOKEN}',
})

def get_with_backoff(url, params=None, attempts=5):
    delay = 1
    for attempt in range(attempts):
        response = session.get(url, params=params, timeout=30)
        if response.status_code in (200, 201, 204):
            return response
        retry_after = response.headers.get('Retry-After')
        remaining = response.headers.get('x-ratelimit-remaining')
        if response.status_code == 403 and remaining == '0':
            reset = int(response.headers.get('x-ratelimit-reset', time.time() + 60))
            time.sleep(max(1, reset - int(time.time())))
        elif response.status_code == 429 or retry_after:
            time.sleep(float(retry_after or delay))
        elif response.status_code >= 500:
            time.sleep(delay)
        else:
            response.raise_for_status()
        delay = min(delay * 2, 60)
    raise RuntimeError(f'GitHub request failed after {attempts} attempts: {url}')

def parse_next(link_header):
    if not link_header:
        return None
    for part in link_header.split(','):
        pieces = part.split(';')
        url = pieces[0].strip().strip('<>')
        if any('rel="next"' in piece for piece in pieces[1:]):
            return url
    return None

def collect_open_issues(owner, repo):
    url = f'{API}/repos/{owner}/{repo}/issues'
    params = {'state': 'open', 'per_page': 100}
    items = []
    page = 0
    while url:
        page += 1
        response = get_with_backoff(url, params=params)
        batch = response.json()
        for item in batch:
            items.append({
                'id': item['id'],
                'number': item['number'],
                'title': item['title'],
                'html_url': item['html_url'],
                'source_page': page,
            })
        url = parse_next(response.headers.get('Link'))
        params = None
    return items

issues = collect_open_issues(OWNER, REPO)
print(f'Collected {len(issues)} issues from GitHub')

Install the only dependency with python -m pip install requests, export GITHUB_TOKEN, then run the file. The source_page field is application-level provenance so an agent can distinguish a complete traversal from a partial sample.

4. The same workflow with cURL

export GITHUB_TOKEN='your-token'
curl --fail-with-body \
  -H 'Accept: application/vnd.github+json' \
  -H 'X-GitHub-Api-Version: 2026-03-10' \
  -H 'User-Agent: github-agent-example/1.0' \
  -H "Authorization: Bearer $GITHUB_TOKEN" \
  'https://api.github.com/repos/octocat/Hello-World/issues?state=open&per_page=100' \
  -D response-headers.txt \
  -o issues-page-1.json

Inspect response-headers.txt for the Link header and rate-limit headers. Follow the exact rel="next" URL returned by GitHub instead of constructing page numbers yourself.

5. Node.js example for an agent tool

const owner = 'octocat';
const repo = 'Hello-World';
const token = process.env.GITHUB_TOKEN;
if (!token) throw new Error('Set GITHUB_TOKEN');

const headers = {
  Accept: 'application/vnd.github+json',
  'X-GitHub-Api-Version': '2026-03-10',
  'User-Agent': 'github-agent-example/1.0',
  Authorization: `Bearer ${token}`
};

function nextUrl(link) {
  if (!link) return null;
  const match = link.split(',').find(part => part.includes('rel="next"'));
  return match ? match.split(';')[0].trim().slice(1, -1) : null;
}

async function collectIssues() {
  let url = `https://api.github.com/repos/${owner}/${repo}/issues?state=open&per_page=100`;
  const issues = [];
  let page = 0;
  while (url) {
    page += 1;
    const response = await fetch(url, { headers });
    if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
    const batch = await response.json();
    issues.push(...batch.map(item => ({
      id: item.id,
      number: item.number,
      title: item.title,
      html_url: item.html_url,
      source_page: page
    })));
    url = nextUrl(response.headers.get('link'));
  }
  return issues;
}

console.log(await collectIssues());

6. Pagination: never assume the first response is complete

Many list responses are paginated. GitHub’s example shows an issues endpoint returning 30 items by default even though the example repository has more than 1,600 open issues. Follow the response’s Link header, which can include next, prev, first, and last URLs.

  • Use per_page only where the endpoint supports it.
  • The maximum for most endpoints is 100, but defaults and maxima can differ.
  • Stop only when there is no rel="next" link.
  • Store endpoint, page URL, retrieval time, item count, and completion status with the agent’s result.
  • For Octokit users, GitHub documents a paginate() helper that follows supported paginated responses.

7. Give an AI agent a safe tool boundary

Expose narrow functions rather than a generic “call GitHub” tool. A read-only tool might accept owner, repo, state, and a bounded label filter. Validate repository names, cap the number of records, and return structured JSON with provenance.

  1. Show the model the target owner and repository before execution.
  2. Require explicit confirmation before any POST, PATCH, PUT, or DELETE operation.
  3. Use separate credentials for read and write tools.
  4. Validate generated summaries against the fetched records and source URLs.
  5. Log request IDs, endpoint names, status codes, page counts, and truncation reasons without logging tokens.

GitHub’s Terms of Service say: “You are responsible for reviewing, testing, and validating any Output before use.” Treat that as a required review step for code, issue edits, pull requests, and other consequential actions.

8. Rate limits, secondary limits, and efficient collection

GitHub’s published primary REST limits, reviewed in the documentation on 2026-09-29, are 60 requests per hour for unauthenticated public-data requests and 5,000 requests per hour for authenticated users. Search endpoints, GraphQL, installations, and secondary limits can differ.

Read x-ratelimit-remaining, x-ratelimit-reset, and retry-after. When remaining is zero, wait until reset. If retry-after is present, wait that duration. For a secondary limit without those indicators, wait at least one minute, increase delays exponentially after repeated failures, and stop after a bounded retry count. Continuing to request while limited can lead to suspension.

Prefer webhooks when the event model fits. If polling is necessary, poll only as often as needed, request only fields and resources you need, use authenticated conditional requests, and send serial requests where possible. GitHub states that a correctly authorized conditional GET returning 304 Not Modified does not count against the primary rate limit.

9. API collection versus website scraping

GitHub’s acceptable-use policy defines scraping as automated extraction through a bot or web crawler and explicitly says, “Scraping does not refer to the collection of information through our API.” API use is governed by the API terms; it is not blanket permission for every purpose.

The policy lists research use of public, non-personal information when resulting publications are open access and archival use among reasons for using service information. It prohibits spam, including unsolicited email or selling personal information, and requires compliance with GitHub’s Privacy Statement, especially for personal information.

Before automating collection, check the current acceptable-use policy, Terms of Service, privacy requirements, repository licenses, and agreements that apply to your account and deployment. Those documents do not decide every jurisdiction or customer contract.

When HTML scraping is technically necessary

Use the API when the data you need is represented by a documented endpoint. If a public page has information unavailable through the API, first confirm that automated extraction is allowed for your purpose. Respect robots guidance, authentication boundaries, access controls, request pacing, and personal-data obligations. Do not bypass CAPTCHAs, bot checks, paywalls, or other access controls.

10. Or skip the browser setup

If your agent’s job is to capture a GitHub page, README, issue, or dashboard as an image or PDF, ScreenshotNeo provides a single HTTP request instead of maintaining a browser. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://github.com/octocat/Hello-World -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://github.com/octocat/Hello-World"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://github.com/octocat/Hello-World' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also has an MCP server so Claude, Cursor, and other MCP clients can call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

11. Reliability and cost checklist

  • Pin and review the API version used by your integration.
  • Use timeouts and bounded retries; never retry destructive mutations blindly.
  • Persist page-level provenance and detect incomplete traversal.
  • Cache immutable or rarely changing data and use conditional requests.
  • Budget for authenticated requests, search limits, secondary limits, and agent model calls separately.
  • Measure records per request, pages per job, retry rate, and truncation rate.
  • Run a dry-run mode for tools that can write to GitHub.

12. Troubleshooting

401 Unauthorized

The token is missing, expired, malformed, or sent with the wrong scheme. Check the environment variable, use Authorization: Bearer TOKEN, and generate a replacement token without exposing it in logs.

403 Forbidden

The token may lack the endpoint permission, the repository may be inaccessible, or a rate limit may have been reached. Inspect x-ratelimit-remaining, x-ratelimit-reset, and the response body. Reduce permissions only after confirming the endpoint’s requirement.

422 Validation failed

An input is invalid or required parameters are missing. Compare every parameter with the endpoint reference and validate owner, repository, state, labels, and dates before sending.

404 Not Found

The path may be wrong, the repository may not exist, or a private resource may be hidden from a token without access. Verify the URL and token’s repository permissions.

Only the first 30 or 100 records appear

The response is paginated. Follow the Link header until there is no rel="next" URL. Do not infer completeness from a successful first response.

Requests keep returning 429 or secondary-limit errors

Stop sending requests, honor retry-after when supplied, otherwise wait at least one minute and back off exponentially. Replace polling with webhooks or conditional requests where possible.

The agent invents repository facts

Return source URLs and raw fields alongside summaries, require citations to fetched records, and reject claims that cannot be matched to the collected data.

HTML scraping gets blocked

Do not attempt to bypass a bot check or CAPTCHA. Re-evaluate whether a documented API endpoint provides the required information and whether your purpose is allowed by GitHub’s current policies.

13. FAQ

How do I use the GitHub API with an AI agent?

Wrap a specific endpoint in a validated tool, authenticate with the smallest required permission, paginate through every result, return provenance, and require review before actions with side effects.

Should an agent use REST or GraphQL?

Use the interface that matches the resource and query shape. REST endpoints have their documented permissions and limits; GraphQL has separate limits. Confirm current documentation before choosing.

Is scraping GitHub allowed?

There is no universal yes or no. GitHub distinguishes scraping from API collection and lists limited permitted reasons, while also requiring privacy compliance. Check the current policy, terms, licenses, and agreements for your use.

Can I share a GitHub token with an agent?

Only expose a narrowly scoped credential through a server-side tool. Never place it in prompts, client-side code, source control, or model-visible logs.

How can I prove an agent collected everything?

Record the endpoint, exact page URLs, page count, item count, retrieval times, rate-limit events, and whether a next link was absent at completion.