ScreenshotNeo

BlogGuides

How API Credits Work

API credits are provider-defined usage units. Learn how calls consume them, why balances run out, and how to control cost, quotas, limits, and expiry.

By the ScreenshotNeo team1 October 20267 min read

API credits are provider-defined units that represent available usage. A credit may mean prepaid money, a request, tokens, compute time, or another allowance. There is no universal conversion: one provider’s credit can cover one request, while another provider deducts credits according to model, input size, output size, or monetary cost.

To understand any API credit balance, find five definitions in that provider’s billing documentation:

  1. Meter: what is measured, such as requests, tokens, dollars, seconds, or operations.
  2. Unit price: how much one measured unit costs.
  3. Allowance: the number of units included or prepaid.
  4. Reset or expiry: when unused balance becomes available again or disappears.
  5. Controls: quotas, spend limits, and rate limits that can reject requests independently of balance.

Credits, quotas, spend limits and rate limits

These terms describe different controls:

Term What it controls Typical failure
Credits Available prepaid or paid usage Insufficient balance or payment required
Quota Aggregate allocation for an account, project or organization Quota exceeded even though another balance may exist
Spend limit Maximum billable amount over a billing period Requests blocked after the cap is reached
Rate limit Requests or tokens allowed during a time window HTTP 429 or throttling during bursts

A request can fail for a rate-limit violation while credits remain. Conversely, a request can be within the rate limit but fail because the account has no prepaid balance. OpenAI documents rate limits and 429 troubleshooting separately from prepaid billing and spend limits; Google Gemini likewise describes prepaid credits deducted according to usage cost. Check the provider’s current documentation before relying on a particular rule.

How providers calculate consumption

Consumption depends on the provider and endpoint. Common meters include:

  • Requests: one successful or attempted call counts as one unit.
  • Input and output tokens: language APIs often meter both separately.
  • Monetary spend: a credit balance is reduced by the calculated price of an operation.
  • Operation type: image generation, file processing, searches and batch jobs may have different prices.
  • Payload and response size: larger inputs or outputs can consume more tokens or bandwidth.
  • Retries: a retry may be metered as another request, even when the first response was an error.

Therefore, “one credit equals one call” is correct only when the provider explicitly says so. A useful accounting equation is:

total_cost = sum(operation_units × unit_price) + applicable_overages

For token-metered APIs, track input and output independently:

token_cost = (input_tokens × input_price) + (output_tokens × output_price)

Do not infer a conversion from the word credit. Read the endpoint’s pricing table and billing terms.

Runnable local credit tracking

The following examples show a provider-neutral ledger. They do not call a billing API; they record the meter values returned by your provider so you can reconcile usage by project, key, model and endpoint.

Python

from dataclasses import dataclass
from decimal import Decimal

@dataclass
class Usage:
    input_units: int = 0
    output_units: int = 0
    requests: int = 0

    def add(self, input_units=0, output_units=0, requests=1):
        self.input_units += input_units
        self.output_units += output_units
        self.requests += requests

usage = Usage()
usage.add(input_units=1200, output_units=350)
usage.add(input_units=800, output_units=190)

input_price = Decimal("0.000001")
output_price = Decimal("0.000002")
cost = (Decimal(usage.input_units) * input_price +
        Decimal(usage.output_units) * output_price)

print({
    "requests": usage.requests,
    "input_units": usage.input_units,
    "output_units": usage.output_units,
    "estimated_cost": str(cost)
})

Node.js

const events = [
  { inputUnits: 1200, outputUnits: 350 },
  { inputUnits: 800, outputUnits: 190 }
];

const totals = events.reduce((sum, event) => ({
  requests: sum.requests + 1,
  inputUnits: sum.inputUnits + event.inputUnits,
  outputUnits: sum.outputUnits + event.outputUnits
}), { requests: 0, inputUnits: 0, outputUnits: 0 });

const inputPrice = 0.000001;
const outputPrice = 0.000002;
const estimatedCost = totals.inputUnits * inputPrice +
  totals.outputUnits * outputPrice;

console.log({ ...totals, estimatedCost });

cURL

curl -sS https://api.example.com/usage \
  -H "Authorization: Bearer $API_KEY" \
  -H "Accept: application/json"

Replace the URL and response fields with the provider’s documented usage endpoint. Never assume an undocumented endpoint or field name.

Why credits run out faster than expected

  1. Token growth: prompts, conversation history and requested output became larger.
  2. Retries: timeouts or 429 responses triggered automatic retries that were also metered.
  3. High-cost operations: a different model or endpoint has a higher unit price.
  4. Parallel jobs: a burst consumed the shared project or organization allowance.
  5. Hidden consumers: another service, environment or API key uses the same balance.
  6. Failed assumptions: errors, partial work or asynchronous jobs are charged according to provider rules.
  7. No reset: the balance is prepaid and does not automatically renew.

Log every request with a timestamp, organization or project, key identifier (never the secret), model, endpoint, retry count, measured units and provider response ID. Compare that log with the provider’s usage dashboard.

Controlling API credit spend

Set a budget policy

  • Set alerts below the expected monthly balance.
  • Use hard spend limits where the provider supports them.
  • Assign separate projects or keys to production, staging and experiments.
  • Restrict keys to the smallest required scope and rotate them when ownership changes.

Control request volume

  • Cap retries and use exponential backoff for 429 responses.
  • Queue bursts and pace requests to the documented rate limit.
  • Deduplicate identical work and cache deterministic responses.
  • Set maximum input and output sizes before sending a request.

Choose the right meter

When comparing providers, record the meter, unit price, included or prepaid allowance, overage behavior, reset or expiration policy, spend controls, rate limits, latency and whether balances are shared across projects or billing accounts. Compare equivalent workloads rather than the advertised credit count.

Do API credits expire or roll over?

There is no industry-wide rule. Credits can reset, expire, roll over, or remain available until consumed, depending on the provider, plan and geography. Some subscriptions renew an allowance each billing period; prepaid balances may follow a separate expiration policy. Confirm the exact terms before promising customers that unused credits persist.

Reliability and performance considerations

  • Separate throttling from balance: a 429 should trigger paced retry logic, while an exhausted balance should stop retries and alert an operator.
  • Make retries safe: use idempotency controls when documented, otherwise deduplicate with your own request ID.
  • Budget for variance: token counts and operation prices can change with payloads, models and endpoint options.
  • Measure latency by operation: a cheaper request that requires many retries can cost more and take longer than a single reliable call.
  • Reconcile periodically: compare local logs with the provider dashboard and investigate unexplained differences.

Applying credit accounting to screenshot APIs

Screenshot services illustrate why the meter must be explicit. A service may count requests, successful captures, rendered pages, image bytes or PDF pages. Before integrating one, verify what happens for failed navigation, bot checks, cache hits, retries, asynchronous jobs and bulk requests.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its API is a single GET request; the response can be PNG, JPEG, WebP or PDF. The service accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture, with each step configurable. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for request options. This call captures a page as WebP:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture with lazy images loaded, CSS element capture, dark mode, device presets and custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. An MCP server provides take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.

Troubleshooting

Symptom Likely cause Fix
HTTP 429 Rate limit or burst exceeded Read the rate-limit headers, reduce concurrency and retry with exponential backoff.
Balance appears positive but calls fail Project quota or spend limit reached Check the project, organization and billing account limits separately.
Credits disappear unexpectedly Shared key, retries or larger payloads Group usage by key, endpoint, model and retry count; compare with the provider dashboard.
Credits reset unexpectedly Allowance is subscription-based Read the plan’s reset and rollover terms and export usage before the reset.
Local totals do not match billing Different rounding, delayed events or unlogged consumers Use provider response IDs, account for asynchronous jobs and reconcile over the same time window.
Automatic retries increase cost Retry policy treats all errors alike Retry transient 429 and 5xx responses only; stop on authentication, validation and exhausted-balance errors.

FAQ

How many API calls does one credit cover?

Only the provider can define that. It may be one request, a token allowance, or a monetary amount divided by operation price.

Are failed calls free?

Not universally. Some providers charge for attempted or partially processed work; others charge only successful operations. Check the endpoint’s billing terms.

Can I transfer credits between projects?

Transfer and sharing rules are provider-specific. Treat balances as scoped to the documented account, organization or project.

Should I buy more credits or raise a limit?

First identify the failure: exhausted balance requires funding, a quota requires allocation approval, and a rate limit requires pacing or a higher limit.

What is the safest way to estimate a monthly balance?

Measure a representative workload, include retries and peak bursts, then add an operational buffer. Recalculate when models, payloads or endpoints change.